← All work
Case study · in production Proptech SaaS — leasing operations Client: Harborview Software *

Every sales tour, scored and coached

AI coaching on real conversations

SaaS · applied AI2 engineersIn production, multi-tenant
Download the case study (PDF) * Client name changed. Engagement details anonymized under NDA; detailed numbers and reference calls available on request.

About Harborview Software

Harborview Software records and analyzes in-person leasing tours for multifamily operators. Thousands of tours a month flow through its platform, and its customers coach, staff, and reward leasing teams based on what it reports — so a wrong claim in a coaching report is not a cosmetic bug, it is a personnel decision made on bad evidence.

The challenge

Operators were sitting on hours of recorded tours nobody had time to hear. Managers wanted every tour scored against their own playbook; associates wanted feedback that quoted what they actually said.

Early LLM prototypes failed in the worst possible way: they scored fluently and occasionally invented quotes that were never said. On poor audio, the transcription layer itself fabricated words. And when four people speak on a walking tour, off-the-shelf diarization merged speakers — attributing one person's words to another. In a product that judges people's work, every one of those errors is disqualifying.

The solution

SurgeX Labs built the pipeline around one rule: no claim without evidence. Speech-to-text runs with per-word confidence — low-confidence spans are bracketed, never asserted. A voiceprint layer matches transcript words to speakers and abstains when attribution is genuinely ambiguous, because a wrong speaker label is worse than none.

Scoring runs against a rubric engine versioned per customer, and every coaching bullet must anchor to a real transcript span — a citation-verification layer rejects anything the model cannot point to. Compliance-sensitive judgments on low-confidence audio are held for human review instead of auto-failing staff. Consent is enforced before any analysis begins.

Every change to prompts or models replays against evaluation sets before it ships. The system went from prototype to multi-tenant production with a two-engineer pod, released in guarded increments behind flags.

The evidence-gated coaching pipeline

TOUR AUDIOASR + PER-WORDCONFIDENCElow confidence → bracketedVOICEPRINTDIARIZATIONambiguous → abstainRUBRIC SCORING(LLM)versioned per customerCITE-VERIFYGATEno claim without evidenceSCORED, CITEDCOACHINGEVAL SETS REPLAY ON EVERY PROMPT OR MODEL CHANGE · CONSENT ENFORCED BEFORE ANALYSIS BEGINS

How it's built

The results

0 fabricated citations in a 371-tour audit
0/36 speaker-role swaps in the attribution audit
[N] tours analyzed monthly across [M] communities [confirm]
2 engineers, prototype to multi-tenant production

* Client name changed. Engagement details anonymized under NDA; detailed numbers and reference calls available on request.

[Pull quote pending client approval — e.g. "Coaching only works if the associate believes the quote. Zero fabricated citations is the whole product."]

— [Name], [Title], Harborview Software

Want a system like this?