The problem
Off-the-shelf transcription merges speakers on noisy walk-and-talk audio — and every downstream AI judgment inherits the error. When four people talk, two get merged into one.
The system
A speaker-embedding pipeline that matches transcript words to voiceprints, splits wrongly-merged turns, and abstains when attribution is genuinely ambiguous — because a wrong speaker label is worse than no label.
How it's built
- ResNet-based speaker embeddings aligned to word-level timestamps
- Serverless GPU deployment (Cloud Run L4) — scales to zero between batches
- Merge-splitter for narration-contaminated turns, gated by lexical pre-filters
- Principle encoded in code: cannot attribute → exclude, never guess
Delivery
Research-to-production in short cycles: offline eval on labeled tours first, then a live shadow lane, then default-on.
Results
- Attribution errors reduced to low single-digit percent of turns on eval audio
- GPU cost per recording measured in cents, not dollars
- Downstream coaching quality complaints traced to diarization dropped
Client anonymized by industry. Detailed numbers and references available on a call.