Diarized Transcripts
Two people on a call, two channels on the recording. Who said what is read off the file rather than guessed from the voices, so a coaching note never quotes the customer back at the rep who did not say it.
Transcript, speaker by speaker
two channels- 0:01Marcus Hale
Adoptiv service desk, this is Marcus.
- 0:04Dana Whitlow
Hi, it is Dana Whitlow at Meridian Freight.
- 0:09Dana Whitlowsame channel, 1.4 s gap
I am calling about the invoice that went out on Tuesday.
- 0:14Marcus Hale
Let me open that. Invoice ending 4192?
- 0:18Dana Whitlow
That is the one.
- 0:21Marcus Hale
It went out at the old rate. I will reissue it today.
Product figures from the platform’s own defaults - not customer averages
How it works.
Channel beats a guessed speaker every time
Where a provider returns both a channel tag and an inferred speaker index, the channel wins. Preferring the inference over a channel that was right there is exactly the fault that swapped agent and customer on calls carrying the answer.
Nothing is left to voiceprint guessing
The four engines that can do it are asked for the channels and never for voiceprint separation. Two transcribe the channels independently with speaker labeling switched off, one sets channel diarization with the labels fixed in channel order, and one asks for multichannel output kept apart while refusing to diarize.
A row closes on a change or a pause
Engines that answer in words rather than phrases are regrouped here: the segment holds while the channel holds and the gap stays under 1.2 seconds. Cross either and it ends, which is why one person can occupy two rows in succession.
One control fixes a call that came out backward
Where the convention was reversed on a particular recording, a single toggle flips the labels for that call and no other. Offsets are stored with the segments, so the flip costs a redraw rather than another run.
One moment in every read.
Every read passes through the same seven. Diarized Transcripts is the lit one, and everything either side of it is a different page in this category.
- 01Source
the call or the thread it reads
- 02Transcribe
audio into words, with speakers
- 03Read
the pass over the whole of it
- 04Judge
the score, the sentiment, the intent
- 05Extract
the fields and follow-ups pulled out
- 06Write
what lands back on the record
- 07Review
a person checking the machine
The specifics.
8 facts- Convention
- Channel 0 is the agent leg and channel 1 the customer. Any other index renders under the generic label Speaker
- Segment break
- A channel change, or on word level engines the same channel pausing past 1.2 seconds
- Each segment
- Channel, text, an offset from the start, and a length where the provider gives one
- Mono audio
- Falls back to plain text with legacy prefix detection. No speaker is inferred
- Not every engine
- Two of the wired speech paths return no channel tags, and those calls take the fallback
- Searching it
- One screen has it. The voicemail list carries a transcript search running on Postgres full text; the call list searches names, agents and numbers and never the words
- Correcting a word
- Not offered. What the engine returned stands; masking strong language is a setting
- Not the same as
- This is the record of who said what. Reading it is what the analysis pass does
What it reads, and what reads it.
Nothing here invents its input. These are where the material comes from, and where the verdict goes afterward.
More in Intelligence
12 capabilitiesAI that proposes edits to the record - an insight becomes a field once you accept it.
Nothing on an existing lead moves by itself. What the call gave up arrives as proposed edits, each with the value it would replace and the words it came from.
Finding yesterday's bad calls is a filter, not an afternoon. Positive, neutral or negative on every analyzed call, kept in a field of its own you can sort on.
How the agent sounded, how the customer sounded, one word for the pair. Chosen words rather than a fixed list, so a call can close politely and grudgingly.
A month of calls gets reviewed by whoever had time. Every call carries five numbers instead, and the overall is judged in its own right rather than averaged.
Nobody re-listens to check the disclosure went out. Every call is scored against what compliance means on your floor, with a missing one excluded, not passed.
What the customer asked for and what your agent promised, kept as two lists, because they are two different obligations and only one of them is yours to keep.
A promise made out loud is worth nothing until it sits on a day. Callbacks and meetings come off the call with the time said, real the moment you accept.
A collections call and an admissions call are not judged the same way. Fifteen verticals ship knowing the difference. What comes back still reads the same.
A thread is read only when somebody on it is one of your leads. No match and nothing is sent anywhere, so mail that is not about a customer is left alone.
A draft or a rewrite that lands in the composer and stops there. The subject is a suggestion, every word editable, and nothing leaves until a person sends it.
The field group you would otherwise build one field at a time. Describe what you track or paste a spreadsheet's columns, then approve or refuse each field.
Type the list you want and it builds the filter and the sort. Then it says in one sentence what it understood: a filter nobody typed has to say what it did.
The rest of the platform.
Five more categories, all on the same record and the same bill. Each card names three of its capabilities, so you can tell from here whether it is worth opening.
Intelligence
See diarized transcripts on your own floor.
Thirty minutes, your numbers and your data. We will set diarized transcripts up live and you can decide from the thing itself rather than from this page.
14-day trial · no card · migration included