Skip to content
All posts
Guides

What speaker diarization is, and the two ways it gets 'who spoke when' wrong

Your transcript says Speaker 3 and there were only two people in the room. That is not a random glitch. Automatic speaker labelling fails in two recognisable shapes, each with its own cause, and knowing which one you are looking at tells you exactly what to fix.

Kalima Team8 min read

Speaker diarization is the part of a transcription system that works out who spoke when. It does not know anyone's name and it never did. It listens for changes of voice, cuts the audio into turns, converts each turn into a numeric fingerprint, then groups similar fingerprints together, which is why the transcript comes back as Speaker 1, Speaker 2 and Speaker 3 instead of one undivided block of text.

Everything that later looks wrong about those labels comes from that grouping step, and it goes wrong in two recognisable shapes. Once you can tell them apart, you know which passages to check and which you can repair in a couple of minutes.

How the labels get made in the first place

Vendor documentation describes the same four stages, with small differences in naming. Deepgram's explainer (accessed September 2026) sets them out as detection, segmentation, representation and attribution:

  1. Find the speech. Voice activity detection marks which parts of the audio hold a human voice at all, and which are silence, keyboard noise or an air conditioner.
  2. Cut at the voice changes. A model finds the points where the voice characteristics shift and splits the audio into turns there.
  3. Fingerprint each turn. Each turn becomes a speaker embedding, a vector of numbers encoding pitch range, timbre and other acoustic habits.
  4. Group the fingerprints. Similar vectors are clustered, and each cluster gets a number.

Nothing in that chain involves knowing who you are. The label is a cluster with a number printed on it, and the number is assigned in the order the clusters appear. That single fact explains most of what follows.

The two failure shapes: merge and split

Almost every complaint about speaker labels is one of these two, and they have opposite causes and opposite fixes.

MergeSplit
What you seeTwo people share one label. A dialogue reads like a monologue.One person appears as Speaker 2 and Speaker 4.
Why it happensTheir fingerprints are too close to separate: similar pitch and accent, the same room, the same microphone, short turns.Their fingerprint moved during the recording, far enough to look like a new person.
Typical triggerTwo colleagues of the same age and region on one table mic. Back-channels such as "yep" absorbed into whoever was already speaking.Someone raises their voice, gets animated, leans back from the microphone, swaps a headset for speakerphone, or a dialled-in line changes quality.
What fixes itSeparate the voices manually, or re-run the grouping with the correct number of people.Reassign the stray passages, and use one name for both labels.

AssemblyAI's own write-up of the hard cases (accessed September 2026) names the same pair as speaker count errors: systems "merge two people into one, especially with similar-sounding voices" or "split one person's changing voice into multiple speakers".

The practical takeaway is that the label is a guess about acoustics, not a fact about people. A split tells you something changed in the sound. A merge tells you two things sounded alike.

Overlapping speech causes most of the errors

There is one cause that outweighs the rest. Traditional diarization assigns each moment of audio to exactly one cluster, so when two people speak at the same time, one of them has to lose. Federico Landini's 2024 review From Modular to End-to-End Speaker Diarization puts it plainly: until recently every competitive approach was modular, those systems reached state-of-the-art results in most scenarios, and they "had major difficulties dealing with overlapped speech". Newer end-to-end models are built to handle several voices in the same instant, and research such as Speaker Embedding-aware Neural Diarization exists specifically because meeting audio is so heavily overlapped that the one-voice-at-a-time assumption breaks down.

For a meeting owner the consequence is concrete. Crosstalk is where the labels are least trustworthy, and crosstalk clusters around the moments you care about: the interruption, the objection, the quick agreement while someone else is still finishing a sentence. So do not check a transcript top to bottom. Jump to the places where turns are short and fast, because that is where merges live.

A note on the number vendors quote. Diarization error rate, or DER, is the share of speech time that is either missed, falsely marked as speech, or attributed to the wrong speaker. Because it is measured in seconds, a system can drop every one-word "yes, agreed" and still score well: AssemblyAI describes a case that scored 15.1 % DER while the word-level metric for the same output was 30.7 % cpWER. A DER figure therefore tells you very little about whether your minutes will name the right person. Kalima publishes no diarization percentage, for the same reason we publish no accuracy percentage, and how to read a transcription accuracy claim goes through what to ask for instead.

Why the numbers change between recordings

Clustering happens inside one recording. It groups the voices that are present against each other, and that is all it does. It is not identification, because there is no stored profile of you to compare against.

So in Kalima, speaker numbers stay unique across a session, including when you stop and start several recordings in it, but automatic diarization does not carry identity across takes. The same person can come back as a different number in the next recording, and that is the system working as designed rather than losing track. The colour palette behaves the same way: it holds eight colours and repeats from the start for a ninth speaker, while the numbers stay unique. Both details are in the speakers and diarization help page.

Repairing labels that are already wrong

You do not need to re-record. In Kalima, open the finished recording, choose Check speakers, then Check voices. Kalima listens to the recording again, groups the passages by voice, and shows you samples to play before anything is applied.

Three things are worth knowing before you run it:

  • The review runs on your device. On the first run a 38 MB voice model downloads and then stays on that device, and no audio leaves it.
  • It needs at least two passages with speech to have something to compare.
  • It takes about 40 seconds per hour of recording, so a two-hour workshop is done in a couple of minutes.

In the review you can correct the number of people, which is the direct fix for both failure shapes: lower it when one person has been split, raise it when two have been merged. Then name each voice, listen to the passages the grouping was unsure about, and reassign individual passages before you apply the result.

Renaming is separate and simpler. Click the speaker badge in any segment header and type a name of up to 50 characters; every segment attributed to that voice updates at once. Names you have already entered appear as quick-pick chips, which is how you give one person the same name across two takes that numbered them differently. Only the session owner can rename speakers. Whatever you set flows into the live view, shared sessions and every export, including SRT, VTT, JSON, text and PDF.

Record in the mode that separates voices best

If the final transcript matters more than watching text appear, use Post-Process+ rather than Realtime. It records first and produces the transcript a few minutes after you stop, and it gives the highest accuracy and the best speaker separation of the two modes. Both cost the same transcription time, so this is a choice about the output, not about your budget. Realtime is right when you need live captions, live translation or viewers following along, and a running Post-Process+ take can be switched to Realtime once, though not back. See recording modes.

In a room, the microphone decides

Diarization can only separate what the microphone captured as separate. One laptop mic in the middle of a table records everyone through a different distance and a different amount of room reverberation, which is raw material for both failure shapes at once.

  • Put the microphone closer to the people, or use one per person where the setup allows it.
  • Ask remote participants to use a headset rather than a speakerphone, and to stay on one device.
  • Leave a beat between turns. A short pause is what the segmentation step looks for.
  • Do not move a shared microphone mid-discussion. The change in level looks like a change of speaker.

If you are transcribing a conversation that includes other people, tell them before you start and get their agreement. In Germany the spoken word is protected by § 201 StGB, and the safe path anywhere is consent from everyone present plus a check of your organisation's own rules and the meeting platform's terms.

Frequently asked questions

Why does my transcript say Speaker 3 when there were only two people?

That is a split. One person's voice changed enough during the recording for the clustering step to read it as a new person, usually after they got louder, moved relative to the microphone, or switched device. Run Check voices, set the number of people to two, and reassign the stray passages. Giving both labels the same name also merges them in the transcript and in every export.

Why did the same person get a different number in the second recording?

Because diarization groups voices within one recording and does not recognise identity across takes. Numbers stay unique across the whole session, so the same person may land on a different number in a later take. Reuse the same name from the quick-pick chips and the transcript reads as one continuous conversation again.

Can transcription software recognise a specific person by their voice?

Diarization cannot, and Kalima does not. Diarization only answers whether two passages were spoken by the same voice within this recording. Matching a voice to a named individual is a different technology, speaker identification, which requires an enrolled voice profile to compare against. In Kalima the names on a transcript are typed in by the session owner.

Does diarization still work with more than eight speakers?

Yes. The eight-colour limit is only the palette, which repeats from the start for the ninth speaker onward while the numbers stay unique. The grouping itself does get harder, because more voices means more chances that two sit close together in the fingerprint space. Panels and conferences are the clearest case for one microphone per speaker.

Will the labels be better if I transcribe afterwards instead of live?

Usually, yes. Post-Process+ has the whole recording available before it decides anything, so it can use the later parts of a conversation to settle a decision about the earlier parts. Realtime has to commit as the words arrive. Both cost the same transcription time in Kalima, so if nobody needs to read along during the conversation, there is no reason to pick Realtime.

Does fixing the speakers send my audio anywhere?

No. The Check voices review runs on your own device with a 38 MB voice model that downloads once and stays there, and no audio leaves the device for that step. Renaming a speaker is a text change in your own transcript.

Try Kalima on your next call.

Live transcription and translation, no bot in the meeting. Free to start.

Start for free

Keep reading.