Skip to main content

Speaker diarization

Diarization answers "who said this" alongside "what was said". Transcript text is grouped into turns, each attributed to a speaker label that stays stable for the length of the session.

It is off by default on every surface. A session that does not ask for it behaves exactly as it does today, and a client that ignores the extra event sees no change.

Realtime transcription​

Enable diarization in a session.update before or during the session, or in the body that mints an ephemeral client secret — a browser client never sends session.update itself, so that is its only opportunity:

{
"type": "session.update",
"session": {
"diarization": { "enabled": true, "max_speakers": 6 }
}
}
FieldMeaning
enabledTurns diarization on for this session. Defaults to false.
max_speakersUpper bound on distinct speakers, 1 to 20. Optional. A value outside that range is clamped rather than rejected, and session.updated reports what was applied.

Transcript events carry the same fields and arrive at the same speed as a session without diarization.

One thing does change, by design: a detected speaker change finalises the current turn, so turns are cut at speaker boundaries rather than only where the transcriber would otherwise have ended them. The words are the same; where one turn ends and the next begins is not. If you need transcription identical to a session without diarization, leave it off for that session.

The speakers event​

Speaker attribution arrives as its own server event once a turn is finalized:

{
"type": "conversation.item.input_audio_transcription.speakers",
"item_id": "item_ab12cd34_3",
"speaker": "speaker_1",
"text": "[Calm]and that's what the numbers showed[/Calm]",
"attribution": "commit",
"segments": [
{
"speaker": "speaker_1",
"segment_start_ms": 46000,
"segment_end_ms": 118600
}
]
}
FieldMeaning
speakerStable label for this session, speaker_0, speaker_1, and so on. null when the speaker could not be determined.
textThe transcript text belonging to this speaker, with enriched markup preserved.
attributionHow the speaker was determined. See below.
segmentsTime span of this turn, on the same clock as the segment timing fields in completed transcripts.
item_idCorrelates with the transcript events for the same turn.

attribution reports how confident the assignment is:

ValueMeaning
commitA speaker change was detected and the transcript was finalized at that point.
endpointThe transcriber finalized the turn on its own, and the text was attributed to whoever held the floor.
unattributedNo speaker could be determined. speaker is null and the text is still returned.

Why speakers arrive separately​

A speaker label is only trustworthy once enough of a voice has been heard. Holding transcript text back until then would delay every word.

So words stream at their usual latency and attribution follows, typically within half a second to two seconds of the speaker changing. Build a live caption on the transcript events and fill in speaker labels as the speakers events arrive.

The speakers event fires when a turn is finalized, not continuously. A speaker who talks without pausing is attributed when that stretch closes, not while it is still in progress.

Enriched markup is preserved​

Diarization does not alter transcript text. Emotion, pause, and emphasis markup survive attribution, so a turn keeps its enrichment:

[speaker_0] [Slow]Good morning and welcome to the third quarter
conference call.[Pause_long] All participants will be in
listen-only mode.[/Slow]
[speaker_1] [Calm]Thank you for joining us today.[Pause_medium] As usual,
I'm going to mention a few highlights.[/Calm]

This holds on both surfaces: the realtime speakers event and upload's turns[].text both carry the marked-up text.

Tag names are not case-consistent in model output: both [Pause_long] and [pause_long] occur. Match tag names case-insensitively.

Transcription never depends on diarization​

If diarization fails, degrades, or is unavailable, transcription continues unaffected. Text is still delivered in full; turns that could not be attributed arrive with speaker set to null and attribution set to unattributed.

session.updated reports the current state:

{
"type": "session.updated",
"session": {
"diarization": { "enabled": true, "status": "active", "max_speakers": 6 }
}
}

The diarization key appears only for sessions that asked for it. If max_speakers had to be clamped, the block also carries max_speakers_requested with the original value.

On a deployment where realtime diarization is switched off, a session that asks for it is told so rather than left guessing:

{ "type": "session.updated",
"session": { "diarization": { "enabled": false, "status": "unavailable" } } }
StatusMeaning
activeDiarization is running for this session.
unavailableDiarization could not start. Transcription is unaffected.
degradedDiarization stopped mid-session. Earlier turns keep their labels; later ones are unattributed.

File upload​

Add diarization=true to a transcription request:

curl https://api.kittenml.com/v1/audio/transcriptions \
-H "Authorization: Bearer sk_kitten_live_..." \
-F file=@meeting.mp3 \
-F model=kittenasr-enhanced-preview \
-F response_format=verbose_json \
-F diarization=true \
-F max_speakers=6
FieldMeaning
diarizationtrue to enable. Defaults to false.
num_speakersExact speaker count, when known. Cannot be combined with min_speakers or max_speakers.
min_speakersLower bound on speaker count.
max_speakersUpper bound on speaker count.

Speaker counts accept 1 to 20. response_format must be json or verbose_json.

The response carries a diarization block alongside the transcript:

{
"text": "Good morning and welcome... Thank you for joining us today...",
"diarization": {
"status": "completed",
"speakers": ["speaker_0", "speaker_1"],
"segments": [
{ "id": 0, "start": 1.13, "end": 44.80, "speaker": "speaker_0" },
{ "id": 1, "start": 46.00, "end": 118.60, "speaker": "speaker_1" }
],
"turns": [
{ "speaker": "speaker_0", "start": 1.13, "end": 44.80,
"text": "[Slow]Good morning and welcome...[/Slow]" },
{ "speaker": "speaker_1", "start": 46.00, "end": 118.60,
"text": "[Calm]Thank you for joining us today...[/Calm]" }
]
}
}
FieldMeaning
segmentsEvery speaker span the diarizer produced, with id. Times only.
turnsThe same spans with consecutive same-speaker segments merged, each carrying the transcript text spoken in it. Present for both json and verbose_json.

Use turns unless you need the raw segmentation. A transcript segment that straddles a speaker change is assigned to whichever turn it overlaps more, so it is neither duplicated nor dropped — attribution is whole-segment, not word-level.

Text that falls in no speaker turn is still returned, in a trailing entry whose speaker, start and end are all null. Transcript text is never dropped to make the speaker split tidy.

When diarization degrades​

Neither case fails the request; the transcript always returns.

What happenedWhat you get
Diarization failed"status": "failed", empty speakers and segments, no turns, and a diarization_failed warning.
Diarization worked but no text could be matched to it"status": "completed" with speakers and segments populated, turns absent, and a diarization_text_unavailable warning.

Treat turns as optional: check for the key rather than assuming it.

Which transports support it​

TransportDiarization
/v1/realtime WebSocket (and /v1/stt/realtime)Supported
File upload /v1/audio/transcriptions, json / verbose_jsonSupported
File upload with stream=true (SSE)Supported — the block arrives on the final transcript.text.done event
WebRTC /v1/realtime/callsNot supported; returns diarization_not_supported_on_transport

Limits​

LimitValue
Speakers tracked, realtime1 to 20. Beyond max_speakers, additional voices are merged into existing labels. session.updated reports the ceiling applied.
Speakers, file upload1 to 20
Attribution granularityWhole turns. Word-level speaker labels are not available.

Accuracy falls as the number of speakers rises, and it falls on narrowband or noisy audio whatever the count. Two or three speakers on clean audio are reliable; a crowded, similar-sounding room is not. Treat speaker labels on many-speaker audio as a strong hint rather than a fact.

Overlapping speech is not separated. When two people talk at once the turn is attributed to one of them, or returned as unattributed.

Errors​

Diarization follows the standard error format.

CodeSurfaceMeaning
diarization_not_enabledFile upload (501)Diarization is not available. On the realtime path this is reported as status: "unavailable" rather than an error.
invalid_diarization_optionsFile upload (400)Speaker-count options out of range, mutually exclusive, or min_speakers greater than max_speakers.
invalid_response_formatFile upload (400)diarization=true combined with text, srt or vtt.
diarization_not_supported_on_transportWebRTC (400)Diarization is not available over WebRTC. Use the realtime WebSocket, or upload the recording.

Speaker-count options sent without diarization=true are rejected rather than ignored.

Warnings​

Warnings never fail a request; they appear in a warnings array beside the transcript.

CodeMeaning
diarization_failedDiarization could not run. The transcript is unaffected.
diarization_text_unavailableSpeaker spans were produced but no transcript text could be matched to them, so turns is absent.