Speaker diarization
Diarization answers "who said this" alongside "what was said". Transcript text is grouped into turns, each attributed to a speaker label that stays stable for the length of the session.
It is off by default on every surface. A session that does not ask for it behaves exactly as it does today, and a client that ignores the extra event sees no change.
Realtime transcription
Enable diarization in a session.update before or during the session, or in
the body that mints an ephemeral client secret — a browser client never sends
session.update itself, so that is its only opportunity:
{
"type": "session.update",
"session": {
"diarization": { "enabled": true, "max_speakers": 6 }
}
}
| Field | Meaning |
|---|---|
enabled | Turns diarization on for this session. Defaults to false. |
max_speakers | Upper bound on distinct speakers, 1 to 20. Optional. A value outside that range is clamped rather than rejected, and session.updated reports what was applied. |
Transcript events carry the same fields and arrive at the same speed as a session without diarization.
One thing does change, by design: a detected speaker change finalises the current turn, so turns are cut at speaker boundaries rather than only where the transcriber would otherwise have ended them. The words are the same; where one turn ends and the next begins is not. If you need transcription identical to a session without diarization, leave it off for that session.
The speakers event
Speaker attribution arrives as its own server event once a turn is finalized:
{
"type": "conversation.item.input_audio_transcription.speakers",
"item_id": "item_ab12cd34_3",
"speaker": "speaker_1",
"text": "[Calm]and that's what the numbers showed[/Calm]",
"attribution": "commit",
"segments": [
{
"speaker": "speaker_1",
"segment_start_ms": 46000,
"segment_end_ms": 118600
}
]
}
| Field | Meaning |
|---|---|
speaker | Stable label for this session, speaker_0, speaker_1, and so on. null when the speaker could not be determined. |
text | The transcript text belonging to this speaker, with enriched markup preserved. |
attribution | How the speaker was determined. See below. |
segments | Time span of this turn, on the same clock as the segment timing fields in completed transcripts. |
item_id | Correlates with the transcript events for the same turn. |
attribution reports how confident the assignment is:
| Value | Meaning |
|---|---|
commit | A speaker change was detected and the transcript was finalized at that point. |
endpoint | The transcriber finalized the turn on its own, and the text was attributed to whoever held the floor. |
unattributed | No speaker could be determined. speaker is null and the text is still returned. |
Why speakers arrive separately
A speaker label is only trustworthy once enough of a voice has been heard. Holding transcript text back until then would delay every word.
So words stream at their usual latency and attribution follows, typically within half a second to two seconds of the speaker changing. Build a live caption on the transcript events and fill in speaker labels as the speakers events arrive.
The speakers event fires when a turn is finalized, not continuously. A speaker who talks without pausing is attributed when that stretch closes, not while it is still in progress.
Enriched markup is preserved
Diarization does not alter transcript text. Emotion, pause, and emphasis markup survive attribution, so a turn keeps its enrichment:
[speaker_0] [Slow]Good morning and welcome to the third quarter
conference call.[Pause_long] All participants will be in
listen-only mode.[/Slow]
[speaker_1] [Calm]Thank you for joining us today.[Pause_medium] As usual,
I'm going to mention a few highlights.[/Calm]
This holds on both surfaces: the realtime speakers event and upload's
turns[].text both carry the marked-up text.
Tag names are not case-consistent in model output: both [Pause_long] and
[pause_long] occur. Match tag names case-insensitively.
Transcription never depends on diarization
If diarization fails, degrades, or is unavailable, transcription continues
unaffected. Text is still delivered in full; turns that could not be attributed
arrive with speaker set to null and attribution set to unattributed.
session.updated reports the current state:
{
"type": "session.updated",
"session": {
"diarization": { "enabled": true, "status": "active", "max_speakers": 6 }
}
}
The diarization key appears only for sessions that asked for it. If
max_speakers had to be clamped, the block also carries
max_speakers_requested with the original value.
On a deployment where realtime diarization is switched off, a session that asks for it is told so rather than left guessing:
{ "type": "session.updated",
"session": { "diarization": { "enabled": false, "status": "unavailable" } } }
| Status | Meaning |
|---|---|
active | Diarization is running for this session. |
unavailable | Diarization could not start. Transcription is unaffected. |
degraded | Diarization stopped mid-session. Earlier turns keep their labels; later ones are unattributed. |
File upload
Add diarization=true to a transcription request:
curl https://api.kittenml.com/v1/audio/transcriptions \
-H "Authorization: Bearer sk_kitten_live_..." \
-F file=@meeting.mp3 \
-F model=kittenasr-enhanced-preview \
-F response_format=verbose_json \
-F diarization=true \
-F max_speakers=6
| Field | Meaning |
|---|---|
diarization | true to enable. Defaults to false. |
num_speakers | Exact speaker count, when known. Cannot be combined with min_speakers or max_speakers. |
min_speakers | Lower bound on speaker count. |
max_speakers | Upper bound on speaker count. |
Speaker counts accept 1 to 20. response_format must be json or
verbose_json.
The response carries a diarization block alongside the transcript:
{
"text": "Good morning and welcome... Thank you for joining us today...",
"diarization": {
"status": "completed",
"speakers": ["speaker_0", "speaker_1"],
"segments": [
{ "id": 0, "start": 1.13, "end": 44.80, "speaker": "speaker_0" },
{ "id": 1, "start": 46.00, "end": 118.60, "speaker": "speaker_1" }
],
"turns": [
{ "speaker": "speaker_0", "start": 1.13, "end": 44.80,
"text": "[Slow]Good morning and welcome...[/Slow]" },
{ "speaker": "speaker_1", "start": 46.00, "end": 118.60,
"text": "[Calm]Thank you for joining us today...[/Calm]" }
]
}
}
| Field | Meaning |
|---|---|
segments | Every speaker span the diarizer produced, with id. Times only. |
turns | The same spans with consecutive same-speaker segments merged, each carrying the transcript text spoken in it. Present for both json and verbose_json. |
Use turns unless you need the raw segmentation. A transcript segment that
straddles a speaker change is assigned to whichever turn it overlaps more, so it
is neither duplicated nor dropped — attribution is whole-segment, not
word-level.
Text that falls in no speaker turn is still returned, in a trailing entry whose
speaker, start and end are all null. Transcript text is never dropped to
make the speaker split tidy.
When diarization degrades
Neither case fails the request; the transcript always returns.
| What happened | What you get |
|---|---|
| Diarization failed | "status": "failed", empty speakers and segments, no turns, and a diarization_failed warning. |
| Diarization worked but no text could be matched to it | "status": "completed" with speakers and segments populated, turns absent, and a diarization_text_unavailable warning. |
Treat turns as optional: check for the key rather than assuming it.
Which transports support it
| Transport | Diarization |
|---|---|
/v1/realtime WebSocket (and /v1/stt/realtime) | Supported |
File upload /v1/audio/transcriptions, json / verbose_json | Supported |
File upload with stream=true (SSE) | Supported — the block arrives on the final transcript.text.done event |
WebRTC /v1/realtime/calls | Not supported; returns diarization_not_supported_on_transport |
Limits
| Limit | Value |
|---|---|
| Speakers tracked, realtime | 1 to 20. Beyond max_speakers, additional voices are merged into existing labels. session.updated reports the ceiling applied. |
| Speakers, file upload | 1 to 20 |
| Attribution granularity | Whole turns. Word-level speaker labels are not available. |
Accuracy falls as the number of speakers rises, and it falls on narrowband or noisy audio whatever the count. Two or three speakers on clean audio are reliable; a crowded, similar-sounding room is not. Treat speaker labels on many-speaker audio as a strong hint rather than a fact.
Overlapping speech is not separated. When two people talk at once the turn is
attributed to one of them, or returned as unattributed.
Errors
Diarization follows the standard error format.
| Code | Surface | Meaning |
|---|---|---|
diarization_not_enabled | File upload (501) | Diarization is not available. On the realtime path this is reported as status: "unavailable" rather than an error. |
invalid_diarization_options | File upload (400) | Speaker-count options out of range, mutually exclusive, or min_speakers greater than max_speakers. |
invalid_response_format | File upload (400) | diarization=true combined with text, srt or vtt. |
diarization_not_supported_on_transport | WebRTC (400) | Diarization is not available over WebRTC. Use the realtime WebSocket, or upload the recording. |
Speaker-count options sent without diarization=true are rejected rather than
ignored.
Warnings
Warnings never fail a request; they appear in a warnings array beside the
transcript.
| Code | Meaning |
|---|---|
diarization_failed | Diarization could not run. The transcript is unaffected. |
diarization_text_unavailable | Speaker spans were produced but no transcript text could be matched to them, so turns is absent. |