Stream text input
WSS wss://api.kittenml.com/v1/tts/realtime
Use input streaming when text is produced over time, such as an LLM response.
Keep one WebSocket open, send text fragments as they arrive, and play each
audio delta immediately. This route is a KittenML extension; the
OpenAI-compatible POST /v1/audio/speech endpoint receives complete input
text and streams only its output.
The API key needs the tts:generate permission. Send it in the WebSocket
upgrade header:
Authorization: Bearer sk_kitten_live_...
Connect
wss://api.kittenml.com/v1/tts/realtime?model=kitten-tts-2-latest&voice=eleanor_somber_female_32&speed=1&response_format=pcm
| Query parameter | Default | Accepted values |
|---|---|---|
model | kitten-tts-mini-0.8 | kitten-tts-2-latest or a Nano, Micro, or Mini 0.8 model ID |
voice | Bella (eleanor_somber_female_32 for Kitten TTS 2) | A voice ID from the selected model's catalog (for example eleanor_somber_female_32), or an owner-scoped saved voice_... ID for Kitten TTS 2 |
speed | 1 | 0.25 through 4.0 |
response_format | pcm | pcm |
mode, temperature, top_p, top_k, min_p, max_new_tokens | From mode (expressive) | Kitten TTS 2 only. The decoding controls, with the same ranges; they apply to every phrase of the session |
For Kitten TTS 2, send model=kitten-tts-2-latest; its earlier model ID still
works, as described in Models.
An invalid decoding value, such as top_p=0.2, closes the socket with an
error event whose param names the field, then code 1008:
wss://api.kittenml.com/v1/tts/realtime?model=kitten-tts-2-latest&voice=maeve_cozy_female_22&mode=expressive&top_k=40&response_format=pcm
Emotion and sound tags work here too.
Input streaming returns headerless 24 kHz, mono, signed little-endian PCM16. Read server events continuously while sending text; this prevents client-side buffers from delaying long streams.
Kitten TTS 2 can return speech.audio.delta before input_text.done. Keep
the reader and writer active concurrently, as the examples do. A private
saved custom voice also works here: set voice to its voice_... ID in the
query. The saved voice may be created through OpenAI's consent-bound lifecycle
or through KittenML's consent-free POST /v1/audio/voices extension:
curl --fail-with-body -sS \
https://api.kittenml.com/v1/audio/voices \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'name=My realtime voice' \
-F 'audio_sample=@reference.wav;type=audio/wav'
Use the returned ID in the connection URL:
wss://api.kittenml.com/v1/tts/realtime?model=kitten-tts-2-latest&voice=voice_0123456789abcdef0123456789abcdef&speed=1&response_format=pcm
Inline reference_audio is not a WebSocket query payload. Save it first to
obtain a voice ID, or use inline cloning with POST /v1/audio/speech.
Receiving audio before input completion proves bidirectional overlap, but does not guarantee uninterrupted playback: phrase boundaries, load, and replica placement can still produce gaps. Applications should monitor arrival cadence and buffer PCM adaptively rather than equating the WebSocket transport with a latency or continuity guarantee. See Stream output for measured gap rates and a starting buffer size.
Client events
Append text as it becomes available:
{"type":"input_text.append","text":"The first part of the response. "}
The server buffers incomplete phrases and applies backpressure. Send a commit to synthesize all currently buffered text while keeping the session open:
{"type":"input_text.commit"}
End the input after the final fragment:
{"type":"input_text.done"}
One session accepts at most 262,144 cumulative input characters. Each append
may contain at most 8,192 characters and each WebSocket frame may contain at
most 64 KiB. Send 50–500 characters at a time, preferably complete phrases,
and do not send more text after input_text.done.
Server events
The connection begins with session.created. Each accepted append produces
input_text.accepted; a commit produces input_text.committed.
Audio arrives in sequence-numbered Base64 deltas:
{
"type": "speech.audio.delta",
"event_id": "event_...",
"audio": "AACAPw...",
"sequence": 0
}
Decode and concatenate the audio values in sequence order. The terminal
event confirms the complete duration and character count:
{
"type": "speech.audio.done",
"event_id": "event_...",
"usage": {
"input_characters": 1280,
"audio_seconds": 92.375
}
}
If an error event arrives or the socket closes before speech.audio.done,
treat the generation as incomplete.
Node.js example
Install ws, then run this server-side example. Native browser WebSockets
cannot set an Authorization header; do not expose a long-lived API key in
browser code.
import {createWriteStream} from 'node:fs';
import WebSocket from 'ws';
const ws = new WebSocket(
'wss://api.kittenml.com/v1/tts/realtime' +
'?model=kitten-tts-2-latest&voice=eleanor_somber_female_32' +
'&response_format=pcm',
{headers: {Authorization: `Bearer ${process.env.KITTENML_API_KEY}`}},
);
// Audio deltas are contiguous 24 kHz, mono, signed 16-bit PCM.
const output = createWriteStream('speech.pcm');
let requestId = 'unknown';
let terminalReceived = false;
ws.on('open', () => {
// Append text as it becomes available, then mark the input complete once.
ws.send(JSON.stringify({type: 'input_text.append', text: 'Hello from '}));
ws.send(JSON.stringify({type: 'input_text.append', text: 'a live text stream.'}));
ws.send(JSON.stringify({type: 'input_text.done'}));
});
ws.on('message', (raw) => {
const event = JSON.parse(raw.toString());
if (event.type === 'session.created') {
requestId = event.request_id || event.session?.id || requestId;
} else if (event.type === 'speech.audio.delta') {
// Decode and append each delta in arrival order.
output.write(Buffer.from(event.audio, 'base64'));
} else if (event.type === 'speech.audio.done') {
// Only this terminal event confirms that the output file is complete.
terminalReceived = true;
output.end();
console.log(`request_id: ${event.request_id || requestId}`);
ws.close();
} else if (event.type === 'error') {
terminalReceived = true;
output.destroy();
const error = event.error || {};
const eventRequestId = event.request_id || error.request_id || requestId;
// Log both identifiers so the failed session can be correlated with support.
console.error(
`${error.message || 'TTS stream failed'} ` +
`(code=${error.code || 'unknown'}, request_id=${eventRequestId})`,
);
ws.close();
}
});
ws.on('close', (code) => {
// A clean socket close is still a failure without speech.audio.done.
if (!terminalReceived) {
output.destroy();
console.error(
`TTS connection closed before speech.audio.done ` +
`(code=connection_closed, close_code=${code}, request_id=${requestId})`,
);
process.exitCode = 1;
}
});
Play the raw output with a player configured for signed 16-bit little-endian PCM, 24 kHz, mono. Runnable Python and JavaScript clients are available in the realtime TTS examples.
Limits and recovery
Input-streaming WebSockets share the same two active TTS slots per organization as HTTP generations. A slot lasts until completion or disconnect. Reconnect and resubmit only the text you still need after a failure; a new connection is a new generation. A session closes after five minutes without a client message or after two hours total, whichever happens first.