Skip to main content

Stream text input

WSS wss://api.kittenml.com/v1/tts/realtime

Use input streaming when text is produced over time, such as an LLM response. Keep one WebSocket open, send text fragments as they arrive, and play each audio delta immediately. This route is a KittenML extension; the OpenAI-compatible POST /v1/audio/speech endpoint receives complete input text and streams only its output.

The API key needs the tts:generate permission. Send it in the WebSocket upgrade header:

Authorization: Bearer sk_kitten_live_...

Connect​

wss://api.kittenml.com/v1/tts/realtime?model=kitten-tts-2-latest&voice=eleanor_somber_female_32&speed=1&response_format=pcm
Query parameterDefaultAccepted values
modelkitten-tts-mini-0.8kitten-tts-2-latest or a Nano, Micro, or Mini 0.8 model ID
voiceBella (eleanor_somber_female_32 for Kitten TTS 2)A voice ID from the selected model's catalog (for example eleanor_somber_female_32), or an owner-scoped saved voice_... ID for Kitten TTS 2
speed10.25 through 4.0
response_formatpcmpcm
mode, temperature, top_p, top_k, min_p, max_new_tokensFrom mode (expressive)Kitten TTS 2 only. The decoding controls, with the same ranges; they apply to every phrase of the session

For Kitten TTS 2, send model=kitten-tts-2-latest; its earlier model ID still works, as described in Models.

An invalid decoding value, such as top_p=0.2, closes the socket with an error event whose param names the field, then code 1008:

wss://api.kittenml.com/v1/tts/realtime?model=kitten-tts-2-latest&voice=maeve_cozy_female_22&mode=expressive&top_k=40&response_format=pcm

Emotion and sound tags work here too.

Input streaming returns headerless 24 kHz, mono, signed little-endian PCM16. Read server events continuously while sending text; this prevents client-side buffers from delaying long streams.

Kitten TTS 2 can return speech.audio.delta before input_text.done. Keep the reader and writer active concurrently, as the examples do. A private saved custom voice also works here: set voice to its voice_... ID in the query. The saved voice may be created through OpenAI's consent-bound lifecycle or through KittenML's consent-free POST /v1/audio/voices extension:

curl --fail-with-body -sS \
https://api.kittenml.com/v1/audio/voices \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'name=My realtime voice' \
-F 'audio_sample=@reference.wav;type=audio/wav'

Use the returned ID in the connection URL:

wss://api.kittenml.com/v1/tts/realtime?model=kitten-tts-2-latest&voice=voice_0123456789abcdef0123456789abcdef&speed=1&response_format=pcm

Inline reference_audio is not a WebSocket query payload. Save it first to obtain a voice ID, or use inline cloning with POST /v1/audio/speech.

Receiving audio before input completion proves bidirectional overlap, but does not guarantee uninterrupted playback: phrase boundaries, load, and replica placement can still produce gaps. Applications should monitor arrival cadence and buffer PCM adaptively rather than equating the WebSocket transport with a latency or continuity guarantee. See Stream output for measured gap rates and a starting buffer size.

Client events​

Append text as it becomes available:

{"type":"input_text.append","text":"The first part of the response. "}

The server buffers incomplete phrases and applies backpressure. Send a commit to synthesize all currently buffered text while keeping the session open:

{"type":"input_text.commit"}

End the input after the final fragment:

{"type":"input_text.done"}

One session accepts at most 262,144 cumulative input characters. Each append may contain at most 8,192 characters and each WebSocket frame may contain at most 64 KiB. Send 50–500 characters at a time, preferably complete phrases, and do not send more text after input_text.done.

Server events​

The connection begins with session.created. Each accepted append produces input_text.accepted; a commit produces input_text.committed.

Audio arrives in sequence-numbered Base64 deltas:

{
"type": "speech.audio.delta",
"event_id": "event_...",
"audio": "AACAPw...",
"sequence": 0
}

Decode and concatenate the audio values in sequence order. The terminal event confirms the complete duration and character count:

{
"type": "speech.audio.done",
"event_id": "event_...",
"usage": {
"input_characters": 1280,
"audio_seconds": 92.375
}
}

If an error event arrives or the socket closes before speech.audio.done, treat the generation as incomplete.

Node.js example​

Install ws, then run this server-side example. Native browser WebSockets cannot set an Authorization header; do not expose a long-lived API key in browser code.

import {createWriteStream} from 'node:fs';
import WebSocket from 'ws';

const ws = new WebSocket(
'wss://api.kittenml.com/v1/tts/realtime' +
'?model=kitten-tts-2-latest&voice=eleanor_somber_female_32' +
'&response_format=pcm',
{headers: {Authorization: `Bearer ${process.env.KITTENML_API_KEY}`}},
);
// Audio deltas are contiguous 24 kHz, mono, signed 16-bit PCM.
const output = createWriteStream('speech.pcm');
let requestId = 'unknown';
let terminalReceived = false;

ws.on('open', () => {
// Append text as it becomes available, then mark the input complete once.
ws.send(JSON.stringify({type: 'input_text.append', text: 'Hello from '}));
ws.send(JSON.stringify({type: 'input_text.append', text: 'a live text stream.'}));
ws.send(JSON.stringify({type: 'input_text.done'}));
});

ws.on('message', (raw) => {
const event = JSON.parse(raw.toString());
if (event.type === 'session.created') {
requestId = event.request_id || event.session?.id || requestId;
} else if (event.type === 'speech.audio.delta') {
// Decode and append each delta in arrival order.
output.write(Buffer.from(event.audio, 'base64'));
} else if (event.type === 'speech.audio.done') {
// Only this terminal event confirms that the output file is complete.
terminalReceived = true;
output.end();
console.log(`request_id: ${event.request_id || requestId}`);
ws.close();
} else if (event.type === 'error') {
terminalReceived = true;
output.destroy();
const error = event.error || {};
const eventRequestId = event.request_id || error.request_id || requestId;
// Log both identifiers so the failed session can be correlated with support.
console.error(
`${error.message || 'TTS stream failed'} ` +
`(code=${error.code || 'unknown'}, request_id=${eventRequestId})`,
);
ws.close();
}
});

ws.on('close', (code) => {
// A clean socket close is still a failure without speech.audio.done.
if (!terminalReceived) {
output.destroy();
console.error(
`TTS connection closed before speech.audio.done ` +
`(code=connection_closed, close_code=${code}, request_id=${requestId})`,
);
process.exitCode = 1;
}
});

Play the raw output with a player configured for signed 16-bit little-endian PCM, 24 kHz, mono. Runnable Python and JavaScript clients are available in the realtime TTS examples.

Limits and recovery​

Input-streaming WebSockets share the same two active TTS slots per organization as HTTP generations. A slot lasts until completion or disconnect. Reconnect and resubmit only the text you still need after a failure; a new connection is a new generation. A session closes after five minutes without a client message or after two hours total, whichever happens first.