Stream speech output
POST https://api.kittenml.com/v1/audio/speech
Output streaming uses the same OpenAI-compatible speech endpoint. The request still contains the complete input text, but the client consumes audio before the complete response has downloaded.
Choose one wire format:
stream_format: "audio"streams raw bytes in the selected audio format and works with the official OpenAI SDK streaming helpers.stream_format: "sse"sends explicit KittenMLspeech.audio.deltaevents containing Base64 audio.
For quality over speed, send "stream": false instead to get one complete
file; see Quality over speed.
Kitten TTS 2 emits progressive audio in production. With
response_format: "pcm", each full SSE delta is 4,800 bytes: 100 ms of 24 kHz
mono PCM16. Gateways and HTTP clients may split or combine raw binary writes,
so binary transport chunks do not preserve those application boundaries.
Choose and adapt the playback buffer from observed network conditions; the API
does not currently guarantee underrun-free realtime playback with a fixed
100 ms or 200 ms buffer.
Short gaps are more likely when many streams are in progress at once and on text with emotion markup. Start with a playback buffer of at least 200 ms, or about 350 ms for marked-up text, and adapt from there. Encoded formats such as MP3 remain byte streams, so their network chunk sizes do not correspond to a fixed amount of playable audio.
To send the input text incrementally as well, use the separate text-input WebSocket.
Binary and SSE output streaming share the HTTP endpoint's 40,000-character, 256 KiB JSON-body, and 30-minute request limits. For predictable latency, keep individual requests around 400–5,000 characters.
Stream binary audio with Python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["KITTENML_API_KEY"],
base_url="https://api.kittenml.com/v1",
)
# The streaming context writes network chunks without buffering the full MP3.
with client.audio.speech.with_streaming_response.create(
model="kitten-tts-2-latest",
input="The first sentence can play while the next sentence is generated.",
voice="eleanor_somber_female_32",
response_format="mp3",
speed=1.0,
) as audio:
# Keep x-request-id for support and telemetry correlation.
request_id = audio.headers.get("x-request-id")
audio.stream_to_file("kitten-output.mp3")
print(f"request_id: {request_id}")
stream_to_file writes bytes as they arrive instead of keeping the complete
audio response in application memory.
Stream binary audio with JavaScript
import {createWriteStream} from "node:fs";
import {Readable} from "node:stream";
import {pipeline} from "node:stream/promises";
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.KITTENML_API_KEY,
baseURL: "https://api.kittenml.com/v1",
});
// withResponse() keeps x-request-id while exposing the streaming response body.
const {data: audio, request_id: requestId} =
await client.audio.speech.create({
model: "kitten-tts-2-latest",
input: "The first sentence can play while the next sentence is generated.",
voice: "eleanor_somber_female_32",
response_format: "mp3",
}).withResponse();
// Pipe arriving bytes to disk instead of buffering the complete MP3.
await pipeline(
Readable.fromWeb(audio.body),
createWriteStream("kitten-output.mp3"),
);
console.log(`request_id: ${requestId}`);
Stream SSE events
Set stream_format to sse when your application needs event boundaries:
# -N disables cURL buffering so each SSE event is printed as it arrives, and
# -L follows the continuation redirect used when capacity is still warming up.
curl -L --fail-with-body -sS -N \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "kitten-tts-2-latest",
"input": "The first sentence can play while the next sentence is generated.",
"voice": "eleanor_somber_female_32",
"response_format": "mp3",
"stream_format": "sse"
}'
Each delta contains one Base64-encoded piece of the selected audio format:
data: {"type":"speech.audio.delta","audio":"SUQzBAAAAA..."}
data: {"type":"speech.audio.delta","audio":"qqqqqqqqqq..."}
data: {"type":"speech.audio.done","usage":{"input_tokens":0,"output_tokens":0,"total_tokens":0,"audio_seconds":8.06,"input_characters":65}}
Decode each audio value and concatenate the bytes in event order. The
combined bytes are one playable file; the server does not create a separate
file for each delta. speech.audio.done confirms that the stream completed.
With Kitten TTS 2, its usage reports the input characters processed and the
seconds of audio generated; the token fields are always 0.
Recover from interruption
If the connection closes before the HTTP response completes or before an SSE
speech.audio.done event, treat the audio as incomplete. Discard or explicitly
mark the partial output, then retry the generation as a new request. Do not
blindly append a retry to the partial file because encoded audio containers may
contain headers and trailers.
Runnable Python and JavaScript clients for both binary and SSE modes are in the output-streaming TTS examples.