Skip to main content

Stream speech output

POST https://api.kittenml.com/v1/audio/speech

Output streaming uses the same OpenAI-compatible speech endpoint. The request still contains the complete input text, but the client consumes audio before the complete response has downloaded.

Choose one wire format:

  • stream_format: "audio" streams raw bytes in the selected audio format and works with the official OpenAI SDK streaming helpers.
  • stream_format: "sse" sends explicit KittenML speech.audio.delta events containing Base64 audio.

For quality over speed, send "stream": false instead to get one complete file; see Quality over speed.

Kitten TTS 2 emits progressive audio in production. With response_format: "pcm", each full SSE delta is 4,800 bytes: 100 ms of 24 kHz mono PCM16. Gateways and HTTP clients may split or combine raw binary writes, so binary transport chunks do not preserve those application boundaries. Choose and adapt the playback buffer from observed network conditions; the API does not currently guarantee underrun-free realtime playback with a fixed 100 ms or 200 ms buffer.

Short gaps are more likely when many streams are in progress at once and on text with emotion markup. Start with a playback buffer of at least 200 ms, or about 350 ms for marked-up text, and adapt from there. Encoded formats such as MP3 remain byte streams, so their network chunk sizes do not correspond to a fixed amount of playable audio.

To send the input text incrementally as well, use the separate text-input WebSocket.

Binary and SSE output streaming share the HTTP endpoint's 40,000-character, 256 KiB JSON-body, and 30-minute request limits. For predictable latency, keep individual requests around 400–5,000 characters.

Stream binary audio with Python​

import os
from openai import OpenAI

client = OpenAI(
api_key=os.environ["KITTENML_API_KEY"],
base_url="https://api.kittenml.com/v1",
)

# The streaming context writes network chunks without buffering the full MP3.
with client.audio.speech.with_streaming_response.create(
model="kitten-tts-2-latest",
input="The first sentence can play while the next sentence is generated.",
voice="eleanor_somber_female_32",
response_format="mp3",
speed=1.0,
) as audio:
# Keep x-request-id for support and telemetry correlation.
request_id = audio.headers.get("x-request-id")
audio.stream_to_file("kitten-output.mp3")

print(f"request_id: {request_id}")

stream_to_file writes bytes as they arrive instead of keeping the complete audio response in application memory.

Stream binary audio with JavaScript​

import {createWriteStream} from "node:fs";
import {Readable} from "node:stream";
import {pipeline} from "node:stream/promises";
import OpenAI from "openai";

const client = new OpenAI({
apiKey: process.env.KITTENML_API_KEY,
baseURL: "https://api.kittenml.com/v1",
});

// withResponse() keeps x-request-id while exposing the streaming response body.
const {data: audio, request_id: requestId} =
await client.audio.speech.create({
model: "kitten-tts-2-latest",
input: "The first sentence can play while the next sentence is generated.",
voice: "eleanor_somber_female_32",
response_format: "mp3",
}).withResponse();

// Pipe arriving bytes to disk instead of buffering the complete MP3.
await pipeline(
Readable.fromWeb(audio.body),
createWriteStream("kitten-output.mp3"),
);
console.log(`request_id: ${requestId}`);

Stream SSE events​

Set stream_format to sse when your application needs event boundaries:

# -N disables cURL buffering so each SSE event is printed as it arrives, and
# -L follows the continuation redirect used when capacity is still warming up.
curl -L --fail-with-body -sS -N \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "kitten-tts-2-latest",
"input": "The first sentence can play while the next sentence is generated.",
"voice": "eleanor_somber_female_32",
"response_format": "mp3",
"stream_format": "sse"
}'

Each delta contains one Base64-encoded piece of the selected audio format:

data: {"type":"speech.audio.delta","audio":"SUQzBAAAAA..."}

data: {"type":"speech.audio.delta","audio":"qqqqqqqqqq..."}

data: {"type":"speech.audio.done","usage":{"input_tokens":0,"output_tokens":0,"total_tokens":0,"audio_seconds":8.06,"input_characters":65}}

Decode each audio value and concatenate the bytes in event order. The combined bytes are one playable file; the server does not create a separate file for each delta. speech.audio.done confirms that the stream completed. With Kitten TTS 2, its usage reports the input characters processed and the seconds of audio generated; the token fields are always 0.

Recover from interruption​

If the connection closes before the HTTP response completes or before an SSE speech.audio.done event, treat the audio as incomplete. Discard or explicitly mark the partial output, then retry the generation as a new request. Do not blindly append a retry to the partial file because encoded audio containers may contain headers and trailers.

Runnable Python and JavaScript clients for both binary and SSE modes are in the output-streaming TTS examples.