Generate complete speech
POST https://api.kittenml.com/v1/audio/speech
Send the complete text in one request and receive one complete audio result. This is the standard OpenAI-compatible speech workflow. The examples below fully read the response before your application opens or processes the saved audio.
Kitten TTS 2 (kitten-tts-2-latest) and all three KittenTTS 0.8 models are
available. A valid API key with the tts:generate permission is required so
the generation can appear in the owning organization's usage history.
Need to play bytes as soon as they arrive? See Stream speech output. If the text itself arrives a piece at a time, see Stream text input.
Before you begin
Create an account-backed key in the
KittenML API key dashboard,
give it the tts:generate permission, and export it in your shell:
export KITTENML_API_KEY="sk_kitten_live_..."
Install the official OpenAI SDK for your language before running the SDK examples:
pip install openai
npm install openai
Request body
| Field | Type | Required | Default | Accepted values |
|---|---|---|---|---|
model | string | No | kitten-tts-mini-0.8 | kitten-tts-2-latest, kitten-tts-nano-0.8, kitten-tts-micro-0.8, or kitten-tts-mini-0.8 |
input | string | Yes | — | Non-empty text, at most 40,000 characters |
voice | string or object | No | Model-specific | A built-in voice, a saved {"id":"voice_..."} voice, or a Kitten TTS 2 inline reference object |
response_format | string | No | mp3 | mp3, wav, flac, aac, opus, or pcm |
speed | number | No | 1.0 | 0.25 through 4.0 |
stream | boolean | No | true | Kitten TTS 2 only: false returns one complete file with higher audio quality. See Quality over speed |
mode | string | No | expressive | Kitten TTS 2 only: stable or expressive. See Decoding controls |
temperature | number | No | From mode | Kitten TTS 2 only: 0.3 through 2.0 |
top_p | number | No | From mode | Kitten TTS 2 only: 0.5 through 1.0 |
top_k | integer | No | From mode | Kitten TTS 2 only: 0 through 200 (0 turns it off) |
min_p | number | No | From mode | Kitten TTS 2 only: 0 through 0.5 |
max_new_tokens | integer | No | None | Kitten TTS 2 only: 500 through 12000 |
For ordinary JSON requests, omit voice to use the selected model's built-in
default; do not send "voice": null.
| JSON request | Effective default |
|---|---|
model: "kitten-tts-2-latest", no voice | eleanor_somber_female_32 |
Any KittenTTS 0.8 model, no voice | Bella |
No model and no voice | kitten-tts-mini-0.8 with Bella |
A request without model is served by kitten-tts-mini-0.8, so every Kitten
TTS 2 request must name kitten-tts-2-latest.
Multipart speech requests are reserved for one-request cloning and require
reference_audio. Use ordinary JSON when you want a default built-in voice.
The JSON body may be at most 256 KiB and the request may run for at most 30
minutes. For predictable latency, keep individual generations around 400–5,000
characters. stream_format may be omitted for this workflow; its
OpenAI-compatible default is audio.
cURL
# Complete-response mode writes the finished MP3 after generation completes.
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "kitten-tts-2-latest",
"input": "Hello! This audio was generated by KittenTTS.",
"voice": "eleanor_somber_female_32",
"response_format": "mp3",
"speed": 1.0
}' \
--output kitten-output.mp3
The command returns after the response is complete and saves it as
kitten-output.mp3. The success body is audio, not JSON; do not pipe it to
jq or print it in a terminal.
-L is not optional. When a generation has to wait for hosting capacity, the
platform answers with a redirect to a continuation URL instead of the audio,
exactly as it does for upload transcription.
Without -L, cURL treats that redirect as the final response, exits 0, and
writes an empty file — a silent failure rather than an error. Every HTTP
client you use with this endpoint must follow redirects; the official OpenAI
SDKs already do.
Python with the OpenAI SDK
import os
from pathlib import Path
from openai import OpenAI
client = OpenAI(
api_key=os.environ["KITTENML_API_KEY"],
base_url="https://api.kittenml.com/v1",
)
# This call waits for the complete audio response.
audio = client.audio.speech.create(
model="kitten-tts-2-latest",
input="Hello from the KittenML API.",
voice="eleanor_somber_female_32",
response_format="mp3",
speed=1.0,
)
# Preserve x-request-id in logs so support can correlate the generation.
audio.write_to_file(Path("kitten-output.mp3"))
print(f"request_id: {audio.response.headers.get('x-request-id')}")
JavaScript with the OpenAI SDK
import {writeFile} from "node:fs/promises";
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.KITTENML_API_KEY,
baseURL: "https://api.kittenml.com/v1",
});
// withResponse() exposes x-request-id alongside the audio response.
const {data: audio, request_id: requestId} =
await client.audio.speech.create({
model: "kitten-tts-2-latest",
input: "Hello from the KittenML API.",
voice: "eleanor_somber_female_32",
response_format: "mp3",
speed: 1.0,
}).withResponse();
// Complete-response mode buffers the finished MP3 before saving it.
await writeFile("kitten-output.mp3", Buffer.from(await audio.arrayBuffer()));
console.log(`request_id: ${requestId}`);
Quality over speed
By default, Kitten TTS 2 is tuned for speed, for the lowest latency. For
quality over speed, add "stream": false. The API then returns one complete
file with higher audio quality, sent once the whole text has been generated.
It works with every response_format and cannot be combined with
stream_format: "sse".
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "kitten-tts-2-latest",
"input": "The harbor lights came on one by one as the fog rolled in.",
"voice": "eleanor_somber_female_32",
"response_format": "wav",
"stream": false
}' \
--output kitten-complete.wav
With the OpenAI SDKs, pass it as an extra field: extra_body={"stream": False}
in Python, or stream: false in JavaScript.
Models and voices
| Model | Parameters | Relative trade-off | Recommended use |
|---|---|---|---|
kitten-tts-2-latest | 1.7B ternary | Kitten TTS 2, the latest hosted architecture | New cloud integrations |
kitten-tts-nano-0.8 | 15M | Lowest compute and fastest generation | High-volume or latency-sensitive speech |
kitten-tts-micro-0.8 | 40M | Balanced speed and quality | General-purpose default |
kitten-tts-mini-0.8 | 80M | Highest quality and highest compute | Voice quality is the priority |
Name Kitten TTS 2 as kitten-tts-2-latest. Integrations that still send its
earlier model ID keep working; see the legacy note in
Models. To see which features each model
supports, compare them in Models.
Kitten TTS 2 has 47 built-in voices. Select one by its voice ID, for example
eleanor_somber_female_32, the default. The voices fall into four groups:
| Voices | What they are |
|---|---|
01 – 09 | Character voices, such as saoirse_joyful_irish_female_07 |
10 – 17 | The eight KittenTTS 0.8 speakers, such as bella_10 |
18 – 37 | Descriptive voices, such as maeve_cozy_female_22 |
38 – 47 | One voice per language, such as spanish_41 |
List voices has the full table. Discover
the live catalog, including language metadata and your organization's active
saved voices, with GET /v1/voices.
The three KittenTTS 0.8 models use the classic eight-voice catalog:
| Voice | Character |
|---|---|
Bella | Warm and expressive |
Jasper | Clear and conversational |
Luna | Calm and smooth |
Bruno | Deep and steady |
Rosie | Bright and friendly |
Hugo | Authoritative |
Kiki | Lively and energetic |
Leo | Relaxed and natural |
Voice IDs are case-sensitive. A Kitten TTS 2 voice is not accepted by a 0.8
model, and a 0.8 voice is not accepted by Kitten TTS 2: Bella selects the
KittenTTS 0.8 voice, and its Kitten TTS 2 counterpart is
bella_10. Descriptions indicate the intended character
rather than guaranteed prosody; preview voices using your own text,
punctuation, and speed before choosing one for production.
Kitten TTS 2 can clone an uploaded reference in the same speech request, or
reuse a saved voice_... ID. See Clone a voice.
Language coverage
Kitten TTS 2 supports Arabic, German, English, Spanish, French, Italian,
Portuguese, Russian, and Simplified Chinese through the matching built-in
voice, 38 through 46, and Hindi through hindi_47.
Select the language by setting voice; there is no separate language request
field.
Written-to-spoken normalization, which expands numbers, dates, currencies,
units, and abbreviations into words, is English-only. It runs for voices whose
language is en in GET /v1/voices, and for the other
language voices when the text itself is English: an English sentence
containing 98.2% sent to german_39 is read with the number in
English words. Text in another language, including Arabic, Cyrillic,
Devanagari, and CJK text, receives only punctuation and character cleanup, so
98.2% or St. in German text is not expanded. Write numbers and
abbreviations out as words in that language, and test dates and domain
terminology in the selected language.
English text also gets light formatting before it is spoken: line breaks inside
a sentence are joined, a sentence end with no space after it (noon.It) still
starts a new sentence, and an ellipsis (...) is read as a short pause.
The hosted KittenTTS 0.8 models and saved custom voices remain English-only.
Emotion and sound tags
KittenTTS 0.8 models read bracketed text aloud as ordinary words. See feature support by model.
Kitten TTS 2 reads three kinds of inline markup in input:
| Markup | Supported values | Where to put it |
|---|---|---|
| Emotion tag | [joyful], [excited], [tender], [sad], [angry], [nervous], [stern], [surprised], [contemplative], [mundane] | Before the sentence it should color |
| Sound tag | <laugh>, <giggle>, <sigh>, <gasp>, <sob>, <scoff>, <growl>, <um>, <gulp>, <pause> | Where the sound should happen |
| Emphasis | (((word or short phrase))) | Around the words to stress |
{
"model": "kitten-tts-2-latest",
"input": "[excited] We shipped it! <laugh> It is (((really))) fast.",
"voice": "maeve_cozy_female_22"
}
Tags are case-insensitive. Tags and the emphasis parentheses are not read
aloud. Only the values in the table are tags: any other bracketed text, such
as [happy] or <laughs>, is read aloud as ordinary words. Emphasis needs
exactly three parentheses on each side around at most 80 characters on one
line.
This feature is experimental. Its effect may be subtle and can alter pacing or introduce unintended vocal events, so it does not guarantee a specific emotion. Preview the exact voice and text before using markup in production.
Decoding controls
KittenTTS 0.8 models accept these fields but ignore them. See feature support by model.
Kitten TTS 2 generates speech as a sequence of audio tokens, drawing each token at random from the model's predictions. These optional fields control that draw. They are the decoding options of the Kitten TTS demo, with the same names, presets, and ranges. KittenTTS 0.8 models do not use them.
| Field | Default | Accepted values | Effect |
|---|---|---|---|
mode | expressive | stable or expressive | Picks the preset for the four values below |
temperature | From mode | 0.3 through 2.0 | Below 1 sharpens the predictions; above 1 flattens them |
top_k | From mode | Integer 0 through 200 | Keeps only the k most likely audio tokens; 0 turns the filter off |
top_p | From mode | 0.5 through 1.0 | Keeps the smallest set of tokens holding this much probability; 1.0 turns the filter off |
min_p | From mode | 0 through 0.5 | Drops tokens less likely than this fraction of the most likely one; 0 turns the filter off |
max_new_tokens | None | Integer 500 through 12000 | Upper limit on audio tokens per text segment; see below |
mode | temperature | top_p | top_k | min_p |
|---|---|---|---|---|
stable | 0.8 | 0.8 | 50 | 0 |
expressive (default) | 0.9 | 0.9 | 50 | 0 |
{
"model": "kitten-tts-2-latest",
"input": "The harbor lights came on one by one.",
"voice": "maeve_cozy_female_22",
"mode": "expressive",
"top_k": 40
}
- A field you send overrides the preset. Every field you leave out comes from
mode, or fromexpressivewhenmodeis also left out. The example above samples at temperature0.9,top_p0.9, andmin_p0withtop_k40.nullcounts as not sent. - The filters run in this order:
temperature,top_k,top_p, thenmin_p. - JSON values must be numbers, not strings.
top_kandmax_new_tokensmust be whole numbers;50and50.0are both accepted. - An out-of-range or mistyped value is rejected before synthesis starts with
400,"code": "invalid_request_error", andparamnaming the field. For example,{"top_p": 0.2}returnstop_p must be a number between 0.5 and 1.0. - The same names work as form fields for one-request cloning and as query parameters for input streaming.
max_new_tokens limits each text segment, not the whole request. The model
produces about 25 audio tokens per second of audio, pauses included. Long input
is spoken one segment at a time: usually a sentence, or a chunk of up to about
530 characters for marked-up text. Each segment already has a limit
proportional to its length, at most about 30 seconds of speech for a sentence
and about 90 seconds for a marked-up chunk. max_new_tokens can only lower
that limit. A value below what a segment needs cuts that segment off
mid-speech (500 tokens is about 20 seconds), and the segment is not
regenerated. Values above the service limit change nothing.
Every request is a fresh take: the same text with the same settings sounds a little different each time. Settings far from the presets reduce quality:
- Temperature above
1.0, especially withtop_k0andtop_pnear1.0, lets unlikely tokens through, which produces garbled or repeated sounds. - Very tight settings, such as temperature
0.3withtop_k1,top_p0.5, ormin_p0.5, approach always taking the single most likely token. The model rarely ends a sentence that way, so a segment can run on to its token limit.
Start from a preset and change one value at a time.
Output formats
| Format | Content type | File extension | Notes |
|---|---|---|---|
mp3 | audio/mpeg | .mp3 | Compact default |
wav | audio/wav | .wav | 24 kHz mono PCM in a WAV container |
flac | audio/flac | .flac | Lossless compressed audio |
aac | audio/aac | .aac | AAC audio |
opus | audio/ogg | .opus or .ogg | Opus in an Ogg container |
pcm | audio/L16; rate=24000; channels=1 | .pcm | Headerless 24 kHz mono, little-endian PCM16 |
The pcm bytes are little-endian. Some media libraries interpret the standard
audio/L16 name as big-endian; use wav for portable PCM interchange, or
explicitly configure raw audio as signed 16-bit little-endian.
Successful response
The response body contains the selected audio format. Useful headers include:
| Header | Meaning |
|---|---|
x-request-id | Identifier for logs and support |
X-Model | The model that served the request, for example kitten-tts-2-latest. A request that omits model is served by KittenTTS 0.8 and reports kitten-tts-mini-0.8. See Models |
X-Stream-Format | audio |
Cache-Control | no-store |
Header names are case-insensitive.
Runnable Python and JavaScript clients are available in the complete-response TTS examples.
Errors and limits
Authentication and model errors use the common KittenML error envelope.
Parameter validation errors use a detail field. See
Errors and retries.
An organization may run two TTS generations concurrently. A request above the
limit returns 429 with Retry-After: 1. See
Rate limits and timeouts.