Skip to main content

Generate complete speech

POST https://api.kittenml.com/v1/audio/speech

Send the complete text in one request and receive one complete audio result. This is the standard OpenAI-compatible speech workflow. The examples below fully read the response before your application opens or processes the saved audio.

Kitten TTS 2 (kitten-tts-2-latest) and all three KittenTTS 0.8 models are available. A valid API key with the tts:generate permission is required so the generation can appear in the owning organization's usage history.

Need to play bytes as soon as they arrive? See Stream speech output. If the text itself arrives a piece at a time, see Stream text input.

Before you begin​

Create an account-backed key in the KittenML API key dashboard, give it the tts:generate permission, and export it in your shell:

export KITTENML_API_KEY="sk_kitten_live_..."

Install the official OpenAI SDK for your language before running the SDK examples:

pip install openai
npm install openai

Request body​

FieldTypeRequiredDefaultAccepted values
modelstringNokitten-tts-mini-0.8kitten-tts-2-latest, kitten-tts-nano-0.8, kitten-tts-micro-0.8, or kitten-tts-mini-0.8
inputstringYes—Non-empty text, at most 40,000 characters
voicestring or objectNoModel-specificA built-in voice, a saved {"id":"voice_..."} voice, or a Kitten TTS 2 inline reference object
response_formatstringNomp3mp3, wav, flac, aac, opus, or pcm
speednumberNo1.00.25 through 4.0
streambooleanNotrueKitten TTS 2 only: false returns one complete file with higher audio quality. See Quality over speed
modestringNoexpressiveKitten TTS 2 only: stable or expressive. See Decoding controls
temperaturenumberNoFrom modeKitten TTS 2 only: 0.3 through 2.0
top_pnumberNoFrom modeKitten TTS 2 only: 0.5 through 1.0
top_kintegerNoFrom modeKitten TTS 2 only: 0 through 200 (0 turns it off)
min_pnumberNoFrom modeKitten TTS 2 only: 0 through 0.5
max_new_tokensintegerNoNoneKitten TTS 2 only: 500 through 12000

For ordinary JSON requests, omit voice to use the selected model's built-in default; do not send "voice": null.

JSON requestEffective default
model: "kitten-tts-2-latest", no voiceeleanor_somber_female_32
Any KittenTTS 0.8 model, no voiceBella
No model and no voicekitten-tts-mini-0.8 with Bella

A request without model is served by kitten-tts-mini-0.8, so every Kitten TTS 2 request must name kitten-tts-2-latest.

Multipart speech requests are reserved for one-request cloning and require reference_audio. Use ordinary JSON when you want a default built-in voice.

The JSON body may be at most 256 KiB and the request may run for at most 30 minutes. For predictable latency, keep individual generations around 400–5,000 characters. stream_format may be omitted for this workflow; its OpenAI-compatible default is audio.

cURL​

# Complete-response mode writes the finished MP3 after generation completes.
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "kitten-tts-2-latest",
"input": "Hello! This audio was generated by KittenTTS.",
"voice": "eleanor_somber_female_32",
"response_format": "mp3",
"speed": 1.0
}' \
--output kitten-output.mp3

The command returns after the response is complete and saves it as kitten-output.mp3. The success body is audio, not JSON; do not pipe it to jq or print it in a terminal.

-L is not optional. When a generation has to wait for hosting capacity, the platform answers with a redirect to a continuation URL instead of the audio, exactly as it does for upload transcription. Without -L, cURL treats that redirect as the final response, exits 0, and writes an empty file — a silent failure rather than an error. Every HTTP client you use with this endpoint must follow redirects; the official OpenAI SDKs already do.

Python with the OpenAI SDK​

import os
from pathlib import Path
from openai import OpenAI

client = OpenAI(
api_key=os.environ["KITTENML_API_KEY"],
base_url="https://api.kittenml.com/v1",
)

# This call waits for the complete audio response.
audio = client.audio.speech.create(
model="kitten-tts-2-latest",
input="Hello from the KittenML API.",
voice="eleanor_somber_female_32",
response_format="mp3",
speed=1.0,
)
# Preserve x-request-id in logs so support can correlate the generation.
audio.write_to_file(Path("kitten-output.mp3"))
print(f"request_id: {audio.response.headers.get('x-request-id')}")

JavaScript with the OpenAI SDK​

import {writeFile} from "node:fs/promises";
import OpenAI from "openai";

const client = new OpenAI({
apiKey: process.env.KITTENML_API_KEY,
baseURL: "https://api.kittenml.com/v1",
});

// withResponse() exposes x-request-id alongside the audio response.
const {data: audio, request_id: requestId} =
await client.audio.speech.create({
model: "kitten-tts-2-latest",
input: "Hello from the KittenML API.",
voice: "eleanor_somber_female_32",
response_format: "mp3",
speed: 1.0,
}).withResponse();

// Complete-response mode buffers the finished MP3 before saving it.
await writeFile("kitten-output.mp3", Buffer.from(await audio.arrayBuffer()));
console.log(`request_id: ${requestId}`);

Quality over speed​

By default, Kitten TTS 2 is tuned for speed, for the lowest latency. For quality over speed, add "stream": false. The API then returns one complete file with higher audio quality, sent once the whole text has been generated. It works with every response_format and cannot be combined with stream_format: "sse".

curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "kitten-tts-2-latest",
"input": "The harbor lights came on one by one as the fog rolled in.",
"voice": "eleanor_somber_female_32",
"response_format": "wav",
"stream": false
}' \
--output kitten-complete.wav

With the OpenAI SDKs, pass it as an extra field: extra_body={"stream": False} in Python, or stream: false in JavaScript.

Models and voices​

ModelParametersRelative trade-offRecommended use
kitten-tts-2-latest1.7B ternaryKitten TTS 2, the latest hosted architectureNew cloud integrations
kitten-tts-nano-0.815MLowest compute and fastest generationHigh-volume or latency-sensitive speech
kitten-tts-micro-0.840MBalanced speed and qualityGeneral-purpose default
kitten-tts-mini-0.880MHighest quality and highest computeVoice quality is the priority

Name Kitten TTS 2 as kitten-tts-2-latest. Integrations that still send its earlier model ID keep working; see the legacy note in Models. To see which features each model supports, compare them in Models.

Kitten TTS 2 has 47 built-in voices. Select one by its voice ID, for example eleanor_somber_female_32, the default. The voices fall into four groups:

VoicesWhat they are
01 – 09Character voices, such as saoirse_joyful_irish_female_07
10 – 17The eight KittenTTS 0.8 speakers, such as bella_10
18 – 37Descriptive voices, such as maeve_cozy_female_22
38 – 47One voice per language, such as spanish_41

List voices has the full table. Discover the live catalog, including language metadata and your organization's active saved voices, with GET /v1/voices.

The three KittenTTS 0.8 models use the classic eight-voice catalog:

VoiceCharacter
BellaWarm and expressive
JasperClear and conversational
LunaCalm and smooth
BrunoDeep and steady
RosieBright and friendly
HugoAuthoritative
KikiLively and energetic
LeoRelaxed and natural

Voice IDs are case-sensitive. A Kitten TTS 2 voice is not accepted by a 0.8 model, and a 0.8 voice is not accepted by Kitten TTS 2: Bella selects the KittenTTS 0.8 voice, and its Kitten TTS 2 counterpart is bella_10. Descriptions indicate the intended character rather than guaranteed prosody; preview voices using your own text, punctuation, and speed before choosing one for production.

Kitten TTS 2 can clone an uploaded reference in the same speech request, or reuse a saved voice_... ID. See Clone a voice.

Language coverage​

Kitten TTS 2 supports Arabic, German, English, Spanish, French, Italian, Portuguese, Russian, and Simplified Chinese through the matching built-in voice, 38 through 46, and Hindi through hindi_47. Select the language by setting voice; there is no separate language request field.

Written-to-spoken normalization, which expands numbers, dates, currencies, units, and abbreviations into words, is English-only. It runs for voices whose language is en in GET /v1/voices, and for the other language voices when the text itself is English: an English sentence containing 98.2% sent to german_39 is read with the number in English words. Text in another language, including Arabic, Cyrillic, Devanagari, and CJK text, receives only punctuation and character cleanup, so 98.2% or St. in German text is not expanded. Write numbers and abbreviations out as words in that language, and test dates and domain terminology in the selected language.

English text also gets light formatting before it is spoken: line breaks inside a sentence are joined, a sentence end with no space after it (noon.It) still starts a new sentence, and an ellipsis (...) is read as a short pause.

The hosted KittenTTS 0.8 models and saved custom voices remain English-only.

Emotion and sound tags​

Kitten TTS 2 only

KittenTTS 0.8 models read bracketed text aloud as ordinary words. See feature support by model.

Kitten TTS 2 reads three kinds of inline markup in input:

MarkupSupported valuesWhere to put it
Emotion tag[joyful], [excited], [tender], [sad], [angry], [nervous], [stern], [surprised], [contemplative], [mundane]Before the sentence it should color
Sound tag<laugh>, <giggle>, <sigh>, <gasp>, <sob>, <scoff>, <growl>, <um>, <gulp>, <pause>Where the sound should happen
Emphasis(((word or short phrase)))Around the words to stress
{
"model": "kitten-tts-2-latest",
"input": "[excited] We shipped it! <laugh> It is (((really))) fast.",
"voice": "maeve_cozy_female_22"
}

Tags are case-insensitive. Tags and the emphasis parentheses are not read aloud. Only the values in the table are tags: any other bracketed text, such as [happy] or <laughs>, is read aloud as ordinary words. Emphasis needs exactly three parentheses on each side around at most 80 characters on one line.

This feature is experimental. Its effect may be subtle and can alter pacing or introduce unintended vocal events, so it does not guarantee a specific emotion. Preview the exact voice and text before using markup in production.

Decoding controls​

Kitten TTS 2 only

KittenTTS 0.8 models accept these fields but ignore them. See feature support by model.

Kitten TTS 2 generates speech as a sequence of audio tokens, drawing each token at random from the model's predictions. These optional fields control that draw. They are the decoding options of the Kitten TTS demo, with the same names, presets, and ranges. KittenTTS 0.8 models do not use them.

FieldDefaultAccepted valuesEffect
modeexpressivestable or expressivePicks the preset for the four values below
temperatureFrom mode0.3 through 2.0Below 1 sharpens the predictions; above 1 flattens them
top_kFrom modeInteger 0 through 200Keeps only the k most likely audio tokens; 0 turns the filter off
top_pFrom mode0.5 through 1.0Keeps the smallest set of tokens holding this much probability; 1.0 turns the filter off
min_pFrom mode0 through 0.5Drops tokens less likely than this fraction of the most likely one; 0 turns the filter off
max_new_tokensNoneInteger 500 through 12000Upper limit on audio tokens per text segment; see below
modetemperaturetop_ptop_kmin_p
stable0.80.8500
expressive (default)0.90.9500
{
"model": "kitten-tts-2-latest",
"input": "The harbor lights came on one by one.",
"voice": "maeve_cozy_female_22",
"mode": "expressive",
"top_k": 40
}
  • A field you send overrides the preset. Every field you leave out comes from mode, or from expressive when mode is also left out. The example above samples at temperature 0.9, top_p 0.9, and min_p 0 with top_k 40. null counts as not sent.
  • The filters run in this order: temperature, top_k, top_p, then min_p.
  • JSON values must be numbers, not strings. top_k and max_new_tokens must be whole numbers; 50 and 50.0 are both accepted.
  • An out-of-range or mistyped value is rejected before synthesis starts with 400, "code": "invalid_request_error", and param naming the field. For example, {"top_p": 0.2} returns top_p must be a number between 0.5 and 1.0.
  • The same names work as form fields for one-request cloning and as query parameters for input streaming.

max_new_tokens limits each text segment, not the whole request. The model produces about 25 audio tokens per second of audio, pauses included. Long input is spoken one segment at a time: usually a sentence, or a chunk of up to about 530 characters for marked-up text. Each segment already has a limit proportional to its length, at most about 30 seconds of speech for a sentence and about 90 seconds for a marked-up chunk. max_new_tokens can only lower that limit. A value below what a segment needs cuts that segment off mid-speech (500 tokens is about 20 seconds), and the segment is not regenerated. Values above the service limit change nothing.

Every request is a fresh take: the same text with the same settings sounds a little different each time. Settings far from the presets reduce quality:

  • Temperature above 1.0, especially with top_k 0 and top_p near 1.0, lets unlikely tokens through, which produces garbled or repeated sounds.
  • Very tight settings, such as temperature 0.3 with top_k 1, top_p 0.5, or min_p 0.5, approach always taking the single most likely token. The model rarely ends a sentence that way, so a segment can run on to its token limit.

Start from a preset and change one value at a time.

Output formats​

FormatContent typeFile extensionNotes
mp3audio/mpeg.mp3Compact default
wavaudio/wav.wav24 kHz mono PCM in a WAV container
flacaudio/flac.flacLossless compressed audio
aacaudio/aac.aacAAC audio
opusaudio/ogg.opus or .oggOpus in an Ogg container
pcmaudio/L16; rate=24000; channels=1.pcmHeaderless 24 kHz mono, little-endian PCM16

The pcm bytes are little-endian. Some media libraries interpret the standard audio/L16 name as big-endian; use wav for portable PCM interchange, or explicitly configure raw audio as signed 16-bit little-endian.

Successful response​

The response body contains the selected audio format. Useful headers include:

HeaderMeaning
x-request-idIdentifier for logs and support
X-ModelThe model that served the request, for example kitten-tts-2-latest. A request that omits model is served by KittenTTS 0.8 and reports kitten-tts-mini-0.8. See Models
X-Stream-Formataudio
Cache-Controlno-store

Header names are case-insensitive.

Runnable Python and JavaScript clients are available in the complete-response TTS examples.

Errors and limits​

Authentication and model errors use the common KittenML error envelope. Parameter validation errors use a detail field. See Errors and retries.

An organization may run two TTS generations concurrently. A request above the limit returns 429 with Retry-After: 1. See Rate limits and timeouts.