Skip to main content

Models

KittenASR Enhanced listens; Kitten TTS 2 and three KittenTTS 0.8 models speak. Model IDs are case-sensitive, so send the exact ID shown below.

Speech to text​

Model IDInputOutput
kittenasr-enhanced-previewSpeech audioClean text plus emotion-aware transcript metadata

The same model ID is used for upload and realtime STT.

GET /v1/models reports this model with the identifier kitten-asr-pro-enhanced. That identifier and kittenasr-enhanced-preview name the same hosted model and are interchangeable in every request.

For migrations from OpenAI clients, whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-realtime-whisper are accepted compatibility aliases for the same hosted KittenASR Enhanced model. New integrations should use kittenasr-enhanced-preview; the aliases are not separate models.

Text to speech​

Two model families speak. Kitten TTS 2 is the current hosted model: 47 voices, nine languages, voice cloning, emotion markup, and decoding controls. KittenTTS 0.8 is the earlier, much smaller family: three sizes, eight English voices, and none of the Kitten TTS 2 features. Both use the same endpoints and the same request shape, so switching is a change of model and voice.

Choose a model​

If you needUse
The best quality, more voices, or anything other than Englishkitten-tts-2-latest
Voice cloning, emotion and sound tags, or decoding controlskitten-tts-2-latest
The smallest model, or the same voices you run on device with the KittenTTS SDKsA KittenTTS 0.8 model

A request without model is served by kitten-tts-mini-0.8, not Kitten TTS 2, so always name the model you want.

Model IDFamilySizeBuilt-in voicesLanguages
kitten-tts-2-latestKitten TTS 21.7B ternary47 voices; default eleanor_somber_female_32English, Arabic, German, Spanish, French, Italian, Portuguese, Russian, Simplified Chinese, Hindi
kitten-tts-mini-0.8KittenTTS 0.880M8 classic voices; default BellaEnglish
kitten-tts-micro-0.8KittenTTS 0.840M8 classic voices; default BellaEnglish
kitten-tts-nano-0.8KittenTTS 0.815M8 classic voices; default BellaEnglish

kitten-tts-2-latest always serves the current Kitten TTS 2 release, so its speed and output can improve between releases.

Feature support​

FeatureKitten TTS 2KittenTTS 0.8
Complete response, binary and SSE streamingYesYes
Realtime WebSocket input streamingYesYes
Output formats mp3, wav, flac, aac, opus, pcmYesYes
speed from 0.25 to 4.0YesYes
Up to 40,000 input characters per HTTP requestYesYes
Numbers, dates, and units read as wordsEnglish voices, and English text on the language voicesEnglish
Voice cloning: one-request reference, saved voices, shared voicesYesNo
List voices with GET /v1/voicesYesNo; the eight voices are fixed
Emotion and sound tags, (((emphasis)))YesNo; bracketed text is read aloud
Decoding controls: mode, temperature, top_p, top_k, min_p, max_new_tokensYesNo; the fields are ignored

Voice names belong to one family: a Kitten TTS 2 voice is rejected by a 0.8 model and a 0.8 voice by Kitten TTS 2. The eight 0.8 speakers are also available in Kitten TTS 2 as bella_10 through leo_17.

For pricing, see the KittenML platform.

kitten-tts-0.9 still works

Kitten TTS 2 was previously published as KittenTTS 0.9, with the model ID kitten-tts-0.9. That ID is now a legacy alias of kitten-tts-2-latest: it is still accepted everywhere a model is named (request bodies, form fields, and the model query parameter), and it selects the same model with the same voices and audio. GET /v1/models lists both IDs. A response names the model with the ID you sent (for example, in the X-Model header), so compare model names against both IDs, never just one. New integrations should use kitten-tts-2-latest.

List models​

curl --fail-with-body -sS \
https://api.kittenml.com/v1/models \
-H "Authorization: Bearer $KITTENML_API_KEY"

Use the catalog above as the supported inference contract. The model-list endpoint currently requires a key with ASR access and paid-ASR eligibility.