Models
KittenASR Enhanced listens; Kitten TTS 2 and three KittenTTS 0.8 models speak. Model IDs are case-sensitive, so send the exact ID shown below.
Speech to text
| Model ID | Input | Output |
|---|---|---|
kittenasr-enhanced-preview | Speech audio | Clean text plus emotion-aware transcript metadata |
The same model ID is used for upload and realtime STT.
GET /v1/models reports this model with the identifier
kitten-asr-pro-enhanced. That identifier and kittenasr-enhanced-preview
name the same hosted model and are interchangeable in every request.
For migrations from OpenAI clients, whisper-1, gpt-4o-transcribe,
gpt-4o-mini-transcribe, and gpt-realtime-whisper are accepted compatibility
aliases for the same hosted KittenASR Enhanced model. New integrations should use
kittenasr-enhanced-preview; the aliases are not separate models.
Text to speech
Two model families speak. Kitten TTS 2 is the current hosted model: 47
voices, nine languages, voice cloning, emotion markup, and decoding controls.
KittenTTS 0.8 is the earlier, much smaller family: three sizes, eight
English voices, and none of the Kitten TTS 2 features. Both use the same
endpoints and the same request shape, so switching is a change of model and
voice.
Choose a model
| If you need | Use |
|---|---|
| The best quality, more voices, or anything other than English | kitten-tts-2-latest |
| Voice cloning, emotion and sound tags, or decoding controls | kitten-tts-2-latest |
| The smallest model, or the same voices you run on device with the KittenTTS SDKs | A KittenTTS 0.8 model |
A request without model is served by kitten-tts-mini-0.8, not Kitten TTS 2,
so always name the model you want.
| Model ID | Family | Size | Built-in voices | Languages |
|---|---|---|---|---|
kitten-tts-2-latest | Kitten TTS 2 | 1.7B ternary | 47 voices; default eleanor_somber_female_32 | English, Arabic, German, Spanish, French, Italian, Portuguese, Russian, Simplified Chinese, Hindi |
kitten-tts-mini-0.8 | KittenTTS 0.8 | 80M | 8 classic voices; default Bella | English |
kitten-tts-micro-0.8 | KittenTTS 0.8 | 40M | 8 classic voices; default Bella | English |
kitten-tts-nano-0.8 | KittenTTS 0.8 | 15M | 8 classic voices; default Bella | English |
kitten-tts-2-latest always serves the current Kitten TTS 2 release, so its
speed and output can improve between releases.
Feature support
| Feature | Kitten TTS 2 | KittenTTS 0.8 |
|---|---|---|
| Complete response, binary and SSE streaming | Yes | Yes |
| Realtime WebSocket input streaming | Yes | Yes |
Output formats mp3, wav, flac, aac, opus, pcm | Yes | Yes |
speed from 0.25 to 4.0 | Yes | Yes |
| Up to 40,000 input characters per HTTP request | Yes | Yes |
| Numbers, dates, and units read as words | English voices, and English text on the language voices | English |
| Voice cloning: one-request reference, saved voices, shared voices | Yes | No |
List voices with GET /v1/voices | Yes | No; the eight voices are fixed |
Emotion and sound tags, (((emphasis))) | Yes | No; bracketed text is read aloud |
Decoding controls: mode, temperature, top_p, top_k, min_p, max_new_tokens | Yes | No; the fields are ignored |
Voice names belong to one family: a Kitten TTS 2 voice is rejected by a 0.8
model and a 0.8 voice by Kitten TTS 2. The eight 0.8 speakers are also
available in Kitten TTS 2 as bella_10 through
leo_17.
For pricing, see the KittenML platform.
kitten-tts-0.9 still worksKitten TTS 2 was previously published as KittenTTS 0.9, with the model ID
kitten-tts-0.9. That ID is now a legacy alias of kitten-tts-2-latest: it is
still accepted everywhere a model is named (request bodies, form fields, and
the model query parameter), and it selects the same model with the same
voices and audio. GET /v1/models lists both IDs. A response names the model
with the ID you sent (for example, in the X-Model header), so compare model
names against both IDs, never just one. New integrations should use
kitten-tts-2-latest.
List models
curl --fail-with-body -sS \
https://api.kittenml.com/v1/models \
-H "Authorization: Bearer $KITTENML_API_KEY"
Use the catalog above as the supported inference contract. The model-list endpoint currently requires a key with ASR access and paid-ASR eligibility.