Speech with a little more personality
Hey, kittens. Welcome in. KittenML makes small speech models that pay attention to more than the words: KittenASR Enhanced preserves expression in a transcript, while Kitten TTS 2 and the KittenTTS 0.8 family turn text into a voice with one of four hosted models.
Every hosted API is available from one hostname:
https://api.kittenml.com
Choose an endpoint
| Protocol | Endpoint | Use it for |
|---|---|---|
| HTTPS | POST /v1/audio/transcriptions | Transcribe a complete audio file |
| WebSocket | WSS /v1/realtime | Stream audio and receive partial transcripts |
| WebRTC | /v1/realtime/client_secrets + /v1/realtime/calls | Stream a browser microphone without exposing a permanent API key |
| HTTPS | POST /v1/audio/speech | Generate complete speech or stream its output |
| WebSocket | WSS /v1/tts/realtime | Send text incrementally and receive PCM audio |
Upload STT, realtime STT, and TTS are the public inference operations.
Realtime STT supports both server-side WebSocket clients and browser WebRTC
clients. TTS supports complete-response and streaming-output use through the
same OpenAI-compatible HTTP endpoint, plus a KittenML WebSocket for incremental
text input. The older WSS /v1/stt/realtime route remains a compatibility
alias for existing STT clients.
All inference endpoints require a KittenML API key. The KittenTTS 0.8 models are free, but authentication is still required.
The fastest way to learn the API is to run an example, change it, and see what happens. Test it. Break it. If something sounds strange, tell us what you heard.
Create or manage an API key on the KittenML platform.
Try speech to text
export KITTENML_API_KEY="sk_kitten_live_..."
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/transcriptions \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F "file=@speech.wav" \
-F "model=kittenasr-enhanced-preview" \
-F "response_format=json"
An abridged response looks like this:
{
"text": "A cold, lucid indifference reigned in his soul.",
"enriched_text": "[low][slow]A (cold), [pause_short] (lucid) (indifference) reigned in his (soul).[/slow][/low]",
"accent": "General American",
"request_id": "<request-id>"
}
text is the clean transcript. enriched_text, accent, and tags are
KittenML extensions that preserve vocal delivery information when the model
detects it.
Try text to speech
# Stream the binary MP3 response directly into a local file.
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "kitten-tts-2-latest",
"input": "Hello from the KittenML API.",
"voice": "eleanor_somber_female_32",
"response_format": "mp3",
"stream_format": "audio"
}' \
--output kitten-output.mp3