Skip to main content

Speech with a little more personality

Hey, kittens. Welcome in. KittenML makes small speech models that pay attention to more than the words: KittenASR Enhanced preserves expression in a transcript, while Kitten TTS 2 and the KittenTTS 0.8 family turn text into a voice with one of four hosted models.

Every hosted API is available from one hostname:

https://api.kittenml.com

Choose an endpoint​

ProtocolEndpointUse it for
HTTPSPOST /v1/audio/transcriptionsTranscribe a complete audio file
WebSocketWSS /v1/realtimeStream audio and receive partial transcripts
WebRTC/v1/realtime/client_secrets + /v1/realtime/callsStream a browser microphone without exposing a permanent API key
HTTPSPOST /v1/audio/speechGenerate complete speech or stream its output
WebSocketWSS /v1/tts/realtimeSend text incrementally and receive PCM audio

Upload STT, realtime STT, and TTS are the public inference operations. Realtime STT supports both server-side WebSocket clients and browser WebRTC clients. TTS supports complete-response and streaming-output use through the same OpenAI-compatible HTTP endpoint, plus a KittenML WebSocket for incremental text input. The older WSS /v1/stt/realtime route remains a compatibility alias for existing STT clients.

All inference endpoints require a KittenML API key. The KittenTTS 0.8 models are free, but authentication is still required.

The fastest way to learn the API is to run an example, change it, and see what happens. Test it. Break it. If something sounds strange, tell us what you heard.

Create or manage an API key on the KittenML platform.

Try speech to text​

export KITTENML_API_KEY="sk_kitten_live_..."

curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/transcriptions \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F "file=@speech.wav" \
-F "model=kittenasr-enhanced-preview" \
-F "response_format=json"

An abridged response looks like this:

{
"text": "A cold, lucid indifference reigned in his soul.",
"enriched_text": "[low][slow]A (cold), [pause_short] (lucid) (indifference) reigned in his (soul).[/slow][/low]",
"accent": "General American",
"request_id": "<request-id>"
}

text is the clean transcript. enriched_text, accent, and tags are KittenML extensions that preserve vocal delivery information when the model detects it.

Try text to speech​

# Stream the binary MP3 response directly into a local file.
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "kitten-tts-2-latest",
"input": "Hello from the KittenML API.",
"voice": "eleanor_somber_female_32",
"response_format": "mp3",
"stream_format": "audio"
}' \
--output kitten-output.mp3

Next steps​