OpenAI compatibility
KittenML supports the common OpenAI SDK workflows for upload transcription, speech generation, and realtime transcription connections.
Set this base URL for OpenAI HTTP SDKs:
https://api.kittenml.com/v1
Do not set the SDK base URL to a complete endpoint such as
/v1/audio/transcriptions; the SDK appends that path itself.
Compatibility matrix
| Capability | Status | Notes |
|---|---|---|
POST /v1/audio/transcriptions | Supported | Works with OpenAI Python and JavaScript clients |
model, file, language, response_format, stream | Supported | Use the documented values |
json, verbose_json, text, srt, vtt | Supported | Includes KittenML metadata in JSON formats |
| Upload SSE | Supported | Finite delta/done stream ending in [DONE] |
POST /v1/audio/speech | Supported | Complete-response SDK patterns and progressive binary consumption are supported; SSE is optional |
model, input, voice, response_format, speed, stream_format | Supported | Use KittenML model IDs and voices |
stream in /v1/audio/speech | KittenML extension | Kitten TTS 2: false returns one complete file for quality over speed. Pass extra_body={"stream": False} in Python or stream: false in JavaScript |
/v1/audio/voice_consents lifecycle | Supported for Kitten TTS 2 | OpenAI-compatible multipart fields and audio.voice_consent objects; requires a durable organization or registered legacy key |
POST /v1/audio/voices with consent | Supported for Kitten TTS 2 | Preserves OpenAI's name, consent, and audio_sample multipart shape and audio.voice response |
POST /v1/audio/voices without consent | KittenML extension | Omit consent to save a reusable voice directly; the response remains an audio.voice object |
DELETE /v1/audio/voices/{voice_id} | KittenML extension | Revokes and deletes one owner-scoped voice without deleting a consent or sibling voices |
Saved custom voice in /v1/audio/speech | Supported for Kitten TTS 2 | Pass {"id":"voice_..."} as voice; complete binary and SSE output are supported |
Inline reference audio in /v1/audio/speech | KittenML extension | Send a multipart reference_audio file without creating consent or voice resources, or use a Base64 JSON voice object |
Saved custom voice in WSS /v1/tts/realtime | KittenML extension | Put the owner-scoped voice_... ID in the query; both consent-bound and consent-free saved voices are supported |
| Realtime event names and common fields | Compatible style | Includes KittenML extension fields |
| Official OpenAI Realtime client connection | Supported for transcription | The SDK constructs /v1/realtime from the KittenML base URL; KittenML implements a transcription subset, not every OpenAI Realtime feature |
| WebRTC realtime transcription | Supported | Uses short-lived client tokens, SDP signaling, microphone media, and an oai-events data channel |
| Streaming TTS audio | Supported | stream_format: audio or stream_format: sse |
| Incremental TTS text input | KittenML extension | Use WSS /v1/tts/realtime; OpenAI's speech endpoint accepts complete input text |
| Word-level timestamps, diarization, confidence scores | Not supported | Segment timestamps are available in verbose STT output |
Compatibility covers the workflows above, not every field, model, endpoint, or behavior available from another provider.
For upload STT, the fields in the matrix are the accepted public contract.
OpenAI options such as prompt, temperature, and word-level timestamp
granularities do not affect KittenML inference and should be omitted. An
unsupported response_format, empty file, or oversized file is rejected.
For TTS, send only the documented fields. Unsupported model IDs are rejected, and invalid input, voice, format, or speed values return field validation errors.
The consent-bound saved-voice resource paths, multipart field names, and
response objects follow OpenAI's custom-voice HTTP contract. Making consent
optional, deleting a single voice by ID, and the one-request reference_audio
form are additive KittenML conveniences. Current OpenAI SDK speech helpers
accept voice={"id":"voice_..."} directly; use a normal HTTP multipart client
for an inline file and ordinary HTTP for lifecycle operations the installed SDK
does not expose as a high-level method.
The OpenAI-compatible speech endpoint receives one complete JSON request. A
complete-response client can buffer the response before using it; a streaming
SDK helper lets the application consume the same binary response progressively.
KittenML additionally provides a bidirectional WebSocket when an LLM or
another producer supplies text over time. That WebSocket accepts built-in
voices and saved owner-scoped voice_... IDs created with or without consent.
It has its own event protocol and is not exposed by OpenAI TTS SDK methods.
Inline reference audio remains an HTTP request form.
KittenML extensions
JSON and realtime transcript results can add:
enriched_textaccenttagsclean_textsegmentsrequest_id
Realtime sessions also accept the KittenML session.close client event and
return session.closed after pending audio has been flushed and the connection
is shutting down. These are optional lifecycle extensions. The
OpenAI-compatible
conversation.item.input_audio_transcription.completed event remains the
authoritative final transcript for each turn.
OpenAI SDK response objects preserve these additional JSON fields.
Realtime clients
The canonical WebSocket URL is:
wss://api.kittenml.com/v1/realtime?model=kittenasr-enhanced-preview
The official Python SDK can open this route when configured with
base_url="https://api.kittenml.com/v1". Standard WebSocket libraries also
work; see the tested Python and Node.js examples. Existing
clients using /v1/stt/realtime continue to work through the compatibility
alias.
Browser applications can instead use the documented WebRTC transcription flow. KittenML follows the familiar short-lived-token and SDP exchange pattern, but uses the KittenML API hostname and documented signaling paths.