Skip to main content

OpenAI compatibility

KittenML supports the common OpenAI SDK workflows for upload transcription, speech generation, and realtime transcription connections.

Set this base URL for OpenAI HTTP SDKs:

https://api.kittenml.com/v1

Do not set the SDK base URL to a complete endpoint such as /v1/audio/transcriptions; the SDK appends that path itself.

Compatibility matrix​

CapabilityStatusNotes
POST /v1/audio/transcriptionsSupportedWorks with OpenAI Python and JavaScript clients
model, file, language, response_format, streamSupportedUse the documented values
json, verbose_json, text, srt, vttSupportedIncludes KittenML metadata in JSON formats
Upload SSESupportedFinite delta/done stream ending in [DONE]
POST /v1/audio/speechSupportedComplete-response SDK patterns and progressive binary consumption are supported; SSE is optional
model, input, voice, response_format, speed, stream_formatSupportedUse KittenML model IDs and voices
stream in /v1/audio/speechKittenML extensionKitten TTS 2: false returns one complete file for quality over speed. Pass extra_body={"stream": False} in Python or stream: false in JavaScript
/v1/audio/voice_consents lifecycleSupported for Kitten TTS 2OpenAI-compatible multipart fields and audio.voice_consent objects; requires a durable organization or registered legacy key
POST /v1/audio/voices with consentSupported for Kitten TTS 2Preserves OpenAI's name, consent, and audio_sample multipart shape and audio.voice response
POST /v1/audio/voices without consentKittenML extensionOmit consent to save a reusable voice directly; the response remains an audio.voice object
DELETE /v1/audio/voices/{voice_id}KittenML extensionRevokes and deletes one owner-scoped voice without deleting a consent or sibling voices
Saved custom voice in /v1/audio/speechSupported for Kitten TTS 2Pass {"id":"voice_..."} as voice; complete binary and SSE output are supported
Inline reference audio in /v1/audio/speechKittenML extensionSend a multipart reference_audio file without creating consent or voice resources, or use a Base64 JSON voice object
Saved custom voice in WSS /v1/tts/realtimeKittenML extensionPut the owner-scoped voice_... ID in the query; both consent-bound and consent-free saved voices are supported
Realtime event names and common fieldsCompatible styleIncludes KittenML extension fields
Official OpenAI Realtime client connectionSupported for transcriptionThe SDK constructs /v1/realtime from the KittenML base URL; KittenML implements a transcription subset, not every OpenAI Realtime feature
WebRTC realtime transcriptionSupportedUses short-lived client tokens, SDP signaling, microphone media, and an oai-events data channel
Streaming TTS audioSupportedstream_format: audio or stream_format: sse
Incremental TTS text inputKittenML extensionUse WSS /v1/tts/realtime; OpenAI's speech endpoint accepts complete input text
Word-level timestamps, diarization, confidence scoresNot supportedSegment timestamps are available in verbose STT output

Compatibility covers the workflows above, not every field, model, endpoint, or behavior available from another provider.

For upload STT, the fields in the matrix are the accepted public contract. OpenAI options such as prompt, temperature, and word-level timestamp granularities do not affect KittenML inference and should be omitted. An unsupported response_format, empty file, or oversized file is rejected.

For TTS, send only the documented fields. Unsupported model IDs are rejected, and invalid input, voice, format, or speed values return field validation errors.

The consent-bound saved-voice resource paths, multipart field names, and response objects follow OpenAI's custom-voice HTTP contract. Making consent optional, deleting a single voice by ID, and the one-request reference_audio form are additive KittenML conveniences. Current OpenAI SDK speech helpers accept voice={"id":"voice_..."} directly; use a normal HTTP multipart client for an inline file and ordinary HTTP for lifecycle operations the installed SDK does not expose as a high-level method.

The OpenAI-compatible speech endpoint receives one complete JSON request. A complete-response client can buffer the response before using it; a streaming SDK helper lets the application consume the same binary response progressively. KittenML additionally provides a bidirectional WebSocket when an LLM or another producer supplies text over time. That WebSocket accepts built-in voices and saved owner-scoped voice_... IDs created with or without consent. It has its own event protocol and is not exposed by OpenAI TTS SDK methods. Inline reference audio remains an HTTP request form.

KittenML extensions​

JSON and realtime transcript results can add:

  • enriched_text
  • accent
  • tags
  • clean_text
  • segments
  • request_id

Realtime sessions also accept the KittenML session.close client event and return session.closed after pending audio has been flushed and the connection is shutting down. These are optional lifecycle extensions. The OpenAI-compatible conversation.item.input_audio_transcription.completed event remains the authoritative final transcript for each turn.

OpenAI SDK response objects preserve these additional JSON fields.

Realtime clients​

The canonical WebSocket URL is:

wss://api.kittenml.com/v1/realtime?model=kittenasr-enhanced-preview

The official Python SDK can open this route when configured with base_url="https://api.kittenml.com/v1". Standard WebSocket libraries also work; see the tested Python and Node.js examples. Existing clients using /v1/stt/realtime continue to work through the compatibility alias.

Browser applications can instead use the documented WebRTC transcription flow. KittenML follows the familiar short-lived-token and SDP exchange pattern, but uses the KittenML API hostname and documented signaling paths.