Clone a voice
KittenTTS 0.8 models have no voice cloning. See feature support by model.
Kitten TTS 2 supports three custom-voice workflows:
| Workflow | Consent resource | Saved voice ID | Best for |
|---|---|---|---|
| One-request reference | Not required | No | Trying a voice, user-provided audio that should apply to one request, or provider-style instant cloning |
| Saved voice, KittenML extension | Optional | Yes, voice_... | Reusing a voice with a direct delete-by-voice-ID lifecycle |
| Saved voice, OpenAI-compatible | Required | Yes, voice_... | Preserving OpenAI's consent-bound multipart request shape |
All three workflows use POST /v1/audio/speech, support complete binary
responses and SSE output, and are available only with Kitten TTS 2,
kitten-tts-2-latest. Saved voices—with or without a consent binding—also
work on the KittenML input-streaming WebSocket. One-request inline reference
audio remains an HTTP request form. Custom voices currently synthesize English.
:::warning Permission Only upload your own voice or a voice for which you have explicit permission. Keep API keys, recordings, consent IDs, and voice IDs out of client-side logs and analytics. :::
Use reference audio once
Send reference_audio as a multipart file alongside the text to synthesize.
This form requires no consent resource and creates no saved voice_...
object:
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'model=kitten-tts-2-latest' \
-F 'input=Speak this sentence using the uploaded reference voice.' \
-F 'reference_audio=@reference.wav;type=audio/wav' \
-F 'response_format=wav' \
--output cloned-speech.wav
Each cURL -F option automatically selects multipart/form-data; do not add a
JSON Content-Type header. The reference_audio=@reference.wav field is what
selects one-request cloning: @ tells cURL to read and upload the local file.
A multipart speech request without reference_audio returns 400. Send an
ordinary JSON request when you want a default built-in voice.
The multipart form accepts model, input, reference_audio,
reference_transcript, response_format, stream_format, speed,
instructions, and the decoding controls
mode, temperature, top_p, top_k, min_p, and max_new_tokens. Any
other field is rejected with 400, and model must be kitten-tts-2-latest.
reference.wav should contain one clear speaker with little background noise.
The maximum upload size is 10 MiB. Accepted formats are WAV, MP3, OGG, AAC,
FLAC, WebM, and MP4.
The hosted API transcribes the reference internally when the transcript is omitted. If you already know the exact words spoken in the recording, send them to skip that transcription:
-F 'reference_transcript=These are the exact words spoken in reference.wav.'
A one-request clone never appears in GET /v1/voices.
JSON form
JSON clients can Base64-encode the file and use KittenML's additive voice
object. With cURL, that object can be sent through --data only after the
audio is Base64-encoded. @reference.wav is cURL multipart syntax and is not
valid JSON; inside JSON it would only be a literal string. Base64 increases
the request size by about one third, so multipart is recommended for local
files on the hosted API.
import base64
import os
from pathlib import Path
import httpx
reference = base64.b64encode(Path("reference.wav").read_bytes()).decode("ascii")
response = httpx.post(
"https://api.kittenml.com/v1/audio/speech",
headers={"Authorization": f"Bearer {os.environ['KITTENML_API_KEY']}"},
json={
"model": "kitten-tts-2-latest",
"input": "Speak this sentence using the uploaded reference voice.",
"voice": {
"reference_audio": {"data": reference, "format": "wav"},
# Optional; omit this field to let KittenML transcribe the clip.
"reference_transcript": "These are the words in reference.wav.",
},
"response_format": "wav",
},
timeout=180,
)
response.raise_for_status()
Path("cloned-speech.wav").write_bytes(response.content)
The voice object also accepts an optional consent ID; a one-request clone
does not need one.
For progressive output, add stream_format=sse to multipart data, or
"stream_format": "sse" to JSON. Decode each Base64
speech.audio.delta in order and require speech.audio.done before treating
the stream as complete. Raw binary output remains available with
stream_format=audio.
Save a reusable voice without consent
KittenML's additive lifecycle makes consent optional on the existing voice
creation route. Omitting it does not change the response shape:
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/voices \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'name=My saved voice' \
-F 'audio_sample=@reference.wav;type=audio/wav'
{
"id": "voice_0123456789abcdef0123456789abcdef",
"object": "audio.voice",
"name": "My saved voice",
"created_at": 1787760000
}
Generate speech with the
returned ID exactly as shown in
Generate speech with the voice ID, and
delete it with DELETE /v1/audio/voices/{voice_id},
which removes that one voice and no other voice or consent.
When a Kitten TTS 2 update changes how cloned voices are prepared, each saved
voice is prepared again on its first request after the update. That request
can take up to about 25 seconds longer than usual. If preparation takes
longer, the API returns 503 with a Retry-After header; retry after that
many seconds and the prepared voice is used. Later requests with the same voice
are not affected, and the voice ID does not change.
Stream text with the consent-free saved voice
Use the returned ID as the voice query parameter when text itself arrives
incrementally. This is a KittenML WebSocket extension rather than an OpenAI
speech endpoint:
wss://api.kittenml.com/v1/tts/realtime?model=kitten-tts-2-latest&voice=voice_0123456789abcdef0123456789abcdef&speed=1&response_format=pcm
Send the API key in the WebSocket upgrade header:
Authorization: Bearer sk_kitten_live_...
Then send input_text.append, optionally input_text.commit, and finally
input_text.done while reading speech.audio.delta events concurrently. The
deltas contain Base64-encoded, headerless 24 kHz mono PCM16. Require the final
speech.audio.done event before treating the output as complete. See
Stream text input for complete Python and Node.js
clients.
response_format on this WebSocket must be pcm; any other value is
rejected with 400 unsupported_audio_format. A custom voice ID here also
requires a long-lived organization or registered legacy key — an ephemeral
ek_... credential is refused with 403 durable_api_key_required.
The voice remains owner-scoped. The API key used by the WebSocket must resolve to the same organization or registered legacy owner that created it.
Delete the consent-free saved voice
curl -L --fail-with-body -sS -X DELETE \
https://api.kittenml.com/v1/audio/voices/voice_0123456789abcdef0123456789abcdef \
-H "Authorization: Bearer $KITTENML_API_KEY"
{
"id": "voice_0123456789abcdef0123456789abcdef",
"deleted": true,
"object": "audio.voice"
}
Deletion revokes the voice before its stored artifacts are removed. A later
speech request receives 404 voice_not_found. IDs are owner-scoped; another
organization or registered legacy key receives the same 404 and cannot learn
whether the ID exists.
OpenAI-compatible consent-bound voice
The original persistent lifecycle continues to follow OpenAI's custom-voice HTTP shape without changes:
- create a consent record with
POST /v1/audio/voice_consents; - create a voice with
POST /v1/audio/voices; - pass
{"id":"voice_..."}as the speech request'svoice; - delete the consent to revoke every dependent voice, or use KittenML's additive delete-by-voice-ID route to remove one voice only.
Saved voices belong to the organization—or migrated legacy key—that created
them. Another owner receives 404 rather than learning whether an ID exists.
Ephemeral ek_... credentials cannot manage persistent biometric resources.
List the current built-ins and the authenticated organization's active saved
voices with GET /v1/voices?model=kitten-tts-2-latest. See
List voices. Revoked or foreign voices are never returned.
GET /v1/audio/voices is not a listing route and answers 405 with
Allow: POST.
1. Create a consent record
The speaker records this phrase exactly:
I am the owner of this voice and I consent to its use to create a synthetic voice.
The consent clip is separate from the reference sample. Record one speaker, without music, for 2–30 seconds.
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/voice_consents \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'name=Studio consent' \
-F 'language=en-US' \
-F 'recording=@consent.wav;type=audio/wav'
The response is an audio.voice_consent object with an ID such as
cons_0123..., plus name, language, and created_at.
GET /v1/audio/voice_consents lists the consents the caller owns, newest
first, in an object: "list" envelope with data, has_more, first_id,
and last_id, and accepts limit (1–100) and after. Unlike
GET /v1/audio/voices, this route is a listing route. Use it to recover a
consent ID you did not keep.
The hosted API transcribes the recording to check the phrase. Send the words yourself to skip that step:
-F 'consent_transcript=I am the owner of this voice and I consent to its use to create a synthetic voice.'
The field is optional.
2. Create the saved voice
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/voices \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'name=My studio voice' \
-F 'consent=cons_0123456789abcdef0123456789abcdef' \
-F 'audio_sample=@reference.wav;type=audio/wav'
The response is an audio.voice object with an ID such as voice_0123....
This request keeps the OpenAI-compatible name, consent, and audio_sample
fields. reference_transcript is an accepted optional field here too, exactly
as it is for a one-request reference: supply it to skip internal
transcription of the sample.
The sample must match the consent speaker. A sample whose speaker similarity
falls below the deployment's threshold is rejected with
422 speaker_verification_failed on audio_sample rather than saved.
3. Generate speech with the voice ID
curl -L --fail-with-body -sS -N \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H 'Content-Type: application/json' \
--data '{
"model": "kitten-tts-2-latest",
"input": "This sentence uses my saved custom voice.",
"voice": {"id": "voice_0123456789abcdef0123456789abcdef"},
"response_format": "pcm",
"stream_format": "sse"
}'
Each SSE frame is a data: line whose JSON carries a type. Expect a run of
speech.audio.delta frames and one closing speech.audio.done; there is no
data: [DONE] sentinel on this route.
The current OpenAI Python SDK accepts the custom voice object directly:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["KITTENML_API_KEY"],
base_url="https://api.kittenml.com/v1",
)
audio = client.audio.speech.create(
model="kitten-tts-2-latest",
input="This sentence uses my saved custom voice.",
voice={"id": "voice_0123456789abcdef0123456789abcdef"},
response_format="wav",
)
audio.write_to_file("custom-voice.wav")
Existing OpenAI SDK users therefore change the API key, base URL, and model name; the speech method and saved voice object stay the same. The current SDK does not expose high-level helpers for every multipart voice-management operation, so create/update/delete lifecycle examples use ordinary HTTP.
Saved custom voices support buffered HTTP audio, progressive HTTP SSE, and
bidirectional WSS /v1/tts/realtime text-input/audio-output streaming. The
same WebSocket contract applies whether the saved voice was created with or
without consent. Realtime availability does not guarantee gap-free playback:
phrase boundaries, load, and replica placement can still produce gaps, so
clients should buffer PCM adaptively. See
Stream text input.
4. Revoke the consent and all dependent voices
curl -L --fail-with-body -sS -X DELETE \
https://api.kittenml.com/v1/audio/voice_consents/cons_0123456789abcdef0123456789abcdef \
-H "Authorization: Bearer $KITTENML_API_KEY"
{
"id": "cons_0123456789abcdef0123456789abcdef",
"deleted": true,
"object": "audio.voice_consent"
}
Subsequent synthesis with a dependent voice returns 404 voice_not_found.
To delete one consent-bound voice while retaining its consent and sibling
voices, use DELETE /v1/audio/voices/{voice_id} instead.
Timing and many clones at once
Saving a voice usually takes a few seconds, and a one-request reference
usually starts speaking within one to two seconds. Several clone requests can
run at the same time: one-request references, saved voices, and consents.
When many arrive together, a request waits in a short queue for a free slot
instead of failing. If the queue is full, or a request has waited about two
minutes, it returns 429 with "code": "custom_voice_creation_busy"; retry
after a short wait, with backoff. A one-request reference is also a speech
request, so your organization's
TTS concurrency limit applies to it.
The sections above cover each hosted workflow from upload to deletion. To see which saved voices your organization has, use List voices or the runnable voice-listing TTS examples.