Skip to main content

Clone a voice

Kitten TTS 2 only

KittenTTS 0.8 models have no voice cloning. See feature support by model.

Kitten TTS 2 supports three custom-voice workflows:

WorkflowConsent resourceSaved voice IDBest for
One-request referenceNot requiredNoTrying a voice, user-provided audio that should apply to one request, or provider-style instant cloning
Saved voice, KittenML extensionOptionalYes, voice_...Reusing a voice with a direct delete-by-voice-ID lifecycle
Saved voice, OpenAI-compatibleRequiredYes, voice_...Preserving OpenAI's consent-bound multipart request shape

All three workflows use POST /v1/audio/speech, support complete binary responses and SSE output, and are available only with Kitten TTS 2, kitten-tts-2-latest. Saved voices—with or without a consent binding—also work on the KittenML input-streaming WebSocket. One-request inline reference audio remains an HTTP request form. Custom voices currently synthesize English.

:::warning Permission Only upload your own voice or a voice for which you have explicit permission. Keep API keys, recordings, consent IDs, and voice IDs out of client-side logs and analytics. :::

Use reference audio once​

Send reference_audio as a multipart file alongside the text to synthesize. This form requires no consent resource and creates no saved voice_... object:

curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'model=kitten-tts-2-latest' \
-F 'input=Speak this sentence using the uploaded reference voice.' \
-F 'reference_audio=@reference.wav;type=audio/wav' \
-F 'response_format=wav' \
--output cloned-speech.wav

Each cURL -F option automatically selects multipart/form-data; do not add a JSON Content-Type header. The reference_audio=@reference.wav field is what selects one-request cloning: @ tells cURL to read and upload the local file. A multipart speech request without reference_audio returns 400. Send an ordinary JSON request when you want a default built-in voice.

The multipart form accepts model, input, reference_audio, reference_transcript, response_format, stream_format, speed, instructions, and the decoding controls mode, temperature, top_p, top_k, min_p, and max_new_tokens. Any other field is rejected with 400, and model must be kitten-tts-2-latest.

reference.wav should contain one clear speaker with little background noise. The maximum upload size is 10 MiB. Accepted formats are WAV, MP3, OGG, AAC, FLAC, WebM, and MP4.

The hosted API transcribes the reference internally when the transcript is omitted. If you already know the exact words spoken in the recording, send them to skip that transcription:

-F 'reference_transcript=These are the exact words spoken in reference.wav.'

A one-request clone never appears in GET /v1/voices.

JSON form​

JSON clients can Base64-encode the file and use KittenML's additive voice object. With cURL, that object can be sent through --data only after the audio is Base64-encoded. @reference.wav is cURL multipart syntax and is not valid JSON; inside JSON it would only be a literal string. Base64 increases the request size by about one third, so multipart is recommended for local files on the hosted API.

import base64
import os
from pathlib import Path

import httpx

reference = base64.b64encode(Path("reference.wav").read_bytes()).decode("ascii")
response = httpx.post(
"https://api.kittenml.com/v1/audio/speech",
headers={"Authorization": f"Bearer {os.environ['KITTENML_API_KEY']}"},
json={
"model": "kitten-tts-2-latest",
"input": "Speak this sentence using the uploaded reference voice.",
"voice": {
"reference_audio": {"data": reference, "format": "wav"},
# Optional; omit this field to let KittenML transcribe the clip.
"reference_transcript": "These are the words in reference.wav.",
},
"response_format": "wav",
},
timeout=180,
)
response.raise_for_status()
Path("cloned-speech.wav").write_bytes(response.content)

The voice object also accepts an optional consent ID; a one-request clone does not need one.

For progressive output, add stream_format=sse to multipart data, or "stream_format": "sse" to JSON. Decode each Base64 speech.audio.delta in order and require speech.audio.done before treating the stream as complete. Raw binary output remains available with stream_format=audio.

KittenML's additive lifecycle makes consent optional on the existing voice creation route. Omitting it does not change the response shape:

curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/voices \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'name=My saved voice' \
-F 'audio_sample=@reference.wav;type=audio/wav'
{
"id": "voice_0123456789abcdef0123456789abcdef",
"object": "audio.voice",
"name": "My saved voice",
"created_at": 1787760000
}

Generate speech with the returned ID exactly as shown in Generate speech with the voice ID, and delete it with DELETE /v1/audio/voices/{voice_id}, which removes that one voice and no other voice or consent.

First use after a model update

When a Kitten TTS 2 update changes how cloned voices are prepared, each saved voice is prepared again on its first request after the update. That request can take up to about 25 seconds longer than usual. If preparation takes longer, the API returns 503 with a Retry-After header; retry after that many seconds and the prepared voice is used. Later requests with the same voice are not affected, and the voice ID does not change.

Use the returned ID as the voice query parameter when text itself arrives incrementally. This is a KittenML WebSocket extension rather than an OpenAI speech endpoint:

wss://api.kittenml.com/v1/tts/realtime?model=kitten-tts-2-latest&voice=voice_0123456789abcdef0123456789abcdef&speed=1&response_format=pcm

Send the API key in the WebSocket upgrade header:

Authorization: Bearer sk_kitten_live_...

Then send input_text.append, optionally input_text.commit, and finally input_text.done while reading speech.audio.delta events concurrently. The deltas contain Base64-encoded, headerless 24 kHz mono PCM16. Require the final speech.audio.done event before treating the output as complete. See Stream text input for complete Python and Node.js clients.

response_format on this WebSocket must be pcm; any other value is rejected with 400 unsupported_audio_format. A custom voice ID here also requires a long-lived organization or registered legacy key — an ephemeral ek_... credential is refused with 403 durable_api_key_required.

The voice remains owner-scoped. The API key used by the WebSocket must resolve to the same organization or registered legacy owner that created it.

curl -L --fail-with-body -sS -X DELETE \
https://api.kittenml.com/v1/audio/voices/voice_0123456789abcdef0123456789abcdef \
-H "Authorization: Bearer $KITTENML_API_KEY"
{
"id": "voice_0123456789abcdef0123456789abcdef",
"deleted": true,
"object": "audio.voice"
}

Deletion revokes the voice before its stored artifacts are removed. A later speech request receives 404 voice_not_found. IDs are owner-scoped; another organization or registered legacy key receives the same 404 and cannot learn whether the ID exists.

The original persistent lifecycle continues to follow OpenAI's custom-voice HTTP shape without changes:

  1. create a consent record with POST /v1/audio/voice_consents;
  2. create a voice with POST /v1/audio/voices;
  3. pass {"id":"voice_..."} as the speech request's voice;
  4. delete the consent to revoke every dependent voice, or use KittenML's additive delete-by-voice-ID route to remove one voice only.

Saved voices belong to the organization—or migrated legacy key—that created them. Another owner receives 404 rather than learning whether an ID exists. Ephemeral ek_... credentials cannot manage persistent biometric resources.

List the current built-ins and the authenticated organization's active saved voices with GET /v1/voices?model=kitten-tts-2-latest. See List voices. Revoked or foreign voices are never returned. GET /v1/audio/voices is not a listing route and answers 405 with Allow: POST.

The speaker records this phrase exactly:

I am the owner of this voice and I consent to its use to create a synthetic voice.

The consent clip is separate from the reference sample. Record one speaker, without music, for 2–30 seconds.

curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/voice_consents \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'name=Studio consent' \
-F 'language=en-US' \
-F 'recording=@consent.wav;type=audio/wav'

The response is an audio.voice_consent object with an ID such as cons_0123..., plus name, language, and created_at.

GET /v1/audio/voice_consents lists the consents the caller owns, newest first, in an object: "list" envelope with data, has_more, first_id, and last_id, and accepts limit (1–100) and after. Unlike GET /v1/audio/voices, this route is a listing route. Use it to recover a consent ID you did not keep.

The hosted API transcribes the recording to check the phrase. Send the words yourself to skip that step:

-F 'consent_transcript=I am the owner of this voice and I consent to its use to create a synthetic voice.'

The field is optional.

2. Create the saved voice​

curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/voices \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F 'name=My studio voice' \
-F 'consent=cons_0123456789abcdef0123456789abcdef' \
-F 'audio_sample=@reference.wav;type=audio/wav'

The response is an audio.voice object with an ID such as voice_0123.... This request keeps the OpenAI-compatible name, consent, and audio_sample fields. reference_transcript is an accepted optional field here too, exactly as it is for a one-request reference: supply it to skip internal transcription of the sample.

The sample must match the consent speaker. A sample whose speaker similarity falls below the deployment's threshold is rejected with 422 speaker_verification_failed on audio_sample rather than saved.

3. Generate speech with the voice ID​

curl -L --fail-with-body -sS -N \
https://api.kittenml.com/v1/audio/speech \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-H 'Content-Type: application/json' \
--data '{
"model": "kitten-tts-2-latest",
"input": "This sentence uses my saved custom voice.",
"voice": {"id": "voice_0123456789abcdef0123456789abcdef"},
"response_format": "pcm",
"stream_format": "sse"
}'

Each SSE frame is a data: line whose JSON carries a type. Expect a run of speech.audio.delta frames and one closing speech.audio.done; there is no data: [DONE] sentinel on this route.

The current OpenAI Python SDK accepts the custom voice object directly:

import os
from openai import OpenAI

client = OpenAI(
api_key=os.environ["KITTENML_API_KEY"],
base_url="https://api.kittenml.com/v1",
)
audio = client.audio.speech.create(
model="kitten-tts-2-latest",
input="This sentence uses my saved custom voice.",
voice={"id": "voice_0123456789abcdef0123456789abcdef"},
response_format="wav",
)
audio.write_to_file("custom-voice.wav")

Existing OpenAI SDK users therefore change the API key, base URL, and model name; the speech method and saved voice object stay the same. The current SDK does not expose high-level helpers for every multipart voice-management operation, so create/update/delete lifecycle examples use ordinary HTTP.

Saved custom voices support buffered HTTP audio, progressive HTTP SSE, and bidirectional WSS /v1/tts/realtime text-input/audio-output streaming. The same WebSocket contract applies whether the saved voice was created with or without consent. Realtime availability does not guarantee gap-free playback: phrase boundaries, load, and replica placement can still produce gaps, so clients should buffer PCM adaptively. See Stream text input.

curl -L --fail-with-body -sS -X DELETE \
https://api.kittenml.com/v1/audio/voice_consents/cons_0123456789abcdef0123456789abcdef \
-H "Authorization: Bearer $KITTENML_API_KEY"
{
"id": "cons_0123456789abcdef0123456789abcdef",
"deleted": true,
"object": "audio.voice_consent"
}

Subsequent synthesis with a dependent voice returns 404 voice_not_found.

To delete one consent-bound voice while retaining its consent and sibling voices, use DELETE /v1/audio/voices/{voice_id} instead.

Timing and many clones at once​

Saving a voice usually takes a few seconds, and a one-request reference usually starts speaking within one to two seconds. Several clone requests can run at the same time: one-request references, saved voices, and consents. When many arrive together, a request waits in a short queue for a free slot instead of failing. If the queue is full, or a request has waited about two minutes, it returns 429 with "code": "custom_voice_creation_busy"; retry after a short wait, with backoff. A one-request reference is also a speech request, so your organization's TTS concurrency limit applies to it.

The sections above cover each hosted workflow from upload to deletion. To see which saved voices your organization has, use List voices or the runnable voice-listing TTS examples.