Skip to main content

Troubleshooting

Model download fails

Confirm the Hugging Face repo is reachable and the process can write to its cache. Retry with an explicit cache directory on local storage:

tts = KittenTTS(model_name, cache_dir="./.models")

Delete only the incomplete model snapshot if a download was interrupted; avoid clearing the entire shared Hugging Face cache unnecessarily.

A voice is rejected

Inspect tts.available_voices and use a display name exactly as returned, or use the raw embedding ID. Voice names are case-sensitive after alias resolution.

CUDA was requested but inference is on CPU

import onnxruntime as ort
print(ort.get_available_providers())

CUDAExecutionProvider must appear. Do not install onnxruntime and onnxruntime-gpu together in the same environment when they conflict; use a fresh virtual environment and the repository GPU requirements.

Output sounds wrong around numbers

Call normalize_text() to inspect the spoken form, then generate with clean_text=True. Add punctuation to give the chunker clear sentence boundaries.

Array shape is unexpected

Generated chunks can contain singleton dimensions. Use audio.squeeze() before passing a chunk to a playback layer that expects one-dimensional mono samples.

Audio plays at the wrong speed

Write and play at 24,000 Hz. The speed parameter controls synthesis pace; it is not a replacement for sample-rate configuration.

Need more detail

Open an issue with OS, Python version, model repo, ONNX Runtime provider list, minimal input text, and the complete traceback: KittenTTS issues.