Skip to main content

Generation and streaming

Generate one result​

let result = try await tts.generate(
"Small models can still have a big voice.",
voice: .luna,
speed: 1.05
)

print(result.samples.count)
print(result.sampleRate) // 24000
print(result.duration)
print(result.effectiveSpeed)

Blank input throws KittenTTSError.emptyInput. Speed is clamped to 0.5–2.0 and then combined with the selected model's voice speed prior.

Stream long input​

for try await chunk in tts.generateStreaming(
articleText,
voice: .jasper,
speed: 1.0
) {
consume(chunk.samples, sampleRate: chunk.sampleRate)
}

The SDK splits input into sentences and yields one complete KittenTTSResult after each local inference. The stream finishes after every sentence or throws the first synthesis error.

Actor behavior​

KittenTTS is an actor. Await its isolated methods and reuse the same instance instead of constructing a model for every sentence. Inference work is detached with user-initiated priority so UI actors are not occupied by the ONNX call.

Long-form playback​

For gap-controlled narration, feed streamed PCM into an audio engine or queue. Calling speak() separately for every sentence is simpler but waits for each full generation and playback cycle.