Skip to main content

Upload transcription

POST https://api.kittenml.com/v1/audio/transcriptions

Already have the whole recording? Drop it here. Upload one audio file as multipart/form-data, then choose one complete response or a finite Server-Sent Events stream.

Send a bearer API key with the asr:transcribe permission.

Request fields​

FieldTypeRequiredDefaultDescription
filefileYes—Non-empty audio file, at most 25 MiB
modelstringNokittenasr-enhanced-previewHosted ASR model ID
languagestringNoAutomatic detectionOptional language hint such as en
response_formatstringNojsonjson, verbose_json, text, srt, or vtt
streamboolean-like stringNofalseSet to true for finite SSE output

Supported containers include WAV, FLAC, MP3, M4A, MP4, OGG, and WebM. The service decodes the file and converts it to mono audio internally.

The plain transcript tells you what was said. JSON responses can also preserve how it was said through enriched_text, accent, and tags.

cURL​

curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/transcriptions \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F "file=@speech.flac" \
-F "model=kittenasr-enhanced-preview" \
-F "response_format=verbose_json"

In cURL, @ reads the file from disk. Relative paths start from your current directory. You can also use an absolute path:

-F "file=@/absolute/path/to/speech.flac"

-L follows the continuation URL used when a transcription crosses the hosting platform's synchronous response window. --fail-with-body makes cURL return a nonzero exit code for HTTP errors while preserving the API error body. -sS hides progress but still prints errors.

Python with the OpenAI SDK​

import os
from openai import OpenAI

client = OpenAI(
api_key=os.environ["KITTENML_API_KEY"],
base_url="https://api.kittenml.com/v1",
)

# Keep the file open until the SDK has completed the multipart upload.
with open("speech.flac", "rb") as audio:
result = client.audio.transcriptions.create(
model="kittenasr-enhanced-preview",
file=audio,
response_format="verbose_json",
)

# request_id lets support correlate this client result with server-side logs.
print(result.text)
print(result.enriched_text)
print(result.request_id)

JavaScript with the OpenAI SDK​

Install the official OpenAI JavaScript SDK:

npm install openai

Then pass a Node.js file stream directly to the transcription method, following the standard OpenAI SDK upload pattern:

import fs from "node:fs";
import OpenAI from "openai";

const openai = new OpenAI({
apiKey: process.env.KITTENML_API_KEY,
baseURL: "https://api.kittenml.com/v1",
});

const transcription = await openai.audio.transcriptions.create({
file: fs.createReadStream("speech.mp3"),
model: "kittenasr-enhanced-preview",
response_format: "verbose_json",
});

console.log(transcription.text);
console.log(transcription.enriched_text);
console.log(transcription.request_id);

baseURL points the standard OpenAI client at KittenML. No custom upload helper, manual FormData, or file buffering is required. Keep request_id with application logs when you need support to trace a request.

Non-streaming responses​

Leave stream unset or set it to false for one response after transcription finishes.

response_formatContent typeResult
jsonapplication/jsonTranscript and KittenML metadata
verbose_jsonapplication/jsonJSON plus duration, language, and segments
texttext/plainClean transcript only
srtapplication/x-subripTimestamped SubRip subtitles
vtttext/vttTimestamped WebVTT subtitles

JSON​

{
"text": "A cold, lucid indifference reigned in his soul.",
"enriched_text": "[low][slow]A (cold), [pause_short] (lucid) (indifference) (reigned) in his (soul).[/slow][/low]",
"accent": "General American",
"tags": [
{"type": "tag", "label": "low", "start": 0, "end": 5},
{"type": "tag", "label": "slow", "start": 5, "end": 11},
{"type": "tag", "label": "pause_short", "start": 21, "end": 34},
{"type": "tag", "label": "slow", "start": 82, "end": 89},
{"type": "tag", "label": "low", "start": 89, "end": 95},
{"type": "stress", "label": "stress", "start": 13, "end": 19},
{"type": "stress", "label": "stress", "start": 35, "end": 42},
{"type": "stress", "label": "stress", "start": 43, "end": 57},
{"type": "stress", "label": "stress", "start": 58, "end": 67},
{"type": "stress", "label": "stress", "start": 75, "end": 81}
],
"request_id": "<request-id>"
}

Verbose JSON​

verbose_json adds task, language, duration, and segments:

{
"task": "transcribe",
"language": "English",
"duration": 4.275,
"text": "A cold, lucid indifference reigned in his soul.",
"enriched_text": "[low][slow]A (cold), [pause_short] (lucid) (indifference) (reigned) in his (soul).[/slow][/low]",
"accent": "General American",
"segments": [
{
"id": 0,
"start": 0.56,
"end": 4.192,
"text": "A cold, lucid indifference reigned in his soul.",
"enriched_text": "[low][slow]A (cold), [pause_short] (lucid) (indifference) (reigned) in his (soul).[/slow][/low]"
}
],
"request_id": "<request-id>"
}

The complete response also contains the same tags array shown in the JSON example.

Text​

A cold, lucid indifference reigned in his soul.

SRT​

1
00:00:00,560 --> 00:00:04,192
A cold, lucid indifference reigned in his soul.

VTT​

WEBVTT

00:00:00.560 --> 00:00:04.192
A cold, lucid indifference reigned in his soul.

Model wording, enrichment labels, accent, and segment boundaries can vary.

Stream an uploaded file with SSE​

Set stream=true and use cURL's -N option to display events immediately:

curl -L --fail-with-body -sS -N \
https://api.kittenml.com/v1/audio/transcriptions \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F "file=@speech.flac" \
-F "model=kittenasr-enhanced-preview" \
-F "stream=true"

The response uses text/event-stream:

data: {"type":"transcript.text.delta","delta":"A cold, lucid","clean_text":"A cold, lucid","enriched_text":"[low]A (cold), lucid[/low]","request_id":"<request-id>"}

data: {"type":"transcript.text.done","text":"A cold, lucid indifference reigned in his soul.","enriched_text":"[low][slow]A (cold), lucid indifference reigned in his soul.[/slow][/low]","request_id":"<request-id>"}

data: [DONE]

There may be zero or more delta events. The terminal transcript.text.done event is authoritative and contains accent, tags, and segments. The server closes the finite response after [DONE]. The abridged events above omit those final metadata fields.

If transcription fails after streaming begins, the stream can emit a structured error event before [DONE].

With stream=true, every accepted response_format value produces this same SSE protocol. Use response_format to select an output representation only when stream=false.

See Response formats and SSE for a longer verified example with four timestamped segments and multiple delta events.

Response fields​

FieldMeaning
textClean transcript
enriched_textTranscript with delivery, pause, and emphasis markup
accentBest-effort accent classification; may be empty
tagsHalf-open character ranges into enriched_text
durationDecoded input duration in seconds (verbose_json)
segmentsTimestamped speech segments; upload start and end are seconds (verbose_json and final SSE event)
request_idIdentifier for logs and support

Upload STT shares the organization's ASR concurrency limit with realtime sessions.

Runnable Python and JavaScript clients are available in the upload STT examples.