Upload transcription
POST https://api.kittenml.com/v1/audio/transcriptions
Already have the whole recording? Drop it here. Upload one audio file as
multipart/form-data, then choose one complete response or a finite
Server-Sent Events stream.
Send a bearer API key with the asr:transcribe permission.
Request fields
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
file | file | Yes | — | Non-empty audio file, at most 25 MiB |
model | string | No | kittenasr-enhanced-preview | Hosted ASR model ID |
language | string | No | Automatic detection | Optional language hint such as en |
response_format | string | No | json | json, verbose_json, text, srt, or vtt |
stream | boolean-like string | No | false | Set to true for finite SSE output |
Supported containers include WAV, FLAC, MP3, M4A, MP4, OGG, and WebM. The service decodes the file and converts it to mono audio internally.
The plain transcript tells you what was said. JSON responses can also preserve
how it was said through enriched_text, accent, and tags.
cURL
curl -L --fail-with-body -sS \
https://api.kittenml.com/v1/audio/transcriptions \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F "file=@speech.flac" \
-F "model=kittenasr-enhanced-preview" \
-F "response_format=verbose_json"
In cURL, @ reads the file from disk. Relative paths start from your current
directory. You can also use an absolute path:
-F "file=@/absolute/path/to/speech.flac"
-L follows the continuation URL used when a transcription crosses the hosting
platform's synchronous response window. --fail-with-body makes cURL return a
nonzero exit code for HTTP errors while preserving the API error body. -sS
hides progress but still prints errors.
Python with the OpenAI SDK
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["KITTENML_API_KEY"],
base_url="https://api.kittenml.com/v1",
)
# Keep the file open until the SDK has completed the multipart upload.
with open("speech.flac", "rb") as audio:
result = client.audio.transcriptions.create(
model="kittenasr-enhanced-preview",
file=audio,
response_format="verbose_json",
)
# request_id lets support correlate this client result with server-side logs.
print(result.text)
print(result.enriched_text)
print(result.request_id)
JavaScript with the OpenAI SDK
Install the official OpenAI JavaScript SDK:
npm install openai
Then pass a Node.js file stream directly to the transcription method, following the standard OpenAI SDK upload pattern:
import fs from "node:fs";
import OpenAI from "openai";
const openai = new OpenAI({
apiKey: process.env.KITTENML_API_KEY,
baseURL: "https://api.kittenml.com/v1",
});
const transcription = await openai.audio.transcriptions.create({
file: fs.createReadStream("speech.mp3"),
model: "kittenasr-enhanced-preview",
response_format: "verbose_json",
});
console.log(transcription.text);
console.log(transcription.enriched_text);
console.log(transcription.request_id);
baseURL points the standard OpenAI client at KittenML. No custom upload
helper, manual FormData, or file buffering is required. Keep request_id
with application logs when you need support to trace a request.
Non-streaming responses
Leave stream unset or set it to false for one response after transcription
finishes.
response_format | Content type | Result |
|---|---|---|
json | application/json | Transcript and KittenML metadata |
verbose_json | application/json | JSON plus duration, language, and segments |
text | text/plain | Clean transcript only |
srt | application/x-subrip | Timestamped SubRip subtitles |
vtt | text/vtt | Timestamped WebVTT subtitles |
JSON
{
"text": "A cold, lucid indifference reigned in his soul.",
"enriched_text": "[low][slow]A (cold), [pause_short] (lucid) (indifference) (reigned) in his (soul).[/slow][/low]",
"accent": "General American",
"tags": [
{"type": "tag", "label": "low", "start": 0, "end": 5},
{"type": "tag", "label": "slow", "start": 5, "end": 11},
{"type": "tag", "label": "pause_short", "start": 21, "end": 34},
{"type": "tag", "label": "slow", "start": 82, "end": 89},
{"type": "tag", "label": "low", "start": 89, "end": 95},
{"type": "stress", "label": "stress", "start": 13, "end": 19},
{"type": "stress", "label": "stress", "start": 35, "end": 42},
{"type": "stress", "label": "stress", "start": 43, "end": 57},
{"type": "stress", "label": "stress", "start": 58, "end": 67},
{"type": "stress", "label": "stress", "start": 75, "end": 81}
],
"request_id": "<request-id>"
}
Verbose JSON
verbose_json adds task, language, duration, and segments:
{
"task": "transcribe",
"language": "English",
"duration": 4.275,
"text": "A cold, lucid indifference reigned in his soul.",
"enriched_text": "[low][slow]A (cold), [pause_short] (lucid) (indifference) (reigned) in his (soul).[/slow][/low]",
"accent": "General American",
"segments": [
{
"id": 0,
"start": 0.56,
"end": 4.192,
"text": "A cold, lucid indifference reigned in his soul.",
"enriched_text": "[low][slow]A (cold), [pause_short] (lucid) (indifference) (reigned) in his (soul).[/slow][/low]"
}
],
"request_id": "<request-id>"
}
The complete response also contains the same tags array shown in the JSON
example.
Text
A cold, lucid indifference reigned in his soul.
SRT
1
00:00:00,560 --> 00:00:04,192
A cold, lucid indifference reigned in his soul.
VTT
WEBVTT
00:00:00.560 --> 00:00:04.192
A cold, lucid indifference reigned in his soul.
Model wording, enrichment labels, accent, and segment boundaries can vary.
Stream an uploaded file with SSE
Set stream=true and use cURL's -N option to display events immediately:
curl -L --fail-with-body -sS -N \
https://api.kittenml.com/v1/audio/transcriptions \
-H "Authorization: Bearer $KITTENML_API_KEY" \
-F "file=@speech.flac" \
-F "model=kittenasr-enhanced-preview" \
-F "stream=true"
The response uses text/event-stream:
data: {"type":"transcript.text.delta","delta":"A cold, lucid","clean_text":"A cold, lucid","enriched_text":"[low]A (cold), lucid[/low]","request_id":"<request-id>"}
data: {"type":"transcript.text.done","text":"A cold, lucid indifference reigned in his soul.","enriched_text":"[low][slow]A (cold), lucid indifference reigned in his soul.[/slow][/low]","request_id":"<request-id>"}
data: [DONE]
There may be zero or more delta events. The terminal transcript.text.done
event is authoritative and contains accent, tags, and segments. The
server closes the finite response after [DONE]. The abridged events above
omit those final metadata fields.
If transcription fails after streaming begins, the stream can emit a
structured error event before [DONE].
With stream=true, every accepted response_format value produces this same
SSE protocol. Use response_format to select an output representation only
when stream=false.
See Response formats and SSE for a longer verified example with four timestamped segments and multiple delta events.
Response fields
| Field | Meaning |
|---|---|
text | Clean transcript |
enriched_text | Transcript with delivery, pause, and emphasis markup |
accent | Best-effort accent classification; may be empty |
tags | Half-open character ranges into enriched_text |
duration | Decoded input duration in seconds (verbose_json) |
segments | Timestamped speech segments; upload start and end are seconds (verbose_json and final SSE event) |
request_id | Identifier for logs and support |
Upload STT shares the organization's ASR concurrency limit with realtime sessions.
Runnable Python and JavaScript clients are available in the upload STT examples.