Browser WebRTC transcription
Use WebRTC when a browser needs to transcribe microphone audio directly. The
browser sends its audio track over WebRTC and receives partial and final
transcript events on the oai-events data channel.
Your permanent KittenML API key must stay on your backend. The browser receives
only a short-lived ek_... token immediately before it starts a call.
Run the complete browser example
The fastest way to start is the complete WebRTC browser example. It includes a small Node.js token server and a browser interface with Start and Stop controls, connection status, and live enriched and clean transcripts.
git clone https://github.com/KittenML/kittenml-api-examples.git
cd kittenml-api-examples/stt/realtime/webrtc
npm install
cp .env.example .env
Add an ASR-enabled key to .env:
KITTENML_API_KEY=sk_kitten_live_...
Start the example:
npm start
Open http://localhost:3000, select Start mic, allow microphone access,
and begin speaking. Select Stop to finalize the transcript.
Never place a permanent sk_kitten_live_... key in browser JavaScript. Protect
your token endpoint with your application's authentication and rate limits.
How the connection works
- The browser asks your backend for a short-lived client token.
- Your backend calls
POST /v1/realtime/client_secretswith its permanent KittenML API key. - The browser adds its microphone track and an
oai-eventsdata channel to anRTCPeerConnection. - The browser sends its SDP offer to
POST /v1/realtime/callsusing the short-lived token. - The browser applies the SDP answer. Microphone audio and transcript events then travel over the negotiated WebRTC connection.
- When recording stops, the browser sends
session.closeand waits for the final transcript before closing the peer connection.
The token and SDP requests use HTTPS. After signaling, audio and data-channel traffic use WebRTC over the negotiated direct ICE or TURN path.
Backend: create a short-lived token
The runnable example uses the OpenAI JavaScript SDK with KittenML's base URL.
The following is the relevant part of
server.mjs:
import OpenAI from "openai";
const MODEL = "kittenasr-enhanced-preview";
// This client runs only on your server. The permanent key is read from the
// environment and is never returned to the browser.
const client = new OpenAI({
apiKey: process.env.KITTENML_API_KEY,
baseURL: "https://api.kittenml.com/v1",
});
app.post("/api/realtime-token", async (request, response) => {
// In production, authenticate `request` as one of your users before
// creating a token.
const token = await client.realtime.clientSecrets.create({
expires_after: {anchor: "created_at", seconds: 300},
session: {
type: "transcription",
model: MODEL,
audio: {
input: {
format: {type: "audio/pcm", rate: 24_000},
transcription: {model: MODEL},
turn_detection: null,
},
},
},
});
// The response contains a short-lived `value` beginning with `ek_...`.
response.json(token);
});
The token's expires_at value is the Unix time after which it cannot open a
new connection. It is not the maximum duration of an already established call.
Mint a new token for every Start or reconnect action, and do not log or persist
the token.
Browser: connect the microphone
The browser requests a temporary token, creates the peer connection, and
exchanges SDP with KittenML. This excerpt follows the runnable example's
public/index.html:
// Ask your own backend for a temporary credential. The permanent API key
// never enters this browser process.
const tokenResponse = await fetch("/api/realtime-token", {method: "POST"});
const token = await tokenResponse.json();
if (!tokenResponse.ok) {
throw new Error(token.error?.message || "Could not create a realtime token");
}
// Ask for microphone access and add the resulting audio track to WebRTC.
const microphone = await navigator.mediaDevices.getUserMedia({audio: true});
const peer = new RTCPeerConnection({
iceServers: [{urls: "stun:stun.cloudflare.com:3478"}],
});
microphone.getTracks().forEach((track) => peer.addTrack(track, microphone));
// Transcript events arrive on this data channel.
const events = peer.createDataChannel("oai-events");
let completedReceived = false;
let speechDetected = false;
events.addEventListener("message", ({data}) => {
const event = JSON.parse(data);
if (event.type === "conversation.item.input_audio_transcription.delta") {
speechDetected = true;
// Delta events contain cumulative snapshots. Replace displayed text;
// do not concatenate each value.
console.log("partial", event.clean_text || event.delta);
} else if (
event.type === "conversation.item.input_audio_transcription.completed"
) {
completedReceived = true;
const finalText = event.transcript || event.clean_text || "";
console.log(finalText.trim() ? "final" : "no speech detected", finalText);
} else if (event.type === "session.closed") {
// session.closed confirms transport shutdown, not transcript success.
if (!completedReceived && speechDetected) {
console.error("WebRTC closed before the authoritative final transcript");
} else if (!completedReceived) {
console.log("no speech detected");
}
} else if (event.type === "error") {
console.error(event.error);
}
});
// Create the browser's offer and wait for ICE candidates to be gathered.
await peer.setLocalDescription(await peer.createOffer());
if (peer.iceGatheringState !== "complete") {
await new Promise((resolve) => {
const ready = () => {
if (peer.iceGatheringState === "complete") {
peer.removeEventListener("icegatheringstatechange", ready);
resolve();
}
};
peer.addEventListener("icegatheringstatechange", ready);
});
}
// Exchange the offer for KittenML's SDP answer using the temporary token.
const answer = await fetch("https://api.kittenml.com/v1/realtime/calls", {
method: "POST",
headers: {
Authorization: `Bearer ${token.value}`,
"Content-Type": "application/sdp",
},
body: peer.localDescription.sdp,
});
if (!answer.ok) {
throw new Error(`WebRTC signaling failed: ${answer.status}`);
}
await peer.setRemoteDescription({
type: "answer",
sdp: await answer.text(),
});
Stop and finalize
Disable the microphone track, ask the server to finalize the session, and wait
for conversation.item.input_audio_transcription.completed before closing the
peer connection:
microphone.getAudioTracks().forEach((track) => {
track.enabled = false;
});
if (events.readyState === "open") {
events.send(JSON.stringify({type: "session.close"}));
}
// Close `peer` after the completed event arrives if your application needs
// the authoritative final transcript.
Treat conversation.item.input_audio_transcription.completed as the
authoritative success signal. session.closed only confirms that the WebRTC
session ended after the close request; it does not prove that transcription
completed. The session.close and session.closed lifecycle events are
KittenML extensions shared by the WebSocket and WebRTC transports. If speech
or a partial transcript was observed but the session closes before a completed
event, report the request as incomplete. If no speech, partial, or completed
event was observed, report No speech detected rather than Finished.
Events and limits
WebRTC returns the same session, partial, completed, and error events documented
in Realtime events. Replace cumulative clean_text and
enriched_text snapshots rather than concatenating them.
WebSocket and WebRTC share the same five active ASR slots per organization.
There is no fixed WebRTC call-duration limit. If SDP admission returns HTTP
429, retry with bounded exponential backoff and jitter.
For production error handling, UI state, cleanup, and a working token server, use the complete KittenML WebRTC browser example as the starting point.