Skip to main content

Browser WebRTC transcription

Use WebRTC when a browser needs to transcribe microphone audio directly. The browser sends its audio track over WebRTC and receives partial and final transcript events on the oai-events data channel.

Your permanent KittenML API key must stay on your backend. The browser receives only a short-lived ek_... token immediately before it starts a call.

Run the complete browser example​

The fastest way to start is the complete WebRTC browser example. It includes a small Node.js token server and a browser interface with Start and Stop controls, connection status, and live enriched and clean transcripts.

git clone https://github.com/KittenML/kittenml-api-examples.git
cd kittenml-api-examples/stt/realtime/webrtc
npm install
cp .env.example .env

Add an ASR-enabled key to .env:

KITTENML_API_KEY=sk_kitten_live_...

Start the example:

npm start

Open http://localhost:3000, select Start mic, allow microphone access, and begin speaking. Select Stop to finalize the transcript.

Keep permanent keys on the server

Never place a permanent sk_kitten_live_... key in browser JavaScript. Protect your token endpoint with your application's authentication and rate limits.

How the connection works​

  1. The browser asks your backend for a short-lived client token.
  2. Your backend calls POST /v1/realtime/client_secrets with its permanent KittenML API key.
  3. The browser adds its microphone track and an oai-events data channel to an RTCPeerConnection.
  4. The browser sends its SDP offer to POST /v1/realtime/calls using the short-lived token.
  5. The browser applies the SDP answer. Microphone audio and transcript events then travel over the negotiated WebRTC connection.
  6. When recording stops, the browser sends session.close and waits for the final transcript before closing the peer connection.

The token and SDP requests use HTTPS. After signaling, audio and data-channel traffic use WebRTC over the negotiated direct ICE or TURN path.

Backend: create a short-lived token​

The runnable example uses the OpenAI JavaScript SDK with KittenML's base URL. The following is the relevant part of server.mjs:

import OpenAI from "openai";

const MODEL = "kittenasr-enhanced-preview";

// This client runs only on your server. The permanent key is read from the
// environment and is never returned to the browser.
const client = new OpenAI({
apiKey: process.env.KITTENML_API_KEY,
baseURL: "https://api.kittenml.com/v1",
});

app.post("/api/realtime-token", async (request, response) => {
// In production, authenticate `request` as one of your users before
// creating a token.
const token = await client.realtime.clientSecrets.create({
expires_after: {anchor: "created_at", seconds: 300},
session: {
type: "transcription",
model: MODEL,
audio: {
input: {
format: {type: "audio/pcm", rate: 24_000},
transcription: {model: MODEL},
turn_detection: null,
},
},
},
});

// The response contains a short-lived `value` beginning with `ek_...`.
response.json(token);
});

The token's expires_at value is the Unix time after which it cannot open a new connection. It is not the maximum duration of an already established call. Mint a new token for every Start or reconnect action, and do not log or persist the token.

Browser: connect the microphone​

The browser requests a temporary token, creates the peer connection, and exchanges SDP with KittenML. This excerpt follows the runnable example's public/index.html:

// Ask your own backend for a temporary credential. The permanent API key
// never enters this browser process.
const tokenResponse = await fetch("/api/realtime-token", {method: "POST"});
const token = await tokenResponse.json();
if (!tokenResponse.ok) {
throw new Error(token.error?.message || "Could not create a realtime token");
}

// Ask for microphone access and add the resulting audio track to WebRTC.
const microphone = await navigator.mediaDevices.getUserMedia({audio: true});
const peer = new RTCPeerConnection({
iceServers: [{urls: "stun:stun.cloudflare.com:3478"}],
});
microphone.getTracks().forEach((track) => peer.addTrack(track, microphone));

// Transcript events arrive on this data channel.
const events = peer.createDataChannel("oai-events");
let completedReceived = false;
let speechDetected = false;

events.addEventListener("message", ({data}) => {
const event = JSON.parse(data);

if (event.type === "conversation.item.input_audio_transcription.delta") {
speechDetected = true;
// Delta events contain cumulative snapshots. Replace displayed text;
// do not concatenate each value.
console.log("partial", event.clean_text || event.delta);
} else if (
event.type === "conversation.item.input_audio_transcription.completed"
) {
completedReceived = true;
const finalText = event.transcript || event.clean_text || "";
console.log(finalText.trim() ? "final" : "no speech detected", finalText);
} else if (event.type === "session.closed") {
// session.closed confirms transport shutdown, not transcript success.
if (!completedReceived && speechDetected) {
console.error("WebRTC closed before the authoritative final transcript");
} else if (!completedReceived) {
console.log("no speech detected");
}
} else if (event.type === "error") {
console.error(event.error);
}
});

// Create the browser's offer and wait for ICE candidates to be gathered.
await peer.setLocalDescription(await peer.createOffer());
if (peer.iceGatheringState !== "complete") {
await new Promise((resolve) => {
const ready = () => {
if (peer.iceGatheringState === "complete") {
peer.removeEventListener("icegatheringstatechange", ready);
resolve();
}
};
peer.addEventListener("icegatheringstatechange", ready);
});
}

// Exchange the offer for KittenML's SDP answer using the temporary token.
const answer = await fetch("https://api.kittenml.com/v1/realtime/calls", {
method: "POST",
headers: {
Authorization: `Bearer ${token.value}`,
"Content-Type": "application/sdp",
},
body: peer.localDescription.sdp,
});
if (!answer.ok) {
throw new Error(`WebRTC signaling failed: ${answer.status}`);
}

await peer.setRemoteDescription({
type: "answer",
sdp: await answer.text(),
});

Stop and finalize​

Disable the microphone track, ask the server to finalize the session, and wait for conversation.item.input_audio_transcription.completed before closing the peer connection:

microphone.getAudioTracks().forEach((track) => {
track.enabled = false;
});

if (events.readyState === "open") {
events.send(JSON.stringify({type: "session.close"}));
}

// Close `peer` after the completed event arrives if your application needs
// the authoritative final transcript.

Treat conversation.item.input_audio_transcription.completed as the authoritative success signal. session.closed only confirms that the WebRTC session ended after the close request; it does not prove that transcription completed. The session.close and session.closed lifecycle events are KittenML extensions shared by the WebSocket and WebRTC transports. If speech or a partial transcript was observed but the session closes before a completed event, report the request as incomplete. If no speech, partial, or completed event was observed, report No speech detected rather than Finished.

Events and limits​

WebRTC returns the same session, partial, completed, and error events documented in Realtime events. Replace cumulative clean_text and enriched_text snapshots rather than concatenating them.

WebSocket and WebRTC share the same five active ASR slots per organization. There is no fixed WebRTC call-duration limit. If SDP admission returns HTTP 429, retry with bounded exponential backoff and jitter.

For production error handling, UI state, cleanup, and a working token server, use the complete KittenML WebRTC browser example as the starting point.