Skip to main content
Open a WebSocket, push audio, receive text as it is recognised. There are two streaming endpoints. They speak the same protocol — same authentication, same audio, same messages — so switching is a URL change. See Which endpoint should I use? for the differences.

Authenticating

Never do either of these from a browser you do not control — the key is visible to anyone who opens devtools. Stream from your server, or proxy through it.

Languages

Pass BCP-47 codes, up to 4, either as ?languages= on the URL or as the languages array in the auth message. The auth message wins if both are set. Pass auto on its own to detect the language automatically. The option behaves differently on the two endpoints:
  • On /stream it restricts recognition to the codes you pass.
  • On /v2/stream it biases recognition toward them. The engine still follows a speaker who switches to another language mid-call.
To check which codes are supported, and how stable each one is, call GET /languages:
On /v2/stream, pass the languages you expect even though it can auto-detect. Without a hint the first few seconds of a call are noticeably less accurate.

Sending audio

Wait for ready, then send binary chunks:
  • 16 kHz, 16-bit signed PCM, mono — no WAV header, no compression
  • 20 ms per chunk — 640 bytes
To convert a file for testing: ffmpeg -i in.mp3 -ac 1 -ar 16000 -f s16le out.pcm

Ending the stream

When the speaker is done, send an end message and keep the socket open until the last final result arrives — usually well under a second. Closing the socket immediately can lose the final utterance, especially on /v2/stream.

Messages you receive

object
Authentication succeeded. Start sending audio.
languages echoes what the engine was configured with, so you can confirm your option was applied.
object
A result. transcript holds the text; isFinal says whether it is settled.
Interim results arrive fast and may be revised — good for showing live captions. Final results are stable — use those for anything you store or act on.language is the detected language of that result on /stream; it is empty on /v2/stream. confidence is always 0 on interims and on /v2/stream.
object
Something went wrong. error explains what; code is machine-readable.

A complete example

Render finals and interims differently — settled text in your normal colour, the interim tail dimmed. Users read a caption that visibly firms up as accurate; one that silently rewrites itself reads as broken.

Which endpoint should I use?

Both take the same audio and return the same messages. The difference is the recognition engine behind them. /stream is the standard engine. It is generally available, finalises phrases quickly, and tags every result with the detected language. The languages option is a hard filter. Use it when you need stable, predictable behaviour and per-result language codes. /v2/stream is the next-generation engine, currently in preview. It streams interim text continuously, roughly twice a second, so captions feel much more responsive. It produces cleaner, more sentence-like finals and handles callers who switch language mid-conversation. The languages option is a hint, not a filter. It does not return per-result language codes or confidence scores, and without a language hint the opening seconds of a call are less accurate. If in doubt, test both on your own recordings with the languages you expect and compare.

Keeping the stream healthy

  • Send steadily. Bursting a buffer is worse than a paced 20 ms cadence.
  • Reconnect on close. Networks drop; keep the transcript you already finalised and resume.
  • Send end before you close, so the last result is flushed.
  • Long calls are fine. The service reconnects to the engine behind the scenes on long sessions; you will not see a gap.