See Which endpoint should I use? for the
differences.
Authenticating
- Subprotocol (recommended)
- First message (fallback)
Pass the key as a WebSocket subprotocol. It never appears in a URL, so it stays
out of proxy and server logs. Options such as languages go in the query string.
Languages
Pass BCP-47 codes, up to 4, either as?languages= on the URL or as the
languages array in the auth message. The auth message wins if both are set.
Pass auto on its own to detect the language automatically.
The option behaves differently on the two endpoints:
- On
/streamit restricts recognition to the codes you pass. - On
/v2/streamit biases recognition toward them. The engine still follows a speaker who switches to another language mid-call.
GET /languages:
Sending audio
Wait forready, then send binary chunks:
- 16 kHz, 16-bit signed PCM, mono — no WAV header, no compression
- 20 ms per chunk — 640 bytes
ffmpeg -i in.mp3 -ac 1 -ar 16000 -f s16le out.pcm
Ending the stream
When the speaker is done, send anend message and keep the socket open until
the last final result arrives — usually well under a second. Closing the socket
immediately can lose the final utterance, especially on /v2/stream.
Messages you receive
object
Authentication succeeded. Start sending audio.
languages echoes what the engine was configured with, so you can confirm
your option was applied.object
A result. Interim results arrive fast and may be revised — good for showing live
captions. Final results are stable — use those for anything you store or
act on.
transcript holds the text; isFinal says whether it is settled.language is the detected language of that result on /stream; it is empty
on /v2/stream. confidence is always 0 on interims and on /v2/stream.object
Something went wrong.
error explains what; code is machine-readable.A complete example
Which endpoint should I use?
Both take the same audio and return the same messages. The difference is the recognition engine behind them./stream is the standard engine. It is generally available, finalises
phrases quickly, and tags every result with the detected language. The
languages option is a hard filter. Use it when you need stable, predictable
behaviour and per-result language codes.
/v2/stream is the next-generation engine, currently in preview. It streams
interim text continuously, roughly twice a second, so captions feel much more
responsive. It produces cleaner, more sentence-like finals and handles callers
who switch language mid-conversation. The languages option is a hint, not a
filter. It does not return per-result language codes or confidence scores, and
without a language hint the opening seconds of a call are less accurate.
If in doubt, test both on your own recordings with the languages you expect and
compare.
Keeping the stream healthy
- Send steadily. Bursting a buffer is worse than a paced 20 ms cadence.
- Reconnect on close. Networks drop; keep the transcript you already finalised and resume.
- Send
endbefore you close, so the last result is flushed. - Long calls are fine. The service reconnects to the engine behind the scenes on long sessions; you will not see a gap.