Skip to main content
The SDK is a convenience. If you are not in a browser — a mobile app, a server bridge, your own telephony stack — connect to the WebSocket yourself.

Getting a URL

The response carries websocketUrl:
The field is websocketUrl. It is short-lived and single-use — create a fresh one per call rather than caching it. No additional authentication is needed on the socket itself; the URL is the credential, which is also why it must not be shared or logged.

Audio

Binary frames carry audio; text frames carry JSON. For pcm, send 16-bit mono at the inputSampleRate you requested.

Messages you receive

Transcripts

A transcript message is one piece of one utterance. Here is what one short exchange looks like on the socket, in the order the frames arrive:
Three things to take from that:
  • The agent streams. You get its words as it says them, then a closing message with the whole sentence.
  • The caller arrives as the final version. Their sentence comes as one message, already final: true, so there is nothing to assemble.
  • Order by ordinal. The caller’s utterance always has a lower ordinal than the reply it prompted, so sort by ordinal rather than by arrival.
Each field exists to let you handle exactly that:
string
required
user or agent. Decides which side of the conversation the text goes on.
integer
required
Which utterance this piece belongs to. All the agent’s delta pieces above share ordinal 0, so you know to join them. It is also the conversation order: the caller’s ordinal 1 belongs before the agent’s 2, whatever order the messages arrive in.
string
The utterance so far, in full. Present on the first piece and the last one. When you see it, replace whatever you have for that ordinal.
string
A few more characters for the same ordinal. When you see it, append. A message has text or delta, never both.
boolean
required
false means more pieces for this ordinal are coming, true means this is the last one. Use it to stop showing a typing indicator, or to know a caller turn is complete before you act on it.
string
default:"voice"
voice if it was spoken, text if you injected it with user_text_message. Useful if you want to style typed input differently.
Put together, a handler is a dozen lines:

Messages you send

Client tool calls

An invocation arrives as:
Do the work, then reply with the same invocationId:
The invocationId is how a result is matched to its call. Send a different one — or none — and the agent waits for an answer that never arrives, which the caller hears as the agent freezing mid-sentence.
Report failures rather than staying silent, so the agent can explain and move on:
Using the SDK? registerTool handles all of this — invocation matching, results, and errors — so you never touch invocationId yourself.

Injecting a message

Push text into a live call as though the caller had spoken it. This is how a deferred tool result gets back into the conversation:

Ending

Close the socket to hang up, or watch for the hangup message when the agent ends it. For the agent to be able to end a call gracefully, assign the hangUp system tool — without it, a finished agent simply waits.