Skip to main content
Transcription is available on its own, without building a voice agent. Send audio in, get text back.
Authenticate with the same key you use for the rest of the API, in the X-API-Key header.

Languages

Every endpoint accepts a languages option: BCP-47 codes, up to 4, or auto to detect the language. On the REST endpoints pass it as the X-Languages header or ?languages= query parameter. The default is Finnish and Swedish (fi-FI,sv-SE) except on /v2/stream, which auto-detects.

Supported languages

GET /languages lists every supported code with a status of stable or preview. No key needed. Add ?codes= to check specific ones before you build on them.
v1 covers /transcribe, /batch and /stream; v2 covers /v2/stream.

Transcribe a file

Up to 10 MB and 60 seconds of audio. For anything longer, use /batch.

Response

string
The full text.
array
The text split by the engine, each with the detected language.
integer
Audio length in milliseconds, estimated from the file. This is what you are billed on.
integer
How long recognition took, in milliseconds.

Transcribe long audio

POST /batch takes the same headers as /transcribe, accepts up to 500 MB and 8 hours, and returns an operation to poll. Add X-Diarization: true to get a speakerTurns array saying who said what.
Results are kept for 24 hours. Direct uploads are capped at 32 MB by the platform; for larger files call POST /batch/init, PUT the audio to the returned uploadUrl, then POST /batch/{operationId}/start. The full flow is in the API reference at https://stt.omnia-voice.com/.

Audio formats

Send the matching Content-Type.
Quality in beats quality out. 16 kHz mono is plenty for speech and keeps uploads small; a noisy 48 kHz stereo recording transcribes worse than a clean 16 kHz one, not better.

Billing

4 credits per minute ($0.04), rounded up, one-minute minimum. See Credits.

Live streaming

Transcribe audio as it is spoken, with interim results.