Transcribe audio
POST /v1/stt/transcriptions — transcribe an audio file in one request.
Request
POST https://router.speko.dev/v1/stt/transcriptions
The body is multipart/form-data with exactly two parts, in this order:
request— a JSON document (at most 1 MiB)audio— WAV containing PCM s16le, or raw PCM s16le with anaudioobject in therequestpart
Part order is load-bearing: the idempotency content hash is defined over the parts in this order, and extra or reordered parts are rejected.
curl -s https://router.speko.dev/v1/stt/transcriptions \
-H "Authorization: Bearer $SPEKO_API_KEY" \
-H "Idempotency-Key: $(uuidgen)" \
-F 'request={"routing":{"mode":"auto","objective":"quality"},"language":"en"};type=application/json' \
-F 'audio=@interview.wav;type=audio/wav'The request is held open until the transcript is ready. For recordings longer than a few minutes, or larger than the selected model's limit, use transcription jobs instead — they take audio of any length and deliver to a webhook.
The request part
routingobjectThe routing object. Omit for auto/balanced.
languagestringOptional language hint, for example en or es-MX. Without it the provider
detects the language where it can.
optionsobjectOptional transcription options.
audioobjectRequired only for raw PCM uploads. Set encoding to pcm_s16le and provide
sample_rate_hz and channels. For WAV uploads, the Router derives these
values from the container.
The options object
Options are provider-neutral asks. They fail closed: a request routed to a
model that cannot honor an option is refused with 400 capability_unsupported
naming the option, never answered 200 with the feature silently missing —
a diarization request served by a model that cannot diarize is
indistinguishable from a one-speaker recording, so nobody would ever notice.
GET /v1/models advertises each STT model's capabilities; in automatic
routing, models lacking a requested capability are skipped.
diarizationbooleanLabel each segment with the speaker who said it. Segments gain a speaker
field carrying the provider's own label verbatim ("0", "A",
"speaker_1").
keywordsstring[]Up to 100 terms of at most 64 characters to bias recognition toward — names,
products, jargon. Translated to each vendor's own mechanism (Deepgram
keyterm, AssemblyAI keyterms_prompt, ElevenLabs keyterms, Soniox
context terms, OpenAI prompt text).
noise_reductionbooleanAsk the provider to clean the audio before transcribing. No currently routed STT model advertises it, so setting it is refused today.
word_timestampsbooleanReturn per-word timings. The response gains a words array of
{ text, start_ms, end_ms }, in reading order, measured from the start of
the audio — what a subtitle builder needs. Served by gemini-3.5-transcribe
today (pin it with explicit routing; GET /v1/models?path=batch shows
word_timestamps per model). Gemini cannot combine it with keywords, so
that pair is refused capability_unsupported naming keywords.
provider_optionsobjectA vendor's own settings, namespaced by provider name and allow-listed per
provider — for example {"deepgram": {"numerals": true}}. Settings for a
provider the request does not reach are ignored, so you can tune several
providers and let routing decide. Scalars only; at most 8 providers and 16
keys each. Settings the Router owns (model, language, encoding,
diarize, credentials, …) are rejected by name.
{
"routing": {"mode": "explicit", "provider": "deepgram", "model": "nova-3"},
"language": "en",
"options": {
"diarization": true,
"keywords": ["Speko", "Router"],
"provider_options": {"deepgram": {"numerals": true, "smart_format": true}}
}
}Response
{
"text": "Good morning everyone, let's get started.",
"segments": [
{ "text": "Good morning everyone, let's get started.", "start_ms": 0, "end_ms": 2900, "speaker": "0" }
],
"route": { "provider": "deepgram", "model": "nova-3", "region": "us-east-1", "attempt_id": "ratt_..." },
"usage": { "duration_ms": 3000 }
}With "options": {"word_timestamps": true} on gemini-3.5-transcribe:
{
"text": "Salom dunyo",
"segments": [{ "text": "Salom dunyo", "start_ms": 120, "end_ms": 1000 }],
"words": [
{ "text": "Salom", "start_ms": 120, "end_ms": 480 },
{ "text": "dunyo", "start_ms": 510, "end_ms": 1000 }
],
"route": { "provider": "gemini", "model": "gemini-3.5-transcribe", "region": "us-east-1", "attempt_id": "ratt_..." },
"usage": { "duration_ms": 1000 }
}textstringThe full transcript. May legitimately be empty for silent audio.
segmentsobject[]Time-aligned utterances. start_ms/end_ms are non-negative, ordered
ranges. speaker is present only when diarization was requested. Some
models return whole-text only, in which case segments is omitted.
wordsobject[]Per-word timings, present only when you asked for word_timestamps.
start_ms/end_ms are non-negative, ordered, and measured from the start
of the audio. speaker appears on each word only when diarization was
also requested and the provider labels speakers.
routeobjectWhich provider, model, and Speko region served the request, and the attempt
ID. model is the catalog id you can pin; when the provider's pre-recorded
API uses a different model name internally (Soniox stt-async-v5 behind
stt-rt-v5, ElevenLabs scribe_v2 behind scribe_v2_realtime), you still
see the catalog id.
usageobjectduration_ms — the audio duration metered: the duration the provider reports
processing, capped at the duration of the audio you sent as parsed from the
container, never a caller-declared value.
The same route facts are echoed in the Speko-Request-ID, Speko-Provider,
Speko-Model and Speko-Region response headers.
How your audio is transcribed
Every STT model in the catalog is served one of two ways, and
GET /v1/models?path=batch tells you which:
- Pre-recorded route. The upload is sent as one file to the provider's
batch/asynchronous transcription API — Deepgram
/v1/listen, AssemblyAI's Transcripts API, Soniox async, ElevenLabs Scribe, OpenAI/v1/audio/transcriptions, and so on. The whole recording is transcribed in one call, bounded by that API's documented limits. These are the models listed underpath=batch, and the only models transcription jobs will use. - Realtime-only. A few models exist only as live sockets (Deepgram Flux,
Cartesia
ink-2, OpenAIgpt-live-transcribe, Palabra). The Router still transcribes short uploads on them by replaying the audio into the socket, bounded to 60 seconds — the most a realtime socket reliably keeps up with when fed faster than real time. They are omitted frompath=batchand refused for jobs.
Either way the Router checks that the transcript actually covers the audio:
a response whose last timed segment stops well short of the recording while
speech continues is refused as a retryable provider_error rather than
returned — and, in automatic mode, triggers failover.
Limits
There is no universal 25 MiB transcription cap. Each model's limit is the bound
of its provider's pre-recorded API on one upload — decoded PCM bytes, duration,
or both. Query GET /v1/models?path=batch and read batch_audio_limits for the
authoritative values in your region. Current limits:
| Provider / model | One upload accepts | Notes |
|---|---|---|
Deepgram nova-3, nova-2 | 2 GiB and 2 hours | Deepgram bounds processing time, not duration; 2 h keeps well inside it |
AssemblyAI universal-3-5-pro | 2.2 GB and 10 hours | Mono only |
Soniox stt-rt-v5 | 5 hours | Served by stt-async-v5; no byte limit published |
ElevenLabs scribe_v2_realtime | 2 GiB and 10 hours | Served by scribe_v2; mono only |
OpenAI gpt-transcribe, gpt-4o-transcribe, gpt-realtime-whisper | 25 MB and 30 minutes | 24 kHz mono only — about 8 minutes of audio per upload; gpt-realtime-whisper is served by whisper-1 |
Cartesia ink-whisper | 1 GiB and 2 hours | Mono only; ink-2 has no batch API |
xAI grok-stt | 500 MB and 2 hours | Mono only |
Speechmatics standard, enhanced | 1 GiB and 4 hours | Mono only |
Azure MAI-Transcribe-2 | 314,556,416 decoded PCM bytes (≈300 MB) and 5 hours | Mono only; batch-only (no streaming route) |
Meta muse-voice-transcribe-1.0 | 33,538,048 decoded PCM bytes and 10 minutes | Mono 16 or 24 kHz |
Fish Audio transcribe-1, transcribe-1-pro | 46,120,960 decoded PCM bytes and 3 minutes | Mono only; batch-only (no streaming route). transcribe-1-pro labels speakers when diarization is set and keeps emotion and sound cues such as [laughter] in the text |
Gemini gemini-3.5-transcribe, gemini-3.5-transcribe-live | 15,712,256 decoded PCM bytes and 8 minutes | Mono 16 kHz; Gemini's 20 MiB request limit, allowing for base64 expansion. gemini-3.5-transcribe is batch-only |
Modulate velma-2-stt-streaming, velma-2-stt-streaming-english-v2 | 100 MB and 100 minutes | About 52 minutes of 16 kHz mono |
Smallest pulse | 250 MB and 10 minutes | Mono only |
Hamsa s3 | 500 MB and 60 minutes | Mono 16 kHz; jobs only (the provider fetches audio by URL) |
Alibaba qwen3-asr-flash-realtime | 2 GiB and 12 hours | Mono 8/16 kHz; jobs only (the provider fetches audio by URL) |
Google chirp_3 | 60 seconds and 7,847,936 decoded PCM bytes | Google's 10 MB synchronous request limit, allowing for base64 expansion |
Deepgram flux-*, Cartesia ink-2, OpenAI gpt-live-transcribe, Palabra default | 60 seconds | Realtime-only; not listed under path=batch |
On this endpoint a recording over the limit is refused, not split:
- With automatic routing, models that cannot accept the recording are skipped. If no compatible model remains, the Router returns
413 payload_too_largebefore admission or provider contact. - With explicit routing, a recording over the selected model's limit returns
413 payload_too_largebefore admission or provider contact — the hint names the limit and suggestsGET /v1/models?path=batch. - Recordings that exceed every model's limit belong on transcription jobs, which chunk them.
Other bounds:
- A deployment may apply an additional safety ceiling. This is an operational override, not a provider capability, so it is not included in
batch_audio_limits. - Batch audio is staged on temporary disk. If the Router task has no staging capacity, it returns retryable
429 concurrency_exhausted; this is capacity pressure, not a file-size limit. - The canonical WAV transport has a format boundary of 4,294,967,258 decoded PCM bytes (about 4 GiB) for every provider.
- This endpoint accepts WAV (PCM s16le) and described raw PCM only. Anything else — including MP3, M4A, OGG, FLAC and WebM — returns
415 unsupported_media; so does a rate or channel layout the pinned model does not accept — the hint lists what it does. Transcription jobs decode compressed formats for you. - The whole request body must complete within the 2-minute read deadline.
- Metering is by audio duration, rounded up to whole seconds. There is deliberately no usage header on STT responses — the usage object in the body is authoritative.
- In auto mode a provider failure before results triggers transparent failover.