Speko Docs
Speech to text

Transcribe audio

POST /v1/stt/transcriptions — transcribe an audio file in one request.

Request

POST https://router.speko.dev/v1/stt/transcriptions

The body is multipart/form-data with exactly two parts, in this order:

  1. request — a JSON document (at most 1 MiB)
  2. audio — WAV containing PCM s16le, or raw PCM s16le with an audio object in the request part

Part order is load-bearing: the idempotency content hash is defined over the parts in this order, and extra or reordered parts are rejected.

curl -s https://router.speko.dev/v1/stt/transcriptions \
  -H "Authorization: Bearer $SPEKO_API_KEY" \
  -H "Idempotency-Key: $(uuidgen)" \
  -F 'request={"routing":{"mode":"auto","objective":"quality"},"language":"en"};type=application/json' \
  -F 'audio=@interview.wav;type=audio/wav'

The request is held open until the transcript is ready. For recordings longer than a few minutes, or larger than the selected model's limit, use transcription jobs instead — they take audio of any length and deliver to a webhook.

The request part

routingobject

The routing object. Omit for auto/balanced.

languagestring

Optional language hint, for example en or es-MX. Without it the provider detects the language where it can.

optionsobject
audioobject

Required only for raw PCM uploads. Set encoding to pcm_s16le and provide sample_rate_hz and channels. For WAV uploads, the Router derives these values from the container.

The options object

Options are provider-neutral asks. They fail closed: a request routed to a model that cannot honor an option is refused with 400 capability_unsupported naming the option, never answered 200 with the feature silently missing — a diarization request served by a model that cannot diarize is indistinguishable from a one-speaker recording, so nobody would ever notice. GET /v1/models advertises each STT model's capabilities; in automatic routing, models lacking a requested capability are skipped.

diarizationboolean

Label each segment with the speaker who said it. Segments gain a speaker field carrying the provider's own label verbatim ("0", "A", "speaker_1").

keywordsstring[]

Up to 100 terms of at most 64 characters to bias recognition toward — names, products, jargon. Translated to each vendor's own mechanism (Deepgram keyterm, AssemblyAI keyterms_prompt, ElevenLabs keyterms, Soniox context terms, OpenAI prompt text).

noise_reductionboolean

Ask the provider to clean the audio before transcribing. No currently routed STT model advertises it, so setting it is refused today.

word_timestampsboolean

Return per-word timings. The response gains a words array of { text, start_ms, end_ms }, in reading order, measured from the start of the audio — what a subtitle builder needs. Served by gemini-3.5-transcribe today (pin it with explicit routing; GET /v1/models?path=batch shows word_timestamps per model). Gemini cannot combine it with keywords, so that pair is refused capability_unsupported naming keywords.

provider_optionsobject

A vendor's own settings, namespaced by provider name and allow-listed per provider — for example {"deepgram": {"numerals": true}}. Settings for a provider the request does not reach are ignored, so you can tune several providers and let routing decide. Scalars only; at most 8 providers and 16 keys each. Settings the Router owns (model, language, encoding, diarize, credentials, …) are rejected by name.

{
  "routing": {"mode": "explicit", "provider": "deepgram", "model": "nova-3"},
  "language": "en",
  "options": {
    "diarization": true,
    "keywords": ["Speko", "Router"],
    "provider_options": {"deepgram": {"numerals": true, "smart_format": true}}
  }
}

Response

{
  "text": "Good morning everyone, let's get started.",
  "segments": [
    { "text": "Good morning everyone, let's get started.", "start_ms": 0, "end_ms": 2900, "speaker": "0" }
  ],
  "route": { "provider": "deepgram", "model": "nova-3", "region": "us-east-1", "attempt_id": "ratt_..." },
  "usage": { "duration_ms": 3000 }
}

With "options": {"word_timestamps": true} on gemini-3.5-transcribe:

{
  "text": "Salom dunyo",
  "segments": [{ "text": "Salom dunyo", "start_ms": 120, "end_ms": 1000 }],
  "words": [
    { "text": "Salom", "start_ms": 120, "end_ms": 480 },
    { "text": "dunyo", "start_ms": 510, "end_ms": 1000 }
  ],
  "route": { "provider": "gemini", "model": "gemini-3.5-transcribe", "region": "us-east-1", "attempt_id": "ratt_..." },
  "usage": { "duration_ms": 1000 }
}
textstring

The full transcript. May legitimately be empty for silent audio.

segmentsobject[]

Time-aligned utterances. start_ms/end_ms are non-negative, ordered ranges. speaker is present only when diarization was requested. Some models return whole-text only, in which case segments is omitted.

wordsobject[]

Per-word timings, present only when you asked for word_timestamps. start_ms/end_ms are non-negative, ordered, and measured from the start of the audio. speaker appears on each word only when diarization was also requested and the provider labels speakers.

routeobject

Which provider, model, and Speko region served the request, and the attempt ID. model is the catalog id you can pin; when the provider's pre-recorded API uses a different model name internally (Soniox stt-async-v5 behind stt-rt-v5, ElevenLabs scribe_v2 behind scribe_v2_realtime), you still see the catalog id.

usageobject

duration_ms — the audio duration metered: the duration the provider reports processing, capped at the duration of the audio you sent as parsed from the container, never a caller-declared value.

The same route facts are echoed in the Speko-Request-ID, Speko-Provider, Speko-Model and Speko-Region response headers.

How your audio is transcribed

Every STT model in the catalog is served one of two ways, and GET /v1/models?path=batch tells you which:

  • Pre-recorded route. The upload is sent as one file to the provider's batch/asynchronous transcription API — Deepgram /v1/listen, AssemblyAI's Transcripts API, Soniox async, ElevenLabs Scribe, OpenAI /v1/audio/transcriptions, and so on. The whole recording is transcribed in one call, bounded by that API's documented limits. These are the models listed under path=batch, and the only models transcription jobs will use.
  • Realtime-only. A few models exist only as live sockets (Deepgram Flux, Cartesia ink-2, OpenAI gpt-live-transcribe, Palabra). The Router still transcribes short uploads on them by replaying the audio into the socket, bounded to 60 seconds — the most a realtime socket reliably keeps up with when fed faster than real time. They are omitted from path=batch and refused for jobs.

Either way the Router checks that the transcript actually covers the audio: a response whose last timed segment stops well short of the recording while speech continues is refused as a retryable provider_error rather than returned — and, in automatic mode, triggers failover.

Limits

There is no universal 25 MiB transcription cap. Each model's limit is the bound of its provider's pre-recorded API on one upload — decoded PCM bytes, duration, or both. Query GET /v1/models?path=batch and read batch_audio_limits for the authoritative values in your region. Current limits:

Provider / modelOne upload acceptsNotes
Deepgram nova-3, nova-22 GiB and 2 hoursDeepgram bounds processing time, not duration; 2 h keeps well inside it
AssemblyAI universal-3-5-pro2.2 GB and 10 hoursMono only
Soniox stt-rt-v55 hoursServed by stt-async-v5; no byte limit published
ElevenLabs scribe_v2_realtime2 GiB and 10 hoursServed by scribe_v2; mono only
OpenAI gpt-transcribe, gpt-4o-transcribe, gpt-realtime-whisper25 MB and 30 minutes24 kHz mono only — about 8 minutes of audio per upload; gpt-realtime-whisper is served by whisper-1
Cartesia ink-whisper1 GiB and 2 hoursMono only; ink-2 has no batch API
xAI grok-stt500 MB and 2 hoursMono only
Speechmatics standard, enhanced1 GiB and 4 hoursMono only
Azure MAI-Transcribe-2314,556,416 decoded PCM bytes (≈300 MB) and 5 hoursMono only; batch-only (no streaming route)
Meta muse-voice-transcribe-1.033,538,048 decoded PCM bytes and 10 minutesMono 16 or 24 kHz
Fish Audio transcribe-1, transcribe-1-pro46,120,960 decoded PCM bytes and 3 minutesMono only; batch-only (no streaming route). transcribe-1-pro labels speakers when diarization is set and keeps emotion and sound cues such as [laughter] in the text
Gemini gemini-3.5-transcribe, gemini-3.5-transcribe-live15,712,256 decoded PCM bytes and 8 minutesMono 16 kHz; Gemini's 20 MiB request limit, allowing for base64 expansion. gemini-3.5-transcribe is batch-only
Modulate velma-2-stt-streaming, velma-2-stt-streaming-english-v2100 MB and 100 minutesAbout 52 minutes of 16 kHz mono
Smallest pulse250 MB and 10 minutesMono only
Hamsa s3500 MB and 60 minutesMono 16 kHz; jobs only (the provider fetches audio by URL)
Alibaba qwen3-asr-flash-realtime2 GiB and 12 hoursMono 8/16 kHz; jobs only (the provider fetches audio by URL)
Google chirp_360 seconds and 7,847,936 decoded PCM bytesGoogle's 10 MB synchronous request limit, allowing for base64 expansion
Deepgram flux-*, Cartesia ink-2, OpenAI gpt-live-transcribe, Palabra default60 secondsRealtime-only; not listed under path=batch

On this endpoint a recording over the limit is refused, not split:

  • With automatic routing, models that cannot accept the recording are skipped. If no compatible model remains, the Router returns 413 payload_too_large before admission or provider contact.
  • With explicit routing, a recording over the selected model's limit returns 413 payload_too_large before admission or provider contact — the hint names the limit and suggests GET /v1/models?path=batch.
  • Recordings that exceed every model's limit belong on transcription jobs, which chunk them.

Other bounds:

  • A deployment may apply an additional safety ceiling. This is an operational override, not a provider capability, so it is not included in batch_audio_limits.
  • Batch audio is staged on temporary disk. If the Router task has no staging capacity, it returns retryable 429 concurrency_exhausted; this is capacity pressure, not a file-size limit.
  • The canonical WAV transport has a format boundary of 4,294,967,258 decoded PCM bytes (about 4 GiB) for every provider.
  • This endpoint accepts WAV (PCM s16le) and described raw PCM only. Anything else — including MP3, M4A, OGG, FLAC and WebM — returns 415 unsupported_media; so does a rate or channel layout the pinned model does not accept — the hint lists what it does. Transcription jobs decode compressed formats for you.
  • The whole request body must complete within the 2-minute read deadline.
  • Metering is by audio duration, rounded up to whole seconds. There is deliberately no usage header on STT responses — the usage object in the body is authoritative.
  • In auto mode a provider failure before results triggers transparent failover.

On this page