Speko Docs

Limits

Size caps, timeouts, and concurrency limits on the Router.

Size limits

LimitValueExceeded →
JSON request body1 MiB413 payload_too_large
STT request multipart part1 MiB413 payload_too_large
Batch STT audio (/v1/stt/transcriptions)Selected model's advertised batch limit, plus any deployment safety ceiling413 payload_too_large
Transcription job upload (/v1/stt/transcriptions/jobs)2 GiB (deployment default), applied to the upload and to the decoded audio of a compressed upload; the model's limit is handled by chunking413 payload_too_large
WebSocket frame1 MiBstream error
Idempotency-Key header256 bytes400 invalid_request

Batch transcription has no universal 25 MiB cap. Limits are provider- and model-specific, and constrain decoded PCM bytes, recording duration, or both; they are the bounds of each provider's pre-recorded API. See Batch transcription limits for the current table and query GET /v1/models?path=batch for the authoritative catalog in your serving region. Recordings over a model's limit belong on transcription jobs, which split them into provider-sized chunks.

Time limits

LimitValue
Request body read deadline2 minutes
WebSocket session.configure deadline10 seconds after connect
Keepalive cadence (WS pings, SSE : heartbeat comments)every 20 seconds

There is deliberately no overall wall-clock cap on streaming sessions — liveness is enforced by heartbeats and by budget and lease extension.

Concurrency

Three independent layers can refuse work with 429 concurrency_exhausted (retryable):

  1. Per-organization slots — your account's concurrent-request limit, enforced at admission. Contact Speko to raise it.
  2. Per-edge stream slots — a local cap on any single Router edge task. Retrying lands on capacity elsewhere.
  3. Batch-audio staging capacity — temporary disk reserved for in-flight transcription uploads. Retrying can succeed after capacity is released or on another task.

Transcription jobs add a fourth, per-organization ceiling: at most 64 jobs and 32 GiB of decoded audio in flight at once. See job limits.

429 rate_limited is distinct: it is request-rate pressure rather than slot exhaustion. Back off in both cases.

Regional behavior

router.speko.dev is anycast (AWS Global Accelerator, static IPv4 addresses); connections land on the nearest healthy region and stay there for their lifetime. /v1/models reflects the serving region's catalog, and every response names its region in Speko-Region. If a region is drained for maintenance, new connections shift automatically; in-flight streams on a draining task are given a generous window to complete.

On this page