Limits
Size caps, timeouts, and concurrency limits on the Router.
Size limits
| Limit | Value | Exceeded → |
|---|---|---|
| JSON request body | 1 MiB | 413 payload_too_large |
STT request multipart part | 1 MiB | 413 payload_too_large |
Batch STT audio (/v1/stt/transcriptions) | Selected model's advertised batch limit, plus any deployment safety ceiling | 413 payload_too_large |
Transcription job upload (/v1/stt/transcriptions/jobs) | 2 GiB (deployment default), applied to the upload and to the decoded audio of a compressed upload; the model's limit is handled by chunking | 413 payload_too_large |
| WebSocket frame | 1 MiB | stream error |
Idempotency-Key header | 256 bytes | 400 invalid_request |
Batch transcription has no universal 25 MiB cap. Limits are provider- and
model-specific, and constrain decoded PCM bytes, recording duration, or both;
they are the bounds of each provider's pre-recorded API. See
Batch transcription limits for the current table and
query GET /v1/models?path=batch for the authoritative catalog in your serving
region. Recordings over a model's limit belong on
transcription jobs, which split them into provider-sized
chunks.
Time limits
| Limit | Value |
|---|---|
| Request body read deadline | 2 minutes |
WebSocket session.configure deadline | 10 seconds after connect |
Keepalive cadence (WS pings, SSE : heartbeat comments) | every 20 seconds |
There is deliberately no overall wall-clock cap on streaming sessions — liveness is enforced by heartbeats and by budget and lease extension.
Concurrency
Three independent layers can refuse work with 429 concurrency_exhausted (retryable):
- Per-organization slots — your account's concurrent-request limit, enforced at admission. Contact Speko to raise it.
- Per-edge stream slots — a local cap on any single Router edge task. Retrying lands on capacity elsewhere.
- Batch-audio staging capacity — temporary disk reserved for in-flight transcription uploads. Retrying can succeed after capacity is released or on another task.
Transcription jobs add a fourth, per-organization ceiling: at most 64 jobs and 32 GiB of decoded audio in flight at once. See job limits.
429 rate_limited is distinct: it is request-rate pressure rather than slot exhaustion. Back off in both cases.
Regional behavior
router.speko.dev is anycast (AWS Global Accelerator, static IPv4 addresses); connections land on the nearest healthy region and stay there for their lifetime. /v1/models reflects the serving region's catalog, and every response names its region in Speko-Region. If a region is drained for maintenance, new connections shift automatically; in-flight streams on a draining task are given a generous window to complete.