Speko Docs
Text to speech

Streaming synthesis

GET /v1/tts/stream — incremental text in, audio out, over WebSocket.

Connecting

Open a WebSocket to wss://router.speko.dev/v1/tts/stream with Authorization and Idempotency-Key headers on the upgrade, then send session.configure as the first frame (within 10 seconds):

{
  "type": "session.configure",
  "routing": { "mode": "auto", "objective": "latency" },
  "audio": { "encoding": "pcm_s16le", "sample_rate_hz": 24000, "channels": 1 }
}

Optionally include "voice" to pin a provider voice ID. As with one-shot synthesis, streaming TTS is currently English-only — see the note on the speech endpoint.

Client frames

FrameMeaning
session.configureFirst frame; fixes routing, audio format, and the idempotency hash
input.append{"type":"input.append","text":"..."} — add text to synthesize (text non-empty)
input.commitFlush: synthesize everything appended so far
input.cancelStop synthesizing the current utterance
session.closeEnd the session

Binary frames from the client are rejected — TTS input is text.

Server frames

Audio arrives as binary frames between utterance.started and utterance.done markers:

FrameMeaning
session.readyAlways first: {"type":"session.ready","request_id":"...","route":{...}}
utterance.started{"type":"utterance.started","sequence":1} — sequence is the 1-based utterance index
(binary frames)Audio for the current utterance, in the configured format
utterance.timingsWord-level time alignment for the current utterance, on routes that measure it
utterance.done{"type":"utterance.done","sequence":1}
usage.updatedRunning character usage
session.closedClean end, with final usage
errorTerminal failure — standard error envelope

Exactly one terminal frame (session.closed or error) ends every session.

Word timings

Routes that measure synthesis timing emit utterance.timings for each utterance, after its audio and immediately before utterance.done:

{
  "type": "utterance.timings",
  "sequence": 1,
  "granularity": "word",
  "spans": [
    { "text": "Welcome", "start_ms": 111, "end_ms": 440 },
    { "text": "to", "start_ms": 520, "end_ms": 520 },
    { "text": "today.", "start_ms": 920, "end_ms": 1240 }
  ]
}
sequencenumber

The utterance these spans belong to, matching the index utterance.started carried.

granularitystring

word on every route that emits timings today. Readings from an engine that measures per character are grouped into words by the Router, so one shape covers every provider.

spansobject[]

Time-aligned spans with start_ms/end_ms measured in milliseconds from the start of this utterance. end_ms is an end time, not a duration.

Spans cover the text you sent rather than a provider's own rewrite of it, so they line up with the string you already hold. Match them to your text word by word: joining the spans back together will not reproduce your input exactly.

The frame follows the audio it describes, so it cannot drive a highlight that tracks playback as the audio arrives. A very long utterance may be split across several frames, all carrying the same sequence; concatenate them in arrival order.

Each span is well-formed on its own, but spans are never checked against each other. Identical start and end times turn up in ordinary output, and engines may return adjacent spans that overlap, so a consumer assuming a strictly increasing timeline will break on real responses.

An utterance stopped with input.cancel produces no timings. Its spans describe audio you never received, so the Router discards them rather than attribute them to the utterance that follows.

Timings are not something you request: there is no field or header that turns them on. What decides whether you get them is the provider behind the route.

ProviderEngine measuresYou receive
Cartesiawordsword spans
ElevenLabscharactersword spans
Sonioxcharactersword spans

Every other TTS provider in the catalog sends nothing today. The catalog moves, so read capabilities.word_timings from GET /v1/models rather than pinning that list: a route without the flag sends no utterance.timings frame at all, rather than an empty one.

Metering and budgets

Characters are counted when an input.append is accepted — before synthesis, and regardless of whether you later cancel. input.cancel stops audio, not billing, for text already accepted.

Streaming sessions admit against an initial character budget and extend it transparently as you append more text. If your organization's credit cannot cover an extension, the next append terminates the stream with budget_exhausted; if the session's lease cannot be renewed, it terminates with lease_expired. Both codes appear only on established streams.

Liveness

The Router pings every 20 seconds. Frames are capped at 1 MiB.

On this page