Speko Docs

Quickstart

Call Router — stream a live transcript, synthesize, and generate in five minutes.

1. Get an API key

Create a key in the dashboard's Gateway keys page and export it:

export SPEKO_API_KEY=sk_speko_...

2. List routable models

curl -s https://router.speko.dev/v1/models \
  -H "Authorization: Bearer $SPEKO_API_KEY"
{
  "models": [
    {
      "id": "deepgram:flux-general-en",
      "provider": "deepgram",
      "kind": "stt",
      "capabilities": { "tools": false, "structured_output": false, "cached_input": false, "reasoning": false },
      "regions": ["us-east-1"]
    }
  ],
  "catalog_digest": "sha256:..."
}

3. Stream a live transcript

Voice agents transcribe a caller who is still speaking, so the Router's main transcription surface is a WebSocket rather than an upload. Open wss://router.speko.dev/v1/stt/stream and send session.configure as the first frame. Push audio as binary frames, read hypotheses back as they arrive, then close the session.

Both headers go on the upgrade request, which browsers cannot set — connect from your backend, or from whatever already holds the call audio.

import WebSocket from 'ws';
import { randomUUID } from 'node:crypto';

const socket = new WebSocket('wss://router.speko.dev/v1/stt/stream', {
  headers: {
    Authorization: `Bearer ${process.env.SPEKO_API_KEY}`,
    'Idempotency-Key': randomUUID(),
  },
});

socket.on('open', () => {
  // The first frame, within 10 seconds of connecting.
  socket.send(JSON.stringify({
    type: 'session.configure',
    routing: { mode: 'auto', objective: 'latency' },
    audio: { encoding: 'pcm_s16le', sample_rate_hz: 16000, channels: 1 },
    language: 'en',
  }));

  // Then the live audio itself, as binary frames.
  callAudio.on('data', (chunk) => socket.send(chunk));
  callAudio.on('end', () => {
    // Ask for the remaining finals, then end the session.
    socket.send(JSON.stringify({ type: 'input.commit' }));
    socket.send(JSON.stringify({ type: 'session.close' }));
  });
});

socket.on('message', (raw, isBinary) => {
  if (isBinary) return;
  const frame = JSON.parse(raw.toString());
  if (frame.type === 'transcript.delta') process.stdout.write(frame.text);
  // One per finalized utterance, so a long call emits many of these.
  if (frame.type === 'transcript.final') console.log('\n✓', frame.text);
  // Exactly one terminal frame ends the session: session.closed or error.
  if (frame.type === 'session.closed' || frame.type === 'error') socket.close();
});
import asyncio, json, os, uuid
import websockets

async def main():
    async with websockets.connect(
        "wss://router.speko.dev/v1/stt/stream",
        additional_headers={
            "Authorization": f"Bearer {os.environ['SPEKO_API_KEY']}",
            "Idempotency-Key": str(uuid.uuid4()),
        },
    ) as socket:
        # The first frame, within 10 seconds of connecting.
        await socket.send(json.dumps({
            "type": "session.configure",
            "routing": {"mode": "auto", "objective": "latency"},
            "audio": {"encoding": "pcm_s16le", "sample_rate_hz": 16000, "channels": 1},
            "language": "en",
        }))

        async def send_audio():
            async for chunk in call_audio():   # live audio, as binary frames
                await socket.send(chunk)
            # Ask for the remaining finals, then end the session.
            await socket.send(json.dumps({"type": "input.commit"}))
            await socket.send(json.dumps({"type": "session.close"}))

        async def read_frames():
            async for raw in socket:
                frame = json.loads(raw)
                if frame["type"] == "transcript.delta":
                    print(frame["text"], end="", flush=True)
                elif frame["type"] == "transcript.final":
                    # One per finalized utterance, not one per session.
                    print("\n✓", frame["text"])
                elif frame["type"] in ("session.closed", "error"):
                    break   # exactly one terminal frame ends the session

        await asyncio.gather(send_audio(), read_frames())

asyncio.run(main())

session.ready always arrives first and names the route that was picked:

{ "type": "session.ready", "request_id": "rreq_...", "route": { "provider": "deepgram", "model": "flux-general-en", "region": "us-east-1", "attempt_id": "ratt_..." } }

Then transcript.delta frames carry interim hypotheses and transcript.final carries finalized text. See streaming transcription for the full frame set, the liveness rules, and the terminal-frame contract.

For a recording you already hold, Transcribe audio is the batch path. To put a phone call on the other end of this socket, see phone agents.

4. Synthesize speech

The response body is raw audio; the character count you were metered for comes back in the Speko-Usage-Characters header.

curl -s https://router.speko.dev/v1/tts/speech \
  -H "Authorization: Bearer $SPEKO_API_KEY" \
  -H "Idempotency-Key: $(uuidgen)" \
  -H "Content-Type: application/json" \
  -d '{
    "routing": { "mode": "auto", "objective": "latency" },
    "input": "Your table is ready.",
    "audio": { "encoding": "pcm_s16le", "sample_rate_hz": 24000, "channels": 1 }
  }' \
  --output speech.pcm

5. Generate a model response

max_output_tokens is required — an unbounded generation cannot be priced up front.

curl -s https://router.speko.dev/v1/llm/responses \
  -H "Authorization: Bearer $SPEKO_API_KEY" \
  -H "Idempotency-Key: $(uuidgen)" \
  -H "Content-Type: application/json" \
  -d '{
    "routing": { "mode": "explicit", "provider": "anthropic", "model": "claude-sonnet-5" },
    "input": [
      { "type": "message", "role": "user", "content": [{ "type": "text", "text": "Say hello in French." }] }
    ],
    "max_output_tokens": 100
  }'
{
  "id": "resp_rreq_...",
  "route": { "provider": "anthropic", "model": "claude-sonnet-5", "region": "us-east-1", "attempt_id": "ratt_..." },
  "output": [
    { "type": "message", "role": "assistant", "content": [{ "type": "text", "text": "Bonjour !" }] }
  ],
  "stop_reason": "stop",
  "usage": { "input_tokens": 21, "output_tokens": 5 }
}

Next

On this page