Speko Docs
Concepts

Sessions vs one-shot

When to mint a session and when to call a one-shot endpoint directly.

Speko supports two integration shapes. Pick based on whether your call is real-time interactive or single-turn batch.

One-shot endpoints

POST /v1/transcribe, POST /v1/synthesize, POST /v1/complete. Each is a single round-trip:

  • Caller sends input + intent.
  • Speko picks a provider, runs the call (with failover), returns the result.

Use for:

  • Batch transcription of recorded audio.
  • Server-side TTS for notifications, IVR prompts, exports.
  • LLM completions in a non-voice flow.

No state is held between calls. There is nothing to clean up.

Real-time sessions

POST /v1/sessions creates one of two different media paths. Set mode explicitly so your client knows which response shape to expect.

Provider-direct S2S

mode: "s2s" reserves a bounded Speko entitlement and returns a scoped, short-lived credential for OpenAI Realtime, Gemini Live, or xAI Grok Voice. The browser passes that response to RealtimeVoiceConversation and connects directly to the selected provider.

The response has transport: "provider_direct", providerTransport, endpoint, credential, telemetry, and reservation fields. It does not contain LiveKit credentials:

your backend ── create session ──▶ Speko control plane
                                       │
                                       └── scoped credential + entitlement
                                                     │
browser ───────────── WebRTC / WebSocket ───────────▶ provider

OpenAI uses WebRTC. Gemini Live and xAI use provider WebSockets. Speko stores content-free lifecycle and billing evidence. The OpenAI sideband observes provider control events but never receives media.

For S2S, ttlSeconds requests the maximum session horizon (default 1,800s, max 3,600s). The returned reservation.authorizedDurationSeconds may be lower because provider credentials can have shorter lifetimes. Gemini returns renewable slices; the browser client renews and resumes them while the requested horizon and credit authorization permit it.

Idempotency-Key is required. Reuse it after an ambiguous bootstrap timeout to recover the same reservation without minting a duplicate credential.

See Provider-direct speech-to-speech for a complete browser integration.

Cascade

mode: "cascade" returns a transportToken, transportUrl, and room name. The browser uses VoiceConversation to join LiveKit; Speko dispatches an agent worker into the same room to run the routed STT → LLM → TTS pipeline.

Two separate time controls apply to cascade sessions:

  • ttlSeconds (default 900s, max 86400s) bounds the join token only. It does not limit an already-connected session.
  • maxDurationSeconds (default 3600s, min 30s, max 14400s) is the hard cap on session lifetime, measured from session creation. Speko force-ends the LiveKit room at the deadline. A call hosted on a speech-to-speech model (runMode: "s2s", e.g. GPT-Live) is capped at 3600s regardless of the requested value.

The agent worker leaves when the room empties or the max-duration deadline hits. The session row is retained for usage and audit.

Ending a session

For a cascade session, POST /v1/calls/{sessionId}/end tears down the LiveKit room and runs the normal close and metering flow. It is idempotent; ending an already-ended session returns status: "already_ended". SDK: client.sessions.end(sessionId) or client.calls.end(sessionId).

For provider-direct S2S, call RealtimeVoiceConversation.endSession() (or close the low-level realtime handle) to close the provider connection. The credential and entitlement stay time-bounded even if the client disappears without a final telemetry event.

Phone sessions

POST /v1/sessions/phone creates the same kind of voice session, then dials a PSTN destination over LiveKit SIP. Inbound calls follow the reverse path: a registered phone number receives the carrier webhook, Speko creates the voice session, hydrates the linked agent or dispatch metadata template, and bridges the caller into the room.

Use phone sessions for:

  • Outbound appointment reminders, sales calls, and scheduled callbacks.
  • Inbound receptionists and support lines.
  • Calls that need carrier lifecycle events, forwarded-number metadata, post-call reports, recordings, or live transfers.

See Build a phone agent for the end-to-end phone flow.

Choosing

NeedUse
Transcribe a file, return text/v1/transcribe
Formatted transcript with diarization/v1/transcribe + workload: transcription (guide)
Generate audio for a notification/v1/synthesize
Long-form narration audio/v1/synthesize + workload: narration (guide)
Single LLM reply, no voice/v1/complete
Lowest-latency provider-native voice in a browser/v1/sessions with mode: "s2s" + RealtimeVoiceConversation
Routed STT → LLM → TTS voice agent in a browser/v1/sessions with mode: "cascade" + VoiceConversation
Outbound PSTN voice call/v1/sessions/phone or speko.voice.dial()
Inbound PSTN receptionist/v1/phone-numbers linked to an agent
Inspect reports, events, recordings, or transfers/v1/calls/{id}
Real-time voice in a self-hosted framework worker@spekoai/adapter-livekit directly, when using LiveKit

What sessions don't do

  • They aren't a chat history store. Cascade turn context lives in the worker; provider-direct context lives on the provider connection.
  • They don't proxy audio through the REST API. S2S audio goes directly to the provider; cascade audio uses the LiveKit media transport.
  • Provider-direct S2S reserves and settles a bounded duration entitlement. Cascade usage is recorded from its connected lifetime and underlying routed stages.

Recording

Cascade and phone sessions are recorded by default. The agent worker captures both speakers — caller and agent — into a single mixed-mono Opus file, persisted to Google Cloud Storage at the end of the call. There is no separate "enable recording" call for those modes.

Provider-direct S2S sessions are not recorded by Speko because their media never enters a Speko or LiveKit data path. Their recordingStatus is null, and GET /v1/sessions/{id}/recording has no artifact to return.

What gets captured:

  • Mixed mono Opus. Both sides of the conversation in one file, ~24 kbps. Stereo / per-speaker tracks are not produced — assume one combined audio stream per session.
  • The full call. Recording starts when the first participant joins the room and ends when the room empties (participants leave, the session is ended via the API, or the max-duration deadline hits).
  • Audio only. Tool-call payloads and transcripts are not in the audio file; those live on the session row and the per-turn entries.

How to fetch one:

curl -L \
  -H "Authorization: Bearer $SPEKO_API_KEY" \
  https://api.speko.dev/v1/sessions/$SESSION_ID/recording \
  --output session.opus

The endpoint 302-redirects to a short-lived (5 minute) signed GCS URL — pass -L so curl follows it. The signed URL is single-use within its TTL window; refetch the endpoint to get a fresh one rather than caching the URL itself.

The status field

Each session entry carries a recordingStatus that walks through:

  • pending — the call ended; the agent worker is assembling the file.
  • uploading — the file is being pushed to GCS.
  • ready — fully persisted and downloadable. recordingObjectPath and recordingDurationMs are populated.
  • failed — the upload errored. The recording is unrecoverable; the session itself is unaffected.
  • suppressed — the organization has recordingEnabled set to false, so no file was ever produced. This is the terminal state for opted-out orgs; it's not a transient one.
  • discarded — the call never connected (the session's endReason is sip_dial_rejected or sip_no_answer, or caller_abandoned when the inbound leg was still ringing), so the capture held only an empty room. The file is withheld and GET /v1/sessions/{id}/recording returns 404; there was no conversation to play back.

null on recordingStatus appears on provider-direct S2S sessions and legacy entries that predate recording. Treat it as "not available" and don't expect a download.

Retention

Recordings are kept for 30 days after the session ends, then deleted automatically by a GCS lifecycle policy. There is no in-product way to extend retention — if you need long-term storage, follow the redirect, download the file, and store it in your own bucket. If you need shorter retention, see the per-org opt-out below.

Per-org opt-out

Recording is governed at the organization level by organization.recordingEnabled (default true). Flipping it to false:

  • Stops new sessions from producing files. Their recordingStatus becomes suppressed and GET /v1/sessions/{id}/recording returns 404.
  • Does not retroactively delete prior recordings — those expire on the normal 30-day timer.

The flag is exposed in the dashboard under the Record voice sessions toggle on the Settings page. Per-call opt-out (skipping recording for a single session even when the org default is on) is out of scope for this round.

HIPAA mode

Organizations on a HIPAA-mode plan are forced into recordingEnabled: false regardless of the dashboard toggle, and the toggle is locked. This is the current bridge until an end-to-end customer-managed-key path lands; see the compliance issue tracker for the full story.

API endpoint

GET /v1/sessions/{id}/recording is the only supported way to retrieve a recording. It authenticates with the same bearer key as every other endpoint and 302-redirects to a 5-minute signed URL on success, or 404 when the session is unknown, the recording is not ready, failed, suppressed, or absent for provider-direct S2S. Inspect the parent session entry's recordingStatus to disambiguate. Never construct GCS URLs from recordingObjectPath directly — those URLs are not publicly addressable, and the field exists only as an internal handle.

Transcript

Every finalized STT and LLM turn during a cascade session is persisted server-side. The dashboard's session detail page surfaces them under a Transcript card; the API equivalent is GET /v1/sessions/{id}/transcript, which returns turns sorted by index.

The worker batches and debounces (~200ms) finalized turns and POSTs them to /v1/sessions/{id}/turns. The ingest endpoint is idempotent on (session_id, index) so retries on transient errors are safe.

Interim STT partials are not persisted — only finalized turns make it into the transcript.

Provider-direct S2S transcripts are available to the connected client through provider events, but Speko does not currently persist them. GET /v1/sessions/{id}/transcript therefore returns no S2S turns.

On this page