Speko Docs
Build

Build a phone agent

Inbound and outbound PSTN calls, phone-number routing, lifecycle webhooks, reports, and transfers.

Speko can run the same voice agent over browser sessions and PSTN calls. Phone calls use the platform-hosted LiveKit SIP path: Speko creates or receives the call leg, dispatches the agent worker, records the room, persists transcripts and events, and finalizes a call report after hangup.

Use this guide when you want a receptionist, appointment setter, callback workflow, or support line that can place calls, answer calls, and transfer callers to a human.

Phone surfaces

NeedAPISDK
Place an outbound phone callPOST /v1/sessions/phonespeko.voice.dial()
Buy or import numbers/v1/phone-numbers/*speko.phoneNumbers
Route inbound callsagentId or dispatchMetadataTemplate on a phone numberspeko.phoneNumbers.update()
Inspect calls and reports/v1/calls/{id}speko.calls.get()
Read lifecycle events/v1/calls/{id}/eventsspeko.calls.events()
Transfer live calls/v1/calls/{id}/transfers/*speko.calls.blindTransfer() / speko.calls.warmTransfer()

1. Create an agent

Phone calls work best with a persisted agent because the phone-number row can hydrate its prompt, routing intent, provider preferences, speech settings, and tools every time a call starts. Lifecycle webhooks are configured once at the workspace level and scoped to this agent when needed.

curl -X POST https://api.speko.dev/v1/agents \
  -H "Authorization: Bearer $SPEKO_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "name": "Front desk",
    "systemPrompt": "You are the front-desk receptionist. Greet callers, collect their name and reason for calling, answer common questions, and transfer urgent calls.",
    "intent": { "language": "en-US", "optimizeFor": "latency" }
  }'

Create lifecycle endpoints in Settings → Webhooks or through speko.webhooks. See Workspace webhooks for event selection, agent scopes, tagged routing, signing, retries, and logs. Agent-level webhooks fields are deprecated compatibility fields through the current SDK major version.

Realtime (speech-to-speech) agents on the phone

An agent saved with runMode: "s2s" (dashboard: Architecture → Realtime) answers and places phone calls with one speech-to-speech model instead of the STT → LLM → TTS cascade. Two models can do this, because Speko's voice worker hosts them and a SIP leg can ride what the worker holds:

PinModelVoicesBilled as
openai:gpt-live-1OpenAI GPT-LiveGPT-Live roster (marin, quartz, …)flat per-minute rate, billed per second (default backend)
google:gemini-3.8-liveGemini 3.8 LiveGemini roster (Puck, Kore, …)prompt + response tokens
google:gemini-3.8-live-extended-thinkingGemini 3.8 Live, reasoning tierGemini rosterprompt + response tokens

The other realtime models in the catalog (gpt-realtime-*, the older Gemini Live ids, Grok Voice) are provider-direct — the browser or SDK talks to the vendor itself — and cannot answer a phone.

  • Leave stackPreferences.allowedProviders.s2s empty or pin one of the rows above (a bare openai / google selects that provider's default). An empty allowlist still means GPT-Live, so an agent saved before Gemini was hostable keeps the model it has been running. The API refuses to link a phone number to a realtime agent pinned to a model a phone cannot host (AGENT_S2S_PIN_NOT_PHONE_HOSTABLE), and refuses the agent update that would create that state (S2S_PIN_NOT_PHONE_HOSTABLE).
  • Gemini needs a Google AI Studio key in the workspace (the google-aistudio provider slot, or google when that holds an AI Studio key rather than a Vertex deployment). Tools, end_call and transcripts work the same on both models; GPT-Live reaches them through its backend model, Gemini runs them itself.
  • The session row carries pipelineKind: "s2s". GPT-Live on its default backend model bills the same flat published per-minute rate as a cascade call, with the voice model and backend tokens included. Choosing a different backend (s2s.backendModel) bills the model usage instead — GPT-Live connected seconds plus backend tokens — and so do the Gemini models (prompt + response tokens). A call on your own OpenAI key (BYOK) pays no Speko minute.
  • If a call still cannot host the route at dial time — no credential for the pinned model's provider in the workspace, or the worker is unavailable — the call runs the cascade and says so: an s2s.fallback event on the session timeline, metadata.s2sFallback on the session, and a warning in the logs. It is never silent.
  • A single outbound call can ask for this without an agent: POST /v1/sessions/phone accepts runMode: "s2s" next to intent / systemPrompt, and the leg is hosted on GPT-Live (the default when nothing is pinned). The per-call value also wins over an agent's saved run mode — runMode: "cascade" on a Realtime agent runs the cascade for that call, with no fallback recorded because it was asked for. The same dial-time fallbacks apply.
  • Outbound GPT-Live calls run worker-side answering-machine detection, rolled out per workspace. Until it is turned on for yours, the detector only observes: each verdict is recorded as a worker.duplex.amd_shadow_verdict event and the call carries on as if nothing had been detected. Once on, turnHandling.onMachine applies, with the differences listed under Voicemail handling. Gemini Live calls have no worker-side detection (a call that asks for it by name gets a worker.duplex.amd_unsupported event); carrier AMD still applies to both models.
  • Keypad navigation works on both models: every realtime phone leg can press digits with send_dtmf unless turnHandling.keypad is false, and turnHandling.profile: "ivr" navigates menus from the first turn. Idle check-ins run on phone legs, using the agent's idleRePrompts when it has them — spoken in the model's own words.
  • Not yet on the realtime phone path: warm transfers, mid-call language switching (the agent answers in its configured language), and data capture. Use a cascade agent when a call needs those.

Voice level on phone calls

TTS voices are mastered at very different levels. Some synthesize speech at -12 to -15 dBFS with peaks near full scale, well above the roughly -20 dBFS of ordinary telephone speech, and callees hear them as loud or harsh on a mobile handset. audioOutput.gainDb sets a static gain on the agent's own voice, in dB from -24 to 12. It applies to every call the agent takes, whichever TTS provider speaks, and on phone legs it is applied before the audio reaches the SIP trunk (dashboard: Advanced → Speech & audio → Voice volume (dB)):

curl -X PATCH https://api.speko.dev/v1/agents/$AGENT_ID \
  -H "Authorization: Bearer $SPEKO_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{ "audioOutput": { "gainDb": -6 } }'
  • Negative values make the agent quieter: -6 roughly halves the amplitude. Attenuation never distorts.
  • Positive values boost a voice that is mastered quiet. The boost saturates at full scale, so it clips a voice that already peaks near 0 dBFS.
  • 0 (or null) plays the voice at its own level. Background audio and tool sounds keep their own volume.
  • It also applies on realtime phone legs (GPT-Live, Gemini Live). It does not apply to provider-direct realtime sessions, where the browser talks to the vendor itself.
  • A pre-call webhook can override it per call (audioOutput in the returned config), and POST /v1/sessions accepts it per session. Either replaces the agent's value as a whole.

2. Add a phone number

You can use a platform-managed number or register a customer-owned SIP trunk number.

Managed number search and purchase is currently the US-number path. It requires phone-number business verification and sufficient credits:

const options = await speko.phoneNumbers.searchAvailable({ areaCode: '415' });

const number = await speko.phoneNumbers.create({
  e164: options[0]!.e164,
  direction: 'both',
  agentId: 'agent_123',
  label: 'Main line',
});

For numbers you already own, import the SIP trunk instead:

await speko.phoneNumbers.importSipTrunk({
  e164: '+442071234567',
  sipConnectionInstallationId: '00000000-0000-4000-8000-000000000010',
  direction: 'both',
  agentId: 'agent_123',
  label: 'London front desk',
});

direction controls whether the number can be used for inbound calls, outbound calls, or both. Inbound calls require either an agentId or a dispatchMetadataTemplate.

3. Place outbound calls

Use POST /v1/sessions/phone or speko.voice.dial() for outbound PSTN calls. The response returns the Speko session id, room name, resolved caller ID, and SIP participant handle.

const call = await speko.voice.dial({
  to: '+12015551234',
  from: '+12015550199',
  agentId: 'agent_123',
  firstMessage: 'Hi, this is Ava from Acme. Is now still a good time?',
  telephony: {
    region: 'us-east',
    amd: { mode: 'agent', timeoutSeconds: 8 },
  },
  webhookTags: {
    environment: 'production',
  },
  metadata: {
    campaignId: 'renewal-q2',
    leadId: 'lead_456',
  },
});

console.log(call.sessionId, call.status);

agentId is required unless you pass an ad hoc intent. Per-call fields such as systemPrompt, firstMessage, voice, llm, ttsOptions, sttOptions, constraints, and metadata override or extend the agent defaults for that call.

Pass timezone (an IANA name such as America/Chicago) to set the call's local time. The agent schedules from it, so "Thursday at 1PM" resolves to the right date. It overrides the agent's saved timezone for that call only. POST /v1/sessions accepts the same field.

How outbound calls are billed

Cascade calls, and GPT-Live on its default backend, bill the flat per-minute rate, charged per second; the other realtime models bill model usage instead (see Realtime agents on the phone). For flat-rate calls, the clock starts when the call's session starts, which is before the number is dialed, so ringing time counts:

OutcomeBilled
Answered (by a person or voicemail)Ringing plus the conversation
Rang, nobody answered (sip_no_answer)The ringing time
Refused by the carrier, the phone never rang (sip_dial_rejected)Nothing
Session closed failedNothing

Set maxDurationSeconds to cap how long a single call can run.

Voicemail handling

Outbound calls use native answering-machine detection by default (a Realtime agent's calls differ — see Realtime (GPT-Live) calls below). With greetFirst: true (the default), detection runs while the agent can speak; greetFirst: false briefly holds the opening for a verdict. telephony.amd.mode: 'disabled' disables native detection and transcript-based machine cues. mode: 'carrier' also disables native detection and requires carrier/trunk support; setting that mode alone does not supply carrier verdicts to the worker. A voicemail policy does not enable detection when it is disabled or carrier-managed.

When native detection identifies recordable voicemail, turnHandling.onMachine controls the response. Set it on the agent to inherit it across calls, or override individual fields per call. For example, passing only { profile: 'ivr' } preserves the agent's onMachine and voicemailMessage settings:

const call = await speko.voice.dial({
  to: '+12015551234',
  agentId: 'agent_123',
  turnHandling: {
    onMachine: 'leave_message',
    voicemailMessage:
      "Hi, this is Ava from Acme about your renewal. I'll try again tomorrow, or call us back at 201-555-0199.",
  },
});
onMachineWhat the agent does
agent_decides (default)The verdict is handed to the model as a system note ("this is a voicemail greeting"). Your prompt's own rules apply — for example "on voicemail, call end_call immediately".
hangupEnds the answered call when voicemail is detected, without a farewell or voicemail message. An opening may already have played with greetFirst: true.
leave_messageWaits for a 1.5-second quiet window, then speaks voicemailMessage once in the agent's voice and hangs up after playout. voicemailMessage renders the same {{variables}} as firstMessage, and can be authored per language — see Multilingual agents. If the greeting stays active for 20 seconds, it ends without speaking the message. Quiet detection is a fallback, not beep detection.

An unavailable mailbox (for example, full or not set up) always ends without attempting a message, regardless of onMachine. It is recorded as endReason: "machine_unavailable" with report outcome unavailable.

IVR and call-screening detections continue through the navigation flow. In IVR mode the agent can follow spoken prompts and send keypad digits. Machine silence after exhausted prompts and a separate 120-second machine-phase budget bound calls that make no progress. IVR timeouts are not reported as voicemail.

Calls ended by a voicemail policy use endReason: "voicemail_reached" and report outcome voicemail. This means a mailbox was reached, not that a message was saved: worker.voicemail_delivery separately records local playout success or failure, a greeting timeout, or a missing message. The outcome holds when the mailbox hangs up before or during your message: the call still ends as voicemail_reached, and the delivery result is call_ended.

The call report and the call.report webhook carry the same detail in a voicemail object:

{
  "outcome": "voicemail",
  "voicemail": {
    "detected": true,
    "amd_verdict": "machine-vm",
    "delivery": "delivered"
  }
}
FieldMeaning
detectedA recordable mailbox was reached.
amd_verdictThe latest detection verdict for the line: human, machine-vm, machine-ivr, machine-unavailable or unknown. null when detection did not run.
deliveryThe leave_message result: delivered, missing_message, greeting_timeout, playout_failed, call_ended (the far end hung up while the message played), or not_attempted (the call ended before the message started). null when no message was owed.

When the mailbox hangs up, the report can be finalized before the delivery result arrives. If the result lands while the report is still being built, the first call.report already carries it. If it lands after that call.report was sent, Speko updates the report and sends call.report once more with the corrected voicemail block (same session, new delivery id); the summary and analysis are not re-run. Treat the latest call.report for a session as authoritative. Reports finalized before this field existed carry voicemail: null.

Realtime (GPT-Live) calls

On an outbound GPT-Live call, detection reads the voice model's own transcript of the far end rather than a separate speech-to-text stage, and it is rolled out per workspace. Until it is enabled for yours it only observes: verdicts are recorded as worker.duplex.amd_shadow_verdict events, none of the actions above run, and the call's outcome is decided exactly as it would be without detection. Once enabled, onMachine works as described above, with these differences:

  • A voicemail or unavailable verdict from the classifier alone is acted on only when the far end spoke before the agent did and said at least five words that do not read like a live reply; otherwise it is recorded as worker.amd_action_gated and the call continues. An explicit recorded-mailbox announcement in the far end's speech (for example "leave a message after the tone" or "the mailbox is full") is acted on without this check, at any point in the call, unless it reads like a person talking ("I'll take a message", quoting a greeting).
  • leave_message: GPT-Live cannot read a script, so the model is told to stay silent through the greeting and then to leave voicemailMessage, which it speaks in its own words — close to what you wrote, not verbatim. A long message is shortened first, at a sentence boundary where one exists, to fit the model's instruction limit: about 1,200 characters of English, and fewer in scripts that take more tokens per character — about 420 characters of Chinese, Japanese or Korean. The realtime path does not switch languages mid-call, so only the voicemailMessage authored for the agent's configured language applies.
  • greetFirst: false is not honoured: the opening is never held for a verdict.
  • A detected voicemail left to the model (agent_decides) ends as voicemail_reached when the model ends the call, and at the latest once the mailbox goes silent or the 120-second budget runs out. IVR navigation has no machine-phase budget on a realtime call; the line guards and the one-hour realtime ceiling bound it instead.
  • Gemini Live calls have no worker-side detection. Carrier AMD (telephony.amd.mode: 'carrier') still applies to both models.

4. Receive inbound calls

When a call arrives on a registered number, Speko matches the dialed E.164 number, checks direction, creates a voice session, hydrates the linked agent, dispatches the worker, and bridges the caller into the LiveKit room.

If the caller hangs up while the call is still ringing, before the agent says anything, the session closes with status: "ended" and endReason: "caller_abandoned", and records a sip.caller_abandoned call event. It isn't counted as a failure. When LiveKit reports the leg as still ringing, the capture of the empty room is discarded (artifacts.recording.status: "discarded"); otherwise the recording is kept and served as usual.

For forwarded calls, Speko attempts to preserve the original forwarding source. It looks at Telnyx payload fields such as forwarded_from, original_to, and redirecting_number, plus SIP headers such as Diversion and History-Info. The normalized value is exposed as:

  • forwardedFromNumber in session metadata.
  • forwarded_from_number in pre-call webhook payloads.
  • call events and reports through the session metadata.

Use dispatchMetadataTemplate when you need static metadata or a legacy template alongside the linked agent:

await speko.phoneNumbers.update('pn_123', {
  agentId: 'agent_123',
  dispatchMetadataTemplate: {
    tenant: 'acme',
    line: 'front_desk',
    caller: '{{callerNumber}}',
    dialedNumber: '{{dialedNumber}}',
    forwardedFromNumber: '{{forwardedFromNumber}}',
  },
});

The agent fields are canonical. If agentId and dispatchMetadataTemplate set the same pipeline key, the persisted agent wins and the template fills the gaps.

Template keys split by name: keys from the pipeline vocabulary (systemPrompt, firstMessage, voice, intent, and the other agent-level fields) hydrate the call's pipeline config, while every other key (like tenant or line above) is static metadata — it surfaces on the session's metadata, in call events and reports, and in webhook payloads, but never changes call routing. A handful of server-internal key names (native, proxy, tools, stages, and similar) are reserved and rejected when you save the template.

5. Customize calls before answer

Configure a workspace endpoint subscribed to call.pre_call and scoped to the agent when your application needs to look up the caller, choose a greeting, inject account context, or override the call pipeline before the worker starts speaking. Only one effective pre-call endpoint may match an agent/tag route.

Speko sends:

{
  "type": "call.pre_call",
  "call_id": "session_123",
  "session_id": "session_123",
  "organization_id": "org_123",
  "direction": "inbound",
  "to": "+12015550199",
  "from": "+12015551234",
  "dialed_number": "+12015550199",
  "forwarded_from_number": null,
  "phone_number_id": "pn_123",
  "call_control_id": "sip_participant_123",
  "pipeline_config": {},
  "metadata": {}
}

Return any supported pipeline overrides at the top level or under pipelineConfig / pipeline_config:

{
  "firstMessage": "Hi Maya, thanks for calling Acme.",
  "systemPrompt": "The caller is Maya Chen, a priority customer. Keep responses brief.",
  "metadata": {
    "customerId": "cus_123",
    "plan": "enterprise"
  }
}

Supported pre-call override keys include intent, constraints, voice, systemPrompt, firstMessage, llm, ttsOptions, sttOptions, backgroundAudio, audioOutput, tools, pronunciation and text replacement rules, and idle reprompts.

The response may also carry a top-level toolSecrets object — per-call values for tools only (a short-lived access token scoped to the caller you just looked up, a tenant API base URL). Unlike variables, these never enter the prompt, the transcript, or the worker's dispatch metadata; they are stored encrypted and released only to tool execution (session.secrets.<name> in custom-code tools, {{name}} in webhook tool url/headers). Same limits and semantics as the toolSecrets request field on /v1/sessions — see Per-call secrets. On outbound calls, values sent on the dial request win over the webhook's on a name collision. An invalid toolSecrets object is dropped with a warning (names logged, never values) and the call proceeds.

{
  "firstMessage": "Hi Priya, thanks for calling back.",
  "variables": { "customer_name": "Priya" },
  "toolSecrets": { "access_token": "eyJhbGciOi...", "base_url": "https://acme.example.com" }
}

6. Track lifecycle and reports

Speko records durable lifecycle events for LiveKit, Telnyx, Speko status updates, SIP cause codes, and transfer attempts.

const call = await speko.calls.get('session_123');
const { events } = await speko.calls.events('session_123');
const report = await speko.calls.report('session_123');

Endpoints subscribed to call.status receive call progress and failure information:

{
  "type": "call.status",
  "call_id": "session_123",
  "event_type": "call.hangup",
  "provider": "telnyx",
  "status": "ended",
  "failure_cause": null,
  "sip_status_code": 200,
  "sip_status": "OK",
  "metadata": {}
}

After hangup, Speko finalizes a report with transcript entries, summary, outcome, structured data, cost breakdown, artifacts, metadata, and any scheduled callback created by analysis. Endpoints subscribed to call.report receive the same report payload. Report, analysis, and recording deliveries retry independently; inspect or manually redeliver them through speko.webhooks.deliveries.

recording_url is a presigned download URL you can fetch directly with no auth headers (a GET on the URL streams the audio). It's a signed link valid for up to 7 days from delivery, so download or copy the file to your own storage rather than persisting the URL long-term. To fetch a recording on demand later, call the authenticated GET /v1/calls/{id}/recording endpoint with your API key instead.

A recording usually finishes uploading after the call ends, so the first call.report delivery often arrives with recording_url: null (and artifacts.recording.status: "uploading"). Once the recording finalizes, Speko re-fires call.report for the same call_id with recording_url populated and artifacts.recording.status: "ready". Deduplicate on call_id and treat the later delivery as authoritative for the recording. recording_url stays null when recording is disabled for the org (artifacts.recording.status: "suppressed"), when the upload failed, or when the dial never connected (artifacts.recording.status: "discarded" — a refused or unanswered outbound leg, or an inbound caller who hung up while it was still ringing, captures only an empty room, so no recording is served).

Dedicated analysis and recording webhooks

If handling the call.report re-fire is awkward for your consumer (for example, your handler dedupes on call_id and drops the delivery that carries the recording), subscribe to the dedicated single-purpose webhooks instead. Each fires once per call — and unlike the call.report re-fire, any webhook-standard duplicate (e.g. a redelivery after a timeout) carries identical content, so deduplicating on type + call_id never loses data:

  • call.analysis fires when the LLM analysis completes: summary, outcome, structured_data, custom_data, and an analysis object (status, provider, model, completed_at). No transcript, cost, or recording — and no recording re-fire to dedupe.
  • call.recording fires when the recording reaches a terminal state. recording_url carries the presigned link (same 7-day-TTL semantics as above) when recording.status is "ready", and null when it's "failed" — so your consumer can stop waiting either way. Suppressed recordings (recording disabled for the org) never fire this event.
{
  "type": "call.recording",
  "call_id": "b2f7…",
  "session_id": "b2f7…",
  "organization_id": "org_…",
  "agent_id": "agent_…",
  "recording_url": "https://storage.googleapis.com/…",
  "recording": { "status": "ready", "duration_ms": 21625 },
  "created_at": "2026-07-09T15:15:19.131Z"
}

Subscribe one endpoint to any combination of these events:

await speko.webhooks.create({
  name: 'Call processing',
  url: 'https://example.com/speko/calls',
  events: ['call.report', 'call.analysis', 'call.recording'],
  allAgents: true,
});

Both events are fully additive: call.report keeps its combined payload and its recording re-fire, so existing postCall consumers are unaffected. Failed deliveries retry with the same backoff as call.report.

Custom data extraction

Add extractionFields to a workspace endpoint subscribed to call.report or call.analysis to have Speko pull structured values out of every call. Each field has a name, a type (string, number, boolean, or enum), and a description that tells the analysis model what to extract — enum fields also take an options list. After the call, the model fills each field from the transcript, coerced to its declared type (or null when the call gives no basis for it), and the values arrive as a top-level custom_data object keyed by field name:

{
  "type": "call.report",
  "summary": "Jane booked a table for four.",
  "outcome": "scheduled",
  "structured_data": { "...": "..." },
  "custom_data": {
    "caller_intent": "booking",
    "party_size": 4,
    "wants_callback": true
  }
}

custom_data always carries one key per configured field, so the payload schema stays stable. Manage these fields from Settings → Webhooks in the dashboard.

Limits: up to 30 fields per webhook, each description up to 10,000 characters, with at most 40,000 characters of descriptions combined across all fields. Detailed rubric-style descriptions are encouraged for classification fields — the analysis model follows them closely.

7. Transfer callers

Use the Calls API for operator-driven transfers, or add the built-in transfer_call tool to let the agent transfer based on conversation context.

await speko.calls.blindTransfer('session_123', {
  to: '+12015554321',
  ringingTimeout: 25,
  metadata: { reason: 'billing escalation' },
});

const transfer = await speko.calls.warmTransfer('session_123', {
  from: '+12015550199',
  destinations: [
    { to: '+12015551234', label: 'Front desk' },
    { to: '+12015554321', label: 'Overflow' },
  ],
  screeningPrompt: 'Confirm the recipient can help before bridging.',
  fallback: { strategy: 'take_message' },
  voicemailDetection: { mode: 'agent', timeoutSeconds: 10 },
});

Warm transfer starts a consultation leg first. Complete the transfer when the screened recipient accepts, or cancel with tryNext: true to continue through the destination list.

Next

On this page