Provider-direct speech-to-speech
Mint a short-lived OpenAI, Gemini Live, or xAI session and connect the browser directly to the provider.
Provider-direct speech-to-speech (S2S) keeps Speko in the control and billing path without putting Speko in the media path:
- Your backend calls
POST /v1/sessionswithmode: 's2s'and its Speko API key. - Speko reserves a bounded entitlement and returns a scoped, short-lived provider credential.
- The browser connects directly to OpenAI Realtime, Gemini Live, or xAI Grok Voice.
- Speko stores content-free lifecycle and billing evidence. Audio never traverses a Speko media proxy or LiveKit room.
Use this path for the lowest possible browser-to-model latency. If you need Speko's routed STT → LLM → TTS pipeline, recordings, or server-persisted transcripts, use a VoiceConversation cascade session instead.
Install
npm install @spekoai/client1. Mint an S2S session on your backend
Keep SPEKO_API_KEY on your server. Return the S2S response to the browser unchanged; it contains a scoped provider credential and Speko control URLs, never a provider root key.
import crypto from 'node:crypto';
app.post('/api/realtime-session', async (req, res) => {
const idempotencyKey = req.get('Idempotency-Key') || crypto.randomUUID();
const response = await fetch('https://api.speko.dev/v1/sessions', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.SPEKO_API_KEY}`,
'Content-Type': 'application/json',
'Idempotency-Key': idempotencyKey,
},
body: JSON.stringify({
mode: 's2s',
s2s: {
provider: 'google',
model: 'gemini-3.1-flash-live-preview',
voice: 'Puck',
systemPrompt: 'You are a concise voice assistant.',
},
ttlSeconds: 900,
}),
});
const body = await response.json();
res.status(response.status).json(body);
});Idempotency-Key is required for S2S. Reuse the same value only when retrying an ambiguous bootstrap timeout; Speko will return the same reservation instead of minting a second one.
2. Connect directly from the browser
import { useEffect, useRef, useState } from 'react';
import { RealtimeVoiceConversation } from '@spekoai/client';
export function RealtimePanel() {
const conversationRef = useRef<RealtimeVoiceConversation | null>(null);
const [status, setStatus] = useState('idle');
const [transcript, setTranscript] = useState<string[]>([]);
async function start() {
setStatus('connecting');
const response = await fetch('/api/realtime-session', { method: 'POST' });
if (!response.ok) throw new Error('Could not create an S2S session');
const session = await response.json();
const conversation = await RealtimeVoiceConversation.create({
...session,
onStatusChange: setStatus,
onMessage: ({ source, text, isFinal }) => {
if (isFinal) setTranscript((items) => [...items, `${source}: ${text}`]);
},
onError: (error) => console.error(error),
onDisconnect: () => setStatus('idle'),
});
conversationRef.current = conversation;
}
async function stop() {
await conversationRef.current?.endSession();
conversationRef.current = null;
}
useEffect(() => () => { void conversationRef.current?.endSession(); }, []);
return (
<div>
<button onClick={start} disabled={status !== 'idle'}>Start</button>
<button onClick={stop} disabled={status === 'idle'}>Stop</button>
<p>Status: {status}</p>
<ul>{transcript.map((item, index) => <li key={index}>{item}</li>)}</ul>
</div>
);
}RealtimeVoiceConversation acquires the microphone, performs the provider handshake, plays response audio, normalizes transcript events, reports content-free telemetry, and tears down every media resource in endSession().
Providers and defaults
| Provider | Default model | Default voice | Direct transport | PCM input / output |
|---|---|---|---|---|
| OpenAI | gpt-realtime-2.1 | marin | WebRTC | 24 kHz / 24 kHz |
gemini-3.1-flash-live-preview | Puck | Gemini Live WebSocket | 16 kHz / 24 kHz | |
| xAI | grok-voice-latest | eve | Grok Voice WebSocket | 24 kHz / 24 kHz |
Pin s2s.model to select another model from the same provider. Google also serves gemini-3.8-live (its current Live generation) and gemini-3.8-live-extended-thinking (the same model with background reasoning, which answers later) on the same socket, sample rates and voices as the default above.
Thinking and reasoning
Two models take a thinking dial, and each one is strict about it:
| Field | Applies to | Values | Default |
|---|---|---|---|
s2s.thinkingLevel | google:gemini-3.8-live-extended-thinking | low, medium, high | low |
s2s.backendModel | openai:gpt-live-1 | an OpenAI LLM catalog id | gpt-5.6-luna |
s2s.backendReasoningEffort | openai:gpt-live-1 | minimal, low, medium, high | the backend model's own default |
thinkingLevel is a property of the model, not a preference. The extended-thinking model requires one — Speko defaults it to low so a session that names none still connects — and every other Gemini Live model refuses one, so sending it there is a 400. Google rejects minimal on the extended-thinking model, which is why it is not offered.
GPT-Live's voice model does no reasoning itself: it delegates to a backend Responses model, so backendModel and backendReasoningEffort are the only knobs that change how hard it thinks. Higher effort buys better answers and costs the caller silence while it works.
Both fields ride the s2s block of POST /v1/sessions, the same request the server example above makes:
{
"mode": "s2s",
"s2s": {
"provider": "google",
"model": "gemini-3.8-live-extended-thinking",
"thinkingLevel": "high"
}
}{
"mode": "s2s",
"s2s": {
"provider": "openai",
"model": "gpt-live-1",
"backendModel": "gpt-5.6-terra",
"backendReasoningEffort": "medium"
}
}The TypeScript SDK takes the Gemini option as a realtime.connect parameter:
const session = await speko.realtime.connect({
provider: 'google',
model: 'gemini-3.8-live-extended-thinking',
thinkingLevel: 'high',
});GPT-Live's two options have no realtime.connect equivalent: that model is hosted on Speko's LiveKit worker, and connect returns a provider-direct handle. Send them on POST /v1/sessions as above and join the returned room with VoiceConversation.
You may omit s2s.provider and s2s.model to let Speko choose an eligible S2S route. If you set one, set both. Use constraints.allowedProviders.s2s to limit routing to provider names or exact provider:model pairs.
Duration and billing
For S2S, ttlSeconds is the requested maximum session horizon, not a LiveKit join-token TTL. The API defaults to 1,800 seconds and accepts up to 3,600 seconds. Provider credential limits can make the first authorized slice shorter:
- OpenAI uses a provider-authenticated sideband to bind the WebRTC call, collect provider events, and enforce the authorized deadline without receiving media.
- Gemini Live uses renewable entitlement slices.
RealtimeVoiceConversationreserves the next slice and resumes the Gemini session automatically while it remains inside the requested horizon. - Browser telemetry is operational evidence and may update the dashboard estimate, but it is not trusted to reduce the final charge. When exact provider-side evidence is unavailable, Speko settles the bounded entitlement that was authorized before the credential was minted.
The create response reports reservation.authorizedDurationSeconds, reservation.billing.maximumAmountMicros, and any reservation.billing.renewalUrl. Treat those returned values—not only the requested ttlSeconds—as the active authorization.
Current artifact behavior
- S2S transcript callbacks are available in the browser, but Speko does not yet persist them to
GET /v1/sessions/{id}/transcript. - Speko does not record provider-direct S2S media.
recordingStatusisnull, and the recording endpoint has no artifact for these sessions. endSession()closes the direct provider connection locally. The provider credential and server-side entitlement remain bounded even if a modified client omits the final telemetry event.