connect_realtime
Provider-direct speech-to-speech sessions, async only.
Open a speech-to-speech (S2S) session. Speko reserves the entitlement and mints a short-lived provider credential; the SDK then connects straight to OpenAI Realtime, Gemini Live, or xAI Grok Voice. Audio never traverses a Speko media proxy.
Realtime sessions are only available on AsyncSpeko. OpenAI uses an aiortc WebRTC connection; Gemini Live and xAI use asynchronous provider WebSockets.
import asyncio
from spekoai import AsyncSpeko, RealtimeConnectParams
async def main():
async with AsyncSpeko(api_key=os.environ["SPEKO_API_KEY"]) as speko:
session = await speko.connect_realtime(
RealtimeConnectParams(provider="openai", model="gpt-realtime-2.1"),
)
async with session:
await session.send_audio(pcm_chunk)
async for frame in session:
if frame["type"] == "audio":
play(frame["pcm"])
elif frame["type"] == "transcript":
print(frame["text"])
asyncio.run(main())Signature
await AsyncSpeko.connect_realtime(
params: RealtimeConnectParams,
) -> AsyncRealtimeSessionRealtimeConnectParams
provider'openai' | 'google' | 'xai'requiredmodelstringrequiredProvider-specific model id, such as gpt-realtime-2.1, gemini-3.1-flash-live-preview, or grok-voice-latest.
voicestringVoice id override — interpreted per provider.
system_promptstringtemperaturefloatinput_sample_rate16000 | 24000output_sample_rate16000 | 24000toolslist[RealtimeToolSpec]Tool definitions the assistant may call. Receive tool_call frames and respond with send_tool_result.
metadatadict[str, object]Free-form metadata attached to the session record.
ttl_secondsintRequested maximum session duration. The returned entitlement and provider limits may impose a lower cap.
idempotency_keystringStable bootstrap key to reuse after an ambiguous timeout. The SDK generates one when omitted.
AsyncRealtimeSession
The returned session is both an async context manager and an async iterator.
Properties
| Property | Type | Description |
|---|---|---|
session_id | str | Server-assigned session identifier. |
expires_at | str | ISO-8601 expiry for the delegated credential. |
input_sample_rate | int | Provider input PCM sample rate. |
output_sample_rate | int | Provider output PCM sample rate. |
Methods
send_audio(pcm: bytes)coroutineSend a PCM16 chunk directly to the selected provider.
commit()coroutineSignal an end-of-user-turn to the provider.
interrupt()coroutineCancel the assistant's current response mid-generation.
send_tool_result(call_id: str, output: str)coroutineReturn the result of a previously-issued tool call.
close(code=1000, reason='client_closed')coroutineClose the socket. Safe to call multiple times; the context manager calls it for you on exit.
Frame types
Iterating the session yields dicts tagged by type:
Frame type | Payload fields |
|---|---|
audio | pcm: bytes, sample_rate: 24000 |
transcript | role: 'user' | 'assistant', text: str, final: bool |
tool_call | call_id: str, name: str, arguments: str (JSON) |
usage | input_audio_tokens: int, output_audio_tokens: int |
error | code: str, message: str |
close | reason: str |
Example — tool calling
import json
from spekoai import AsyncSpeko, RealtimeConnectParams, RealtimeToolSpec
async with AsyncSpeko(api_key=os.environ["SPEKO_API_KEY"]) as speko:
session = await speko.connect_realtime(
RealtimeConnectParams(
provider="openai",
model="gpt-realtime-2.1",
tools=[
RealtimeToolSpec(
name="get_weather",
description="Current weather for a city.",
parameters={
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
),
],
),
)
async with session:
async for frame in session:
if frame["type"] == "tool_call" and frame["name"] == "get_weather":
args = json.loads(frame["arguments"])
result = fetch_weather(args["city"])
await session.send_tool_result(frame["call_id"], result)