Speko Docs

VoiceConversation

Primary API — construct, control, and tear down a voice session.

VoiceConversation is the public class exported from @spekoai/client. Always construct it via the static create() factory — the factory awaits connection, so by the time it resolves the session is live.

import { VoiceConversation } from '@spekoai/client';

There is also a legacy namespace export, Conversation, with a single method Conversation.startSession(options) — it's an alias for VoiceConversation.create(options), kept so consumers migrating from other SDKs can use familiar naming.

VoiceConversation.create(options)

static create(options: CreateOptions): Promise<VoiceConversation>

Where CreateOptions is the short-lived token shape:

type CreateOptions = ConversationOptions;

Your backend calls POST /v1/sessions with mode: 'cascade', optionally with an agentId, and forwards only transportToken and transportUrl to the browser. VoiceConversation.create() connects to the LiveKit media transport, publishes the microphone track, sends any overrides over the data channel, fires onConnect, and resolves. It throws a SpekoClientError on connection, network, or microphone failure.

That same POST /v1/sessions call is where per-session behaviour belongs. See per-session variables for prompt values, tool credentials, and webhook routing scoped to one browser session.

Do not send SPEKO_API_KEY to browser code. VoiceConversation no longer accepts agentId, apiKey, or apiBaseUrl; session minting belongs on your server.

ConversationOptions

FieldTypeRequiredDescription
transportTokenstring✅Browser-safe media transport token, returned by your server.
transportUrlstring✅Media transport URL, returned by your server. Pass it straight through — the SDK does not default this so consumers can't ship against the wrong environment.
overridesConversationOverrides?Accepted and published, but no agent reads it — see ConversationOverrides. Use per-session variables instead.
inputDeviceIdstring?Specific microphone deviceId.
outputDeviceIdstring?Specific speaker deviceId. Applied via setSinkId; silently ignored on browsers without support.
audioConstraintsAudioConstraints?echoCancellation, noiseSuppression, autoGainControl. All default true.
onConnect(d: { conversationId }) => voidFired after the mic publishes and status becomes connected.
onDisconnect(d: DisconnectionDetails) => voidFired on server or client disconnect.
onMessage(m: ConversationMessage) => voidInbound transcripts, agent messages, user-message echoes.
onStatusChange(s: ConversationStatus) => voidconnecting → connected → disconnecting → disconnected.
onModeChange(m: ConversationMode) => voidlistening vs speaking, derived from transport active-speaker events.
onError(err: Error) => voidNon-fatal errors (malformed data packets, media device errors, sink-id failures).

conversationToken and livekitUrl are still accepted as legacy aliases for existing callers.

ConversationOverrides

interface ConversationOverrides {
  agent?: {
    prompt?: string;
    firstMessage?: string;
    language?: string;
  };
  tts?: {
    voiceId?: string;
    speed?: number;
  };
}

Overrides have no effect. The SDK serializes them and publishes them over the data channel; no agent worker reads them. Passing overrides changes nothing about the agent's behaviour.

The packet goes out with no data-channel topic, and the worker's one data handler accepts the speko.control topic alone, so it is discarded on arrival. On that topic the handler dispatches two message types, neither of them overrides.

Set prompt, first message, and language with per-session variables instead. Your server sets those when it mints the session, so page JavaScript cannot forge them.

The type and its wire format remain documented; the API is retained pending a decision on browser-side overrides.

AudioConstraints

interface AudioConstraints {
  echoCancellation?: boolean; // default true
  noiseSuppression?: boolean; // default true
  autoGainControl?: boolean; // default true
}

The SDK always routes through createLocalAudioTrack({ ... }) rather than setMicrophoneEnabled(true) so that constraints are applied even when no inputDeviceId is passed — setMicrophoneEnabled silently ignores them in that case.

Instance methods

getId(): string

Returns the transport conversation id. Populated after create() resolves.

isOpen(): boolean

true while the underlying status is connected.

setMicMuted(muted: boolean): Promise<void>

Mute / unmute the local microphone track. Uses the track-level mute API when a track is attached; falls back to LocalParticipant.setMicrophoneEnabled() otherwise.

setVolume(volume: number): void

Set playback volume for every remote audio element (0–1, clamped). Applied immediately to existing elements and to future ones.

sendChatMessage(text: string): Promise<void>

Send typed text to the agent as a user turn. This is the outbound text path that works — prefer it over sendUserMessage.

It is not a data packet. The SDK calls the transport's native sendText on the lk.chat topic, which LiveKit Agents' RoomIO consumes as a user turn, so the agent replies with audio and transcription as it would to speech.

Delivery depends on the framework's text_enabled default. Speko's worker does not set that flag in either direction, so this path rests on a LiveKit default rather than on a platform guarantee.

sendUserMessage(text: string): void

Publish a user_message packet over the reliable data channel.

No agent reads this packet — it is dropped on the same topic mismatch described under ConversationOverrides. Use sendChatMessage for typed user input.

sendContextualUpdate(text: string): void

Publish a contextual_update packet, intended for out-of-band context such as "user switched to the checkout page".

No agent reads this packet either, for the same reason. To give the agent per-session context that does take effect, pass it as a per-session variable when your server mints the session. Context that has to arrive mid-call has no working browser path today.

endSession(): Promise<void>

Initiate clean disconnection. Sets status to disconnecting, asks the transport to disconnect; the disconnect event completes the teardown (stops the mic track, removes audio elements, fires onDisconnect). Idempotent — calling it twice is a no-op.

Teardown invariants

When disconnection completes (whether triggered by endSession(), agent leaving, token expiry, or error), the SDK:

  1. Sets status to disconnected and fires onStatusChange.
  2. Stops the local microphone track so the browser's mic indicator goes away.
  3. Detaches and removes every <audio> element it added to document.body.
  4. Fires onDisconnect with a mapped DisconnectionReason.

Your component's unmount effect should call endSession() so navigating away doesn't leak a live transport session.

On this page