Speko Docs
Translation

Speech-to-speech translation

Translate live speech into speech in another language over one WebSocket.

A speech-to-speech translation session takes a speaker's audio and returns translated speech, with transcripts of both languages. Each model speaks its vendor's native protocol on its own route. Pick the model explicitly: translation models never join automatic routing.

import WebSocket from 'ws';

const ws = new WebSocket(
  'wss://router.speko.dev/v1/realtime/translations?model=gpt-realtime-translate',
  { headers: { Authorization: `Bearer ${process.env.SPEKO_API_KEY}`, 'Idempotency-Key': crypto.randomUUID() } },
);

ws.on('open', () => {
  ws.send(JSON.stringify({
    type: 'session.update',
    session: { audio: { output: { language: 'es' }, input: { transcription: { model: 'gpt-realtime-whisper' } } } },
  }));
  // Stream 24 kHz mono PCM16, base64, continuously (silence included):
  // ws.send(JSON.stringify({ type: 'session.input_audio_buffer.append', audio: chunk.toString('base64') }));
});

ws.on('message', (raw) => {
  const event = JSON.parse(raw);
  if (event.type === 'session.output_audio.delta') play(Buffer.from(event.delta, 'base64'));
  if (event.type === 'session.output_transcript.delta') process.stdout.write(event.delta);
});

When the speaker stops, send {"type":"session.close"} and keep reading until session.closed. The vendor flushes buffered audio and sends the last translation before it closes.

Models

ModelRouteInput / output audioTarget languages
gpt-realtime-translateGET /v1/realtime/translations?model=…PCM16 24 kHz / 24 kHz13: es, pt, fr, ja, ru, zh, de, ko, hi, id, vi, it, en
gemini-3.5-live-translate-previewGET /v1/bidiPCM16 16 kHz / 24 kHz70+ (BCP-47)
qwen3.8-livetranslate-flash-realtimeGET /v1/realtime/translations/qwen?model=…PCM16 16 kHz / 24 kHz29 with speech output, 60 with text

All three detect the source language on their own. GET /v1/models lists them with capabilities.translation: true.

OpenAI: gpt-realtime-translate

The only command is session.update, restricted to the translation shape:

FieldValues
session.audio.output.languageOne of the 13 target languages above. Required.
session.audio.input.transcription{"model":"gpt-realtime-whisper"} for a source transcript, or null
session.audio.input.noise_reduction{"type":"near_field"}, {"type":"far_field"}, or null

Any other field is refused with invalid_request. There is no voice, no instructions, no tools, and no response.create. The translation follows the speaker continuously.

Server eventMeaning
session.output_audio.deltaTranslated speech, base64 PCM16
session.output_transcript.deltaTranslated text
session.input_transcript.deltaSource-language text (needs transcription)
session.closedThe session ended after session.close

Google: gemini-3.5-live-translate-preview

The first frame is the vendor's setup, with the target in generationConfig.translationConfig:

{
  "setup": {
    "model": "models/gemini-3.5-live-translate-preview",
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "translationConfig": { "targetLanguageCode": "es" }
    },
    "inputAudioTranscription": {},
    "outputAudioTranscription": {}
  }
}

After setupComplete, send audio as realtimeInput.audio with mimeType: "audio/pcm;rate=16000". Translated speech arrives in serverContent.modelTurn.parts[].inlineData. The transcripts arrive in serverContent.inputTranscription and serverContent.outputTranscription.

Keep streaming a few seconds of silence after the speaker stops, before audioStreamEnd. The model finalizes an utterance only when its own end-of-speech detection fires. Ending the stream sooner drops the last utterance untranslated.

Alibaba: qwen3.8-livetranslate-flash-realtime

Send exactly one session.update, before any audio. It names the target in session.translation.language:

{ "type": "session.update", "session": { "translation": { "language": "es" }, "output_modalities": ["text", "audio"] } }

The vendor refuses a second session.update and closes the session, so the Router refuses it first with invalid_request. The Router fixes the audio formats to PCM (16 kHz in, 24 kHz out). An update that names another format or sample rate is refused.

Stream audio as input_audio_buffer.append. The translation arrives as response.audio.delta and response.audio_transcript.delta, with one response.done per segment. End with {"type":"session.finish"} and read until session.finished. Closing the socket without it drops the final segment.

Pricing

With Speko-managed provider keys, the Router bills each session at the vendor's list price plus 5%. With BYOK, the vendor bills you directly. See Billing.

ModelVendor list price
gpt-realtime-translate$0.034 per minute of streamed audio
gemini-3.5-live-translate-preview$3.50 per M input audio tokens, $21 per M output audio tokens (about $0.037/min)
qwen3.8-livetranslate-flash-realtime$7.50 per M input audio tokens, $20 per M output text tokens, $30 per M output audio tokens

On this page