Speech-to-speech translation
Translate live speech into speech in another language over one WebSocket.
A speech-to-speech translation session takes a speaker's audio and returns translated speech, with transcripts of both languages. Each model speaks its vendor's native protocol on its own route. Pick the model explicitly: translation models never join automatic routing.
import WebSocket from 'ws';
const ws = new WebSocket(
'wss://router.speko.dev/v1/realtime/translations?model=gpt-realtime-translate',
{ headers: { Authorization: `Bearer ${process.env.SPEKO_API_KEY}`, 'Idempotency-Key': crypto.randomUUID() } },
);
ws.on('open', () => {
ws.send(JSON.stringify({
type: 'session.update',
session: { audio: { output: { language: 'es' }, input: { transcription: { model: 'gpt-realtime-whisper' } } } },
}));
// Stream 24 kHz mono PCM16, base64, continuously (silence included):
// ws.send(JSON.stringify({ type: 'session.input_audio_buffer.append', audio: chunk.toString('base64') }));
});
ws.on('message', (raw) => {
const event = JSON.parse(raw);
if (event.type === 'session.output_audio.delta') play(Buffer.from(event.delta, 'base64'));
if (event.type === 'session.output_transcript.delta') process.stdout.write(event.delta);
});When the speaker stops, send {"type":"session.close"} and keep reading until session.closed. The vendor flushes buffered audio and sends the last translation before it closes.
Models
| Model | Route | Input / output audio | Target languages |
|---|---|---|---|
gpt-realtime-translate | GET /v1/realtime/translations?model=… | PCM16 24 kHz / 24 kHz | 13: es, pt, fr, ja, ru, zh, de, ko, hi, id, vi, it, en |
gemini-3.5-live-translate-preview | GET /v1/bidi | PCM16 16 kHz / 24 kHz | 70+ (BCP-47) |
qwen3.8-livetranslate-flash-realtime | GET /v1/realtime/translations/qwen?model=… | PCM16 16 kHz / 24 kHz | 29 with speech output, 60 with text |
All three detect the source language on their own. GET /v1/models lists them with capabilities.translation: true.
OpenAI: gpt-realtime-translate
The only command is session.update, restricted to the translation shape:
| Field | Values |
|---|---|
session.audio.output.language | One of the 13 target languages above. Required. |
session.audio.input.transcription | {"model":"gpt-realtime-whisper"} for a source transcript, or null |
session.audio.input.noise_reduction | {"type":"near_field"}, {"type":"far_field"}, or null |
Any other field is refused with invalid_request. There is no voice, no instructions, no tools, and no response.create. The translation follows the speaker continuously.
| Server event | Meaning |
|---|---|
session.output_audio.delta | Translated speech, base64 PCM16 |
session.output_transcript.delta | Translated text |
session.input_transcript.delta | Source-language text (needs transcription) |
session.closed | The session ended after session.close |
Google: gemini-3.5-live-translate-preview
The first frame is the vendor's setup, with the target in generationConfig.translationConfig:
{
"setup": {
"model": "models/gemini-3.5-live-translate-preview",
"generationConfig": {
"responseModalities": ["AUDIO"],
"translationConfig": { "targetLanguageCode": "es" }
},
"inputAudioTranscription": {},
"outputAudioTranscription": {}
}
}After setupComplete, send audio as realtimeInput.audio with mimeType: "audio/pcm;rate=16000". Translated speech arrives in serverContent.modelTurn.parts[].inlineData. The transcripts arrive in serverContent.inputTranscription and serverContent.outputTranscription.
Keep streaming a few seconds of silence after the speaker stops, before audioStreamEnd. The model finalizes an utterance only when its own end-of-speech detection fires. Ending the stream sooner drops the last utterance untranslated.
Alibaba: qwen3.8-livetranslate-flash-realtime
Send exactly one session.update, before any audio. It names the target in session.translation.language:
{ "type": "session.update", "session": { "translation": { "language": "es" }, "output_modalities": ["text", "audio"] } }The vendor refuses a second session.update and closes the session, so the Router refuses it first with invalid_request. The Router fixes the audio formats to PCM (16 kHz in, 24 kHz out). An update that names another format or sample rate is refused.
Stream audio as input_audio_buffer.append. The translation arrives as response.audio.delta and response.audio_transcript.delta, with one response.done per segment. End with {"type":"session.finish"} and read until session.finished. Closing the socket without it drops the final segment.
Pricing
With Speko-managed provider keys, the Router bills each session at the vendor's list price plus 5%. With BYOK, the vendor bills you directly. See Billing.
| Model | Vendor list price |
|---|---|
gpt-realtime-translate | $0.034 per minute of streamed audio |
gemini-3.5-live-translate-preview | $3.50 per M input audio tokens, $21 per M output audio tokens (about $0.037/min) |
qwen3.8-livetranslate-flash-realtime | $7.50 per M input audio tokens, $20 per M output text tokens, $30 per M output audio tokens |