Speech-to-text translation
Get a translated transcript alongside the original from a transcription request.
Add options.translation.target_language to a transcription request. Each transcript then carries the translated text beside the original. It works on the streaming socket and on batch uploads.
{
"type": "session.configure",
"routing": { "mode": "explicit", "provider": "soniox", "model": "stt-rt-v5" },
"audio": { "encoding": "pcm_s16le", "sample_rate_hz": 16000, "channels": 1 },
"options": { "translation": { "target_language": "es" } }
}Send it as the first frame on /v1/stt/stream. The same options object works on POST /v1/stt/transcriptions.
Response fields
text stays the original transcript, so a client that ignores translation is unchanged. The translation arrives in a new translation field:
{ "type": "transcript.final", "text": "Hello, how are you?", "translation": "Hola, ¿cómo estás?", "segments": [] }| Where | Field |
|---|---|
transcript.delta | translation: translated text so far, when the chunk has any |
transcript.final | translation: the segment's translation |
Batch TranscriptionResponse | translation: the whole translated transcript |
A translation can land on a later final than its source words. Concatenate translation across finals rather than pairing them one to one. Translated text carries no word timings.
Options
| Field | Type | Required | Meaning |
|---|---|---|---|
options.translation.target_language | string | yes | BCP-47 tag of the output language, such as es or pt-BR |
Supported models
Only models that list capabilities.translation: true in GET /v1/models accept the option. Today that is soniox:stt-rt-v5. Any other route refuses the request with capability_unsupported, before admission.
Pricing
Soniox includes translation in its transcription price, so a translated session bills at the model's normal speech-to-text rate. See Billing.