This page focuses on where Soniox fits in a production voice stack. For full setup steps, credentials, and API details, use the documentation link above.
Overview
Soniox is a real-time speech recognition platform built around a single multilingual model. Instead of running a separate recognizer per language, one Soniox model transcribes whatever is spoken, including mid-sentence switches between languages, over a single streaming connection.
With Bolna's Soniox integration, your voice agents transcribe caller audio using the real-time stt-rt-v5 model, which combines accuracy across 60+ languages with per-token language identification and semantic endpoint detection for natural turn-taking.
Models
Bolna supports one Soniox model:
Soniox v5(stt-rt-v5) - Real-time multilingual model with code-switching, language identification, and semantic endpoint detection.
Supported Languages
Soniox on Bolna supports multilingual auto-detect plus these languages:
English, English (India), Hindi, Bengali, Tamil, Telugu, Gujarati, Kannada, Malayalam, Marathi, and Punjabi.
See the Bolna docs for the language codes you can set on an agent today.
Features & Use Cases
Native Multilingual, One Connection
A single model handles all supported languages and switches between them automatically, so there is no per-language setup and no routing callers through a language menu first.
Code-Switching Ready
Built for real-world speech where callers move between English and a regional language inside the same sentence, like Hinglish, which is normal across Indian markets.
Semantic Endpoint Detection
Soniox detects when a caller has actually finished their turn rather than waiting on a fixed silence timer, so the agent can respond sooner without cutting people off.
Per-Token Language Identification
Every transcribed token carries its detected language, giving downstream logic an accurate real-time view of what the caller is speaking.
Keyword and Context Biasing
Pass product names, brand terms, and free-form call context to improve recognition on the vocabulary your agents actually hear.
Use Case: Multilingual Support Lines
Run one agent for callers across languages instead of maintaining a separate agent or IVR branch per language.
Use Case: Hinglish Sales and Collections
Handle callers who mix English and Hindi naturally on outbound calls without reconfiguring the transcriber per contact.
Use Case: Faster Turn-Taking
Use semantic endpointing to shorten the gap between a caller finishing and the agent replying, keeping conversations from feeling laggy.