Transcriber
Soniox logo

Soniox

Add Soniox real-time speech-to-text to Bolna voice agents for multilingual transcription. One model covers 60+ languages and code-switches mid-sentence on live phone calls.

stttranscribermultilingualsonioxcode-switchinghinglishreal-time
At a glance

How Soniox fits in the stack

Best for

Turning noisy phone audio into text the agent can reason over.

Use this layer when

Recognition accuracy, language coverage, and latency matter most.

Connects to

Telephony audio upstream and your LLM downstream.

Voice Stack

Speech-to-text is the listening layer

Speech-to-text is the listening layer

This provider converts raw audio into text in real time. It shapes how accurately the agent hears the caller and how natural the back-and-forth feels.

TelephonyPhone Network
→
STTListener
→
LLMReasoning
→
TTSVoice
→
ToolsActions

This page focuses on where Soniox fits in a production voice stack. For full setup steps, credentials, and API details, use the documentation link above.

Overview

Soniox is a real-time speech recognition platform built around a single multilingual model. Instead of running a separate recognizer per language, one Soniox model transcribes whatever is spoken, including mid-sentence switches between languages, over a single streaming connection.

With Bolna's Soniox integration, your voice agents transcribe caller audio using the real-time stt-rt-v5 model, which combines accuracy across 60+ languages with per-token language identification and semantic endpoint detection for natural turn-taking.

Models

Bolna supports one Soniox model:

  • Soniox v5 (stt-rt-v5) - Real-time multilingual model with code-switching, language identification, and semantic endpoint detection.

Supported Languages

Soniox on Bolna supports multilingual auto-detect plus these languages:

English, English (India), Hindi, Bengali, Tamil, Telugu, Gujarati, Kannada, Malayalam, Marathi, and Punjabi.

See the Bolna docs for the language codes you can set on an agent today.

Features & Use Cases

Native Multilingual, One Connection
A single model handles all supported languages and switches between them automatically, so there is no per-language setup and no routing callers through a language menu first.

Code-Switching Ready
Built for real-world speech where callers move between English and a regional language inside the same sentence, like Hinglish, which is normal across Indian markets.

Semantic Endpoint Detection
Soniox detects when a caller has actually finished their turn rather than waiting on a fixed silence timer, so the agent can respond sooner without cutting people off.

Per-Token Language Identification
Every transcribed token carries its detected language, giving downstream logic an accurate real-time view of what the caller is speaking.

Keyword and Context Biasing
Pass product names, brand terms, and free-form call context to improve recognition on the vocabulary your agents actually hear.

Use Case: Multilingual Support Lines
Run one agent for callers across languages instead of maintaining a separate agent or IVR branch per language.

Use Case: Hinglish Sales and Collections
Handle callers who mix English and Hindi naturally on outbound calls without reconfiguring the transcriber per contact.

Use Case: Faster Turn-Taking
Use semantic endpointing to shorten the gap between a caller finishing and the agent replying, keeping conversations from feeling laggy.

Browse this layer

Keep exploring the voice stack

Browse

Speech-to-Text

Speech-to-text converts what callers say into text that your LLM can process. Transcription accuracy and latency directly affect how natural a conversation feels. Bolna supports streaming STT providers optimized for telephony audio, including specialized models for Indian languages.

Browse

Telephony

Telephony providers connect your voice agents to the phone network so they can make and receive real calls. Bolna supports managed integrations with major carriers as well as bring-your-own-carrier via SIP trunking, giving you full control over call routing, number provisioning, and cost.

Browse

Large Language Models

The LLM is the brain of your voice agent. It understands what callers say and decides how to respond. Bolna lets you swap between models like GPT-4o, Claude, and DeepSeek without changing your agent configuration, so you can optimize for speed, cost, or reasoning depth.

Browse

Text-to-Speech

Text-to-speech turns your agent responses into spoken audio. Voice quality shapes how callers perceive your brand. Flat, robotic speech kills trust while natural, expressive voices build it. Bolna integrates with the fastest TTS providers so responses sound human and arrive without awkward pauses.

Browse

Tools & Workflows

Tools let your voice agents take action during a call, not just talk. Book a calendar slot, look up an order in Shopify, push a lead into your CRM, or trigger a multi-step automation in Zapier. These integrations turn voice agents from answering machines into workflow engines.

See where Soniox fits in your production workflow

Use the demo to walk through provider selection, stack tradeoffs, and the exact workflow you want Bolna to automate.