Skip to main content

What is Soniox TTS?

Soniox Text-to-Speech is a real-time speech synthesis platform built for conversational AI. Its tts-rt-v2 model streams speech as the reply is being generated, so audio starts playing before the full sentence is ready. One model covers every supported language and voice, so a multilingual agent does not need a different model per language.

Key Features of Soniox TTS

Soniox TTS offers several features that make it well suited to real-time voice agents: Broad Multilingual Coverage from One Model: Speaks 63 languages, including ten Indian languages, without switching models. Real-Time Streaming Synthesis: Audio is produced as the reply arrives rather than after it completes, keeping responses quick in live conversation. Large Named Voice Catalogue: 200 built-in voices per model, addressed by name rather than by an opaque id, spanning a range of accents and speaking styles. Voice Cloning: Clone a voice in your own Soniox account and put it on an agent by passing its id. Delivery Controls: speed adjusts pace between 0.7 and 1.3, and reduce_silence tightens the pauses inserted between sentences.

How Bolna Uses Soniox for TTS

Bolna AI integrates Soniox’s real-time synthesis to produce the spoken side of its voice agents. Here’s how Bolna leverages Soniox TTS: Streaming Speech for Natural Conversation Flow: Bolna speaks Soniox audio as the reply is generated rather than waiting for the whole response, so agents answer without the pause that makes automated calls feel stilted. Responsive Barge-In: When a caller interrupts, the agent stops mid-sentence and drops the rest of the utterance cleanly, so it starts listening instead of talking over them. Multilingual Agents on a Single Voice Provider: Because one Soniox model covers every supported language, Bolna agents can serve callers across languages without changing synthesizer per language. Consistent Quality on Telephony and Web: Bolna delivers Soniox audio at the quality each channel expects, so the same agent sounds right on a phone call and in the browser.

List of Soniox TTS models supported on Bolna AI

Voices

Soniox addresses its built-in voices by name rather than by an opaque id, so voice and voice_id carry the same value - Adrian, not a UUID. Both fields are required in provider_config, so send the name twice. Point voice_id at a cloned voice’s id to use a clone instead; it takes precedence when the two differ. A few of the voices in the catalogue: A voice that is not in the catalogue is rejected when the agent is saved.
The full Soniox catalogue runs to 200 voices per model and changes over time. Confirm what is currently selectable on your account with List voices, or browse and preview them in the Languages tab.

Supported Languages

One tts-rt-v2 model speaks 63 languages, so a multilingual agent can keep a single synthesizer across its whole language range instead of switching providers per language. Indian languages: Hindi (hi), Bengali (bn), Tamil (ta), Telugu (te), Gujarati (gu), Kannada (kn), Malayalam (ml), Marathi (mr), Punjabi (pa), Urdu (ur) Other languages: major European, Asian and Middle Eastern languages are supported too, including English, Spanish, French, German, Arabic, Japanese and Chinese. Set the language in provider_config.language using its plain ISO code: hi, not hi-IN. Browse every language, and the voices available for each, in the Languages tab.
language inside provider_config is required for Soniox. Leaving it out fails with 400 Tools Config > Voice > Language: This field is required, even when the agent’s transcriber already has a language set.
Running an agent in more than one language? Because one Soniox model covers all 63, you can keep the same synthesizer for every language and only change the voice. See multilingual support.

Configuration

speed and reduce_silence are only sent when you set them; left out, each uses Soniox’s own default for the model. speed is validated when the agent is saved, so a value outside 0.7 to 1.3 is rejected at agent setup rather than failing mid-call. See the Create Agent API for every synthesizer field.

Conclusion

Soniox TTS gives Bolna agents real-time streaming speech across a wide language range from a single model, with a large named voice catalogue and cloning for teams that want a distinct voice identity. For related integrations: