> ## Documentation Index
> Fetch the complete documentation index at: https://www.bolna.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Soniox Text to Speech with Bolna Voice AI agents

> Integrate Soniox real-time TTS with your Bolna Voice AI agents for low-latency streaming speech across 63 languages, named voices, and voice cloning.

## What is Soniox TTS?

[Soniox](https://soniox.com/) Text-to-Speech is a real-time speech synthesis platform built for conversational AI. Its `tts-rt-v2` model streams speech as the reply is being generated, so audio starts playing before the full sentence is ready.

One model covers every supported language and voice, so a multilingual agent does not need a different model per language.

## Key Features of Soniox TTS

Soniox TTS offers several features that make it well suited to real-time voice agents:

**Broad Multilingual Coverage from One Model**: Speaks 63 languages, including ten Indian languages, without switching models.

**Real-Time Streaming Synthesis**: Audio is produced as the reply arrives rather than after it completes, keeping responses quick in live conversation.

**Large Named Voice Catalogue**: 200 built-in voices per model, addressed by name rather than by an opaque id, spanning a range of accents and speaking styles.

**Voice Cloning**: Clone a voice in your own Soniox account and put it on an agent by passing its id.

**Delivery Controls**: `speed` adjusts pace between 0.7 and 1.3, and `reduce_silence` tightens the pauses inserted between sentences.

## How Bolna Uses Soniox for TTS

Bolna AI integrates Soniox's real-time synthesis to produce the spoken side of its voice agents. Here's how Bolna leverages Soniox TTS:

**Streaming Speech for Natural Conversation Flow**:
Bolna speaks Soniox audio as the reply is generated rather than waiting for the whole response, so agents answer without the pause that makes automated calls feel stilted.

**Responsive Barge-In**:
When a caller interrupts, the agent stops mid-sentence and drops the rest of the utterance cleanly, so it starts listening instead of talking over them.

**Multilingual Agents on a Single Voice Provider**:
Because one Soniox model covers every supported language, Bolna agents can serve callers across languages without changing synthesizer per language.

**Consistent Quality on Telephony and Web**:
Bolna delivers Soniox audio at the quality each channel expects, so the same agent sounds right on a phone call and in the browser.

## List of Soniox TTS models supported on Bolna AI

| Model | Description |
| - | - |
| `tts-rt-v2` | Current real-time model and the default on Bolna - 63 languages and 200 voices |
| `tts-rt-v1` | Earlier real-time model, deprecated by Soniox. Use `tts-rt-v2` on new agents |

## Voices

Soniox addresses its built-in voices **by name** rather than by an opaque id, so `voice` and `voice_id` carry the same value - `Adrian`, not a UUID. Both fields are required in `provider_config`, so send the name twice. Point `voice_id` at a cloned voice's id to use a clone instead; it takes precedence when the two differ.

A few of the voices in the catalogue:

| Voice | Gender | Character |
| - | - | - |
| `Adrian` | Male | Deep and composed, with crisp articulation and measured pacing |
| `Emma` | Female | Offhand British delivery with a light lift at the ends of lines |
| `Arjun` | Male | Indian English, round and talkative, holds up well for narration |
| `Priya` | Female | Natural Indian accent, warm pacing, helpful and easy to trust |
| `Divya` | Female | Friendly Tamil voice with a natural conversational rhythm |

A voice that is not in the catalogue is rejected when the agent is saved.

<Tip>
  The full Soniox catalogue runs to 200 voices per model and changes over time. Confirm what is currently selectable on your account with [List voices](/docs/api-reference/voice/get_all), or browse and preview them in the [Languages tab](/docs/agent-setup/audio-tab).
</Tip>

## Supported Languages

One `tts-rt-v2` model speaks 63 languages, so a multilingual agent can keep a single synthesizer across its whole language range instead of switching providers per language.

**Indian languages**: Hindi (`hi`), Bengali (`bn`), Tamil (`ta`), Telugu (`te`), Gujarati (`gu`), Kannada (`kn`), Malayalam (`ml`), Marathi (`mr`), Punjabi (`pa`), Urdu (`ur`)

**Other languages**: major European, Asian and Middle Eastern languages are supported too, including English, Spanish, French, German, Arabic, Japanese and Chinese.

Set the language in `provider_config.language` using its plain ISO code: `hi`, not `hi-IN`. Browse every language, and the voices available for each, in the [Languages tab](/docs/agent-setup/audio-tab).

<Warning>
  `language` inside `provider_config` is **required** for Soniox. Leaving it out fails with
  `400 Tools Config > Voice > Language: This field is required`, even when the agent's transcriber already has a language set.
</Warning>

<Tip>
  Running an agent in more than one language? Because one Soniox model covers all 63, you can keep the same synthesizer for every language and only change the voice. See [multilingual support](/docs/customizations/multilingual-languages-support).
</Tip>

## Configuration

```json theme={"system"}
{
  "synthesizer": {
    "provider": "soniox",
    "provider_config": {
      "voice": "Adrian",
      "voice_id": "Adrian",
      "model": "tts-rt-v2",
      "language": "en",
      "speed": 1.0
    },
    "stream": true,
    "buffer_size": 400
  }
}
```

| Setting | Type | Default | Description |
| - | - | - | - |
| `voice` | string | - | Built-in voice name. Required |
| `voice_id` | string | - | The same name again for a built-in voice, or a cloned voice's id. Required |
| `model` | string | `tts-rt-v2` | One of the models above |
| `language` | string | - | Plain ISO code. Required |
| `speed` | float | Soniox default | `0.7` to `1.3` |
| `reduce_silence` | bool | Soniox default | Shortens the pauses Soniox inserts between sentences |
| `stream` | bool | `false` | Enable streaming audio output |
| `buffer_size` | integer | `400` | Characters buffered before the first chunk is sent |

`speed` and `reduce_silence` are only sent when you set them; left out, each uses Soniox's own default for the model. `speed` is validated when the agent is saved, so a value outside 0.7 to 1.3 is rejected at agent setup rather than failing mid-call. See the [Create Agent API](/docs/api-reference/agent/v2/create#synthesizer) for every synthesizer field.

## Conclusion

Soniox TTS gives Bolna agents real-time streaming speech across a wide language range from a single model, with a large named voice catalogue and cloning for teams that want a distinct voice identity.

For related integrations:

* Combine with the [Soniox transcriber](/docs/soniox) for a complete Soniox integration
* Give each language its own voice with [multilingual support](/docs/customizations/multilingual-languages-support)
* Compare against other providers in the [synthesizer comparison](/docs/concepts/choosing-providers#synthesizers-text-to-speech)
* You can also connect your own Soniox account and use it with [Bolna AI](/docs/providers)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.