Skip to main content
The synthesizer turns the agent’s reply into speech. Pick a voice per language on the agent’s Languages tab, or set it in the synthesizer block of the agent config. Each provider below has its own guide with models, configuration and troubleshooting. Not sure which to choose? Choose providers compares them by language, latency and cost. To use your own account with a provider, see Connect provider credentials.

AWS Polly

Learn how to integrate and use AWS Polly TTS with Bolna Voice AI agents including Amazon’s neural, generative and standard models.

Azure

Integrate Microsoft Azure Text-to-Speech with Bolna to create natural, expressive Voice AI agents. Supports neural voices and multilingual output.

Cartesia

Enable Cartesia voices in Bolna Voice AI agents for expressive, customizable AI voices using their latest voice models for multilingual Indian voices.

Deepgram

Integrate and use your Bolna Voice AI agents with high-quality neural voices from Deepgram for natural, human-like conversational experiences.

ElevenLabs

Configure ElevenLabs text-to-speech in Bolna voice agents, including Turbo v2.5, Flash v2.5, Eleven v3 and Eleven v4 Turbo: models, voices, streaming, and cloning.

Kalpa Labs

Configure Kalpa Labs text-to-speech in Bolna voice agents: models, voices, code-switched Hinglish, streaming, and telephony audio.

Maya

Configure Maya Research’s Maya 2 Native text-to-speech in Bolna voice agents: voices, supported languages, mid-call language switching, and audio.

Rime

Integrate Rime TTS with Bolna Voice AI agents for ultra-fast, expressive speech synthesis with sub-200ms latency and diverse conversational voice options.

Sarvam

Integrate and use your Bolna Voice AI agents with high-quality neural voices from Sarvam’s Bulbul models for natural, human-like Indian-language speech.

Smallest

Integrate and use Smallest voices with Bolna Voice AI agents for lightweight and efficient text-to-speech solutions.

Soniox

Integrate Soniox real-time TTS with your Bolna Voice AI agents for low-latency streaming speech across 63 languages, named voices, and voice cloning.