Transcriber
AssemblyAI logo

AssemblyAI

Add AssemblyAI Universal-3.6 Pro speech-to-text to Bolna voice agents. Built for real calls, it covers 32 languages with code-switching and gets short answers like "yes" right.

stttranscribermultilingualassemblyaicode-switchingkeywordsreal-time
At a glance

How AssemblyAI fits in the stack

Best for

Turning noisy phone audio into text the agent can reason over.

Use this layer when

Recognition accuracy, language coverage, and latency matter most.

Connects to

Telephony audio upstream and your LLM downstream.

Voice Stack

Speech-to-text is the listening layer

Speech-to-text is the listening layer

This provider converts raw audio into text in real time. It shapes how accurately the agent hears the caller and how natural the back-and-forth feels.

TelephonyPhone Network
→
STTListener
→
LLMReasoning
→
TTSVoice
→
ToolsActions

This page focuses on where AssemblyAI fits in a production voice stack. For full setup steps, credentials, and API details, use the documentation link above.

Overview

AssemblyAI is a speech recognition platform whose streaming models are trained on real voice-agent and phone conversations rather than clean studio audio. That focus shows up where voice agents usually break: one-word replies, background voices on the line, and callers who change language mid-sentence.

With Bolna's AssemblyAI integration, your voice agents transcribe caller audio using the universal-3-6-pro model, which code-switches across 32 languages and accepts both a keyword list and a free-form description of the call to bias recognition towards your own vocabulary.

Models

Bolna supports two AssemblyAI streaming models:

  • Universal-3.6 Pro (universal-3-6-pro) - Current model, trained on tens of thousands of hours of real voice-agent and phone audio. 32 languages, native code-switching, and precise end-of-turn detection. Recommended default.
  • Universal-3.5 Pro (universal-3-5-pro) - Previous generation, still supported, covering 19 languages.

Universal-3.6 Pro is a meaningful step up on live calls rather than on benchmarks. Against Universal-3.5 Pro it makes 45% fewer errors on short answers like "yes" or "no thanks", rising to 52% fewer in heavy background noise. It leaks 22% less background speech into the transcript, 29% less in heavy noise, and transcribes 74% fewer sessions in the wrong language. Median latency is unchanged, and AssemblyAI charges the same rate for both models, so moving up is a one-field change on your agent.

Supported Languages

Universal-3.6 Pro code-switches across 32 languages:

Afrikaans, Arabic, Cantonese, Catalan, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Hebrew, Hindi, Italian, Japanese, Korean, Mandarin, Marathi, Norwegian, Norwegian Nynorsk, Persian, Portuguese, Romanian, Russian, Spanish, Swedish, Turkish, Urdu, Vietnamese, Xhosa, and Zulu.

Universal-3.5 Pro covers 19 of these. The 13 that are only available on Universal-3.6 Pro are Afrikaans, Cantonese, Estonian, Galician, Korean, Marathi, Norwegian Nynorsk, Persian, Romanian, Russian, Urdu, Xhosa, and Zulu.

See the Bolna docs for the language codes you can set on an agent today.

Features & Use Cases

Accurate on Short Answers
Qualification and confirmation calls turn on single words. Universal-3.6 Pro is tuned for exactly those turns, so a "yes", a "no thanks", or a spoken digit is far less likely to be misheard or dropped.

Holds the Caller's Language
Native code-switching across 32 languages, with far fewer sessions drifting into the wrong language, so a caller who mixes English and a regional language is transcribed as they actually spoke.

Resists Background Voices
Less background speech leaks into the transcript, which matters on calls taken in shops, households, and open offices where the caller is not the only person talking.

Keyword and Context Biasing
AssemblyAI is one of the few providers that accepts both a keyword list and free-form call context together, so brand names, product codes, and the shape of the conversation can all be passed in.

Precise End-of-Turn Detection
The model marks the end of a speaking turn explicitly instead of relying on a fixed silence timer, so the agent replies sooner without talking over callers who pause mid-thought.

Use Case: Lead Qualification and Confirmations
Run outbound calls where the whole outcome is a short yes or no, and get that answer right on noisy mobile lines.

Use Case: Multilingual Support Lines
Serve callers across 32 languages from a single agent, without a language menu in front of the conversation.

Use Case: Brand and Product Vocabulary
Pass your product names and account identifiers as keywords with a sentence of call context, so proper nouns land correctly in the transcript and in anything extracted from it.

Browse this layer

Keep exploring the voice stack

Browse

Speech-to-Text

Speech-to-text converts what callers say into text that your LLM can process. Transcription accuracy and latency directly affect how natural a conversation feels. Bolna supports streaming STT providers optimized for telephony audio, including specialized models for Indian languages.

Browse

Telephony

Telephony providers connect your voice agents to the phone network so they can make and receive real calls. Bolna supports managed integrations with major carriers as well as bring-your-own-carrier via SIP trunking, giving you full control over call routing, number provisioning, and cost.

Browse

Large Language Models

The LLM is the brain of your voice agent. It understands what callers say and decides how to respond. Bolna lets you swap between models like GPT-4o, Claude, and DeepSeek without changing your agent configuration, so you can optimize for speed, cost, or reasoning depth.

Browse

Text-to-Speech

Text-to-speech turns your agent responses into spoken audio. Voice quality shapes how callers perceive your brand. Flat, robotic speech kills trust while natural, expressive voices build it. Bolna integrates with the fastest TTS providers so responses sound human and arrive without awkward pauses.

Browse

Tools & Workflows

Tools let your voice agents take action during a call, not just talk. Book a calendar slot, look up an order in Shopify, push a lead into your CRM, or trigger a multi-step automation in Zapier. These integrations turn voice agents from answering machines into workflow engines.

See where AssemblyAI fits in your production workflow

Use the demo to walk through provider selection, stack tradeoffs, and the exact workflow you want Bolna to automate.