Transcriber
Smallest AI (Pulse) logo

Smallest AI (Pulse)

Add Smallest AI Pulse speech-to-text to Bolna for ultra-fast, low-latency phone call transcription. Built for real-time voice agent conversations.

stttranscriberlow-latencysmallestpulse
At a glance

How Smallest AI (Pulse) fits in the stack

Best for

Turning noisy phone audio into text the agent can reason over.

Use this layer when

Recognition accuracy, language coverage, and latency matter most.

Connects to

Telephony audio upstream and your LLM downstream.

Voice Stack

Speech-to-text is the listening layer

Speech-to-text is the listening layer

This provider converts raw audio into text in real time. It shapes how accurately the agent hears the caller and how natural the back-and-forth feels.

TelephonyPhone Network
STTListener
LLMReasoning
TTSVoice
ToolsActions

This page focuses on where Smallest AI (Pulse) fits in a production voice stack. For full setup steps, credentials, and API details, use the documentation link above.

Overview

Pulse is Smallest AI's speech-to-text model, built for real-time transcription at very low latency. With Bolna's Smallest AI integration, voice agents can transcribe caller audio on phone calls with minimal delay, keeping conversations feeling natural instead of laggy.

Pulse is ideal for teams that want a fast, cost-effective transcriber to pair with Smallest AI's synthesizer for a lightweight, low-latency voice stack end to end.

Features & Use Cases

Ultra-Low Latency
Optimized for real-time streaming transcription so agents can respond the moment a caller stops speaking.

Cost Effective
Lower pricing structure suited to high volume deployments where transcription costs add up fast.

Built for Voice Agents
Tuned for live phone audio rather than offline transcription, handling noisy lines and natural speech patterns.

Pairs with Smallest AI TTS
Combine with Smallest AI's synthesizer for a full lightweight speech stack optimized for speed end to end.

Use Case: High Volume Call Centers
Run large scale outbound or inbound calling operations where transcription latency directly affects conversation quality.

Use Case: Cost-Sensitive Deployments
Keep per-minute transcription costs low for high call volume use cases without sacrificing real-time responsiveness.

Browse this layer

Keep exploring the voice stack

Browse

Speech-to-Text

Speech-to-text converts what callers say into text that your LLM can process. Transcription accuracy and latency directly affect how natural a conversation feels. Bolna supports streaming STT providers optimized for telephony audio, including specialized models for Indian languages.

Browse

Telephony

Telephony providers connect your voice agents to the phone network so they can make and receive real calls. Bolna supports managed integrations with major carriers as well as bring-your-own-carrier via SIP trunking, giving you full control over call routing, number provisioning, and cost.

Browse

Large Language Models

The LLM is the brain of your voice agent. It understands what callers say and decides how to respond. Bolna lets you swap between models like GPT-4o, Claude, and DeepSeek without changing your agent configuration, so you can optimize for speed, cost, or reasoning depth.

Browse

Text-to-Speech

Text-to-speech turns your agent responses into spoken audio. Voice quality shapes how callers perceive your brand. Flat, robotic speech kills trust while natural, expressive voices build it. Bolna integrates with the fastest TTS providers so responses sound human and arrive without awkward pauses.

Browse

Tools & Workflows

Tools let your voice agents take action during a call, not just talk. Book a calendar slot, look up an order in Shopify, push a lead into your CRM, or trigger a multi-step automation in Zapier. These integrations turn voice agents from answering machines into workflow engines.

See where Smallest AI (Pulse) fits in your production workflow

Use the demo to walk through provider selection, stack tradeoffs, and the exact workflow you want Bolna to automate.