1. Stay inside the flat rate
Your workspace has a flat per-minute rate that covers transcription, LLM and voice as long as the agent uses preferred models. One non-preferred component moves that component onto variable, usage-based pricing — the single most common reason two agents with identical volume bill differently.1
Audit the three model choices on every production agent
Check the transcription and voice models on the Languages tab and the LLM on the Intelligence tab against the current preferred list.
2
Justify each exception
Keep a premium model only where a test showed it materially improves outcomes — accuracy in a specific Indian language, for instance. Otherwise revert.
2. Shorten the call
Every component except the LLM bills by duration, so seconds are money.- Cap total duration with
call_terminate, and end dead air with silence-based hangup — End a call automatically - Cut the greeting to one sentence; long welcome messages are billed at both TTS and telephony rates
- Instruct the agent to answer in one or two sentences and ask one question at a time — Write better prompts
- Detect voicemail so the agent does not deliver a full pitch to an answering machine — Call tab
3. Cut tokens
The LLM bills by tokens generated. Trim the system prompt, avoid restating context the agent already has, keep knowledge base documents scoped to what is actually needed, and prefer a smaller preferred model for simple flows — a reminder or confirmation agent rarely needs a frontier model.4. Choose telephony deliberately
Telephony is billed by the minute and varies sharply by country and provider. If you have volume, bringing your own account or SIP trunk is usually cheaper than a bundled rate. Also stop paying for calls nobody answers: tune retry counts and intervals rather than retrying blindly — Retry calls that did not connect.Measure, then cut
Estimates mislead; executions do not. Pull recent calls withGET /v2/agent/{agent_id}/executions, read total_cost and conversation_duration once the status reaches completed, and group by agent.
Track three numbers per agent: median cost per call, median duration, and cost per successful outcome — the third is the one worth optimizing. An agent that costs 20% more per call but converts twice as often is the cheaper agent.
Next steps
Call pricing
How a call is priced
Preferred models
What the flat rate covers
Plan concurrency
Capacity, separately from cost

