> ## Documentation Index
> Fetch the complete documentation index at: https://www.bolna.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Control cost per call on Bolna voice AI

> Reduce Bolna Voice AI spend without hurting quality: stay on preferred models, shorten calls, cut token usage, choose telephony carefully, and measure real cost per call.

Cost per call is driven by four things: which models you run, how long the call lasts, how much the LLM generates, and which carrier delivers it. Attack them in that order.

## 1. Stay inside the flat rate

Your workspace has a flat per-minute rate that covers transcription, LLM and voice **as long as the agent uses [preferred models](/docs/pricing/preferred-models)**. One non-preferred component moves that component onto variable, usage-based pricing — the single most common reason two agents with identical volume bill differently.

<Steps>
  <Step title="Audit the three model choices on every production agent">
    Check the transcription and voice models on the [Languages tab](/docs/agent-setup/audio-tab) and the LLM on the [Intelligence tab](/docs/agent-setup/llm-tab) against the current preferred list.
  </Step>

  <Step title="Justify each exception">
    Keep a premium model only where a test showed it materially improves outcomes — accuracy in a specific Indian language, for instance. Otherwise revert.
  </Step>
</Steps>

## 2. Shorten the call

Every component except the LLM bills by duration, so seconds are money.

* Cap total duration with `call_terminate`, and end dead air with silence-based hangup — [End a call automatically](/docs/outbound/hangup-calls)
* Cut the greeting to one sentence; long welcome messages are billed at both TTS and telephony rates
* Instruct the agent to answer in one or two sentences and ask one question at a time — [Write better prompts](/docs/prompting/introduction)
* Detect voicemail so the agent does not deliver a full pitch to an answering machine — [Call tab](/docs/agent-setup/call-tab)

## 3. Cut tokens

The LLM bills by tokens generated. Trim the system prompt, avoid restating context the agent already has, keep [knowledge base](/docs/knowledge-base) documents scoped to what is actually needed, and prefer a smaller preferred model for simple flows — a reminder or confirmation agent rarely needs a frontier model.

## 4. Choose telephony deliberately

Telephony is billed by the minute and varies sharply by country and provider. If you have volume, [bringing your own account or SIP trunk](/docs/supported-telephony-providers) is usually cheaper than a bundled rate. Also stop paying for calls nobody answers: tune retry counts and intervals rather than retrying blindly — [Retry calls that did not connect](/docs/outbound/auto-retry).

***

## Measure, then cut

Estimates mislead; executions do not. Pull recent calls with [`GET /v2/agent/{agent_id}/executions`](/docs/api-reference/executions/get_executions), read `total_cost` and `conversation_duration` once the status reaches `completed`, and group by agent.

<Warning>
  Cost and duration are only populated once the execution reaches **`completed`**. Reading them at `call-disconnected` returns zeros and will understate your spend.
</Warning>

Track three numbers per agent: median cost per call, median duration, and cost per *successful outcome* — the third is the one worth optimizing. An agent that costs 20% more per call but converts twice as often is the cheaper agent.

***

## Next steps

<CardGroup cols={3}>
  <Card title="Call pricing" icon="magnifying-glass-dollar" href="/docs/pricing/call-pricing">
    How a call is priced
  </Card>

  <Card title="Preferred models" icon="list-check" href="/docs/pricing/preferred-models">
    What the flat rate covers
  </Card>

  <Card title="Plan concurrency" icon="scale-balanced" href="/docs/enterprise/concurrency-management">
    Capacity, separately from cost
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.