gpt-5.4-mini remains the default recommendation for most voice agents: it has low time-to-first-token and strong instruction following at a fraction of the cost of the full models.
Quick config
Supported models
Recommendation: Start with
gpt-5.4-mini. Step up to gpt-5.6-terra or gpt-5.6-luna for newest-generation quality at moderate cost, or gpt-5.6-sol / gpt-5.5 when you need the strongest multi-step reasoning and highest output quality (financial, medical, nuanced escalation).
Key settings
Keep max_tokens short
Voice responses should be 1–3 sentences.max_tokens: 150 is appropriate for most turns. A higher cap doesn’t hurt quality but increases tail latency on long responses.
On GPT-5 models max_tokens is sent as max_completion_tokens and reasoning tokens come out of the same budget. At reasoning_effort above none/minimal, reasoning can consume most of a 150-token cap and truncate the spoken reply, so raise the cap whenever you raise the effort.
Reasoning effort
GPT-5 models reason before answering. Effort controls how much, and it is the main quality-versus-latency dial on the LLM leg of a call. Leave it unset and the model gets the lowest-latency effort it supports, which is what most voice agents want. Every model accepts a different subset, and an unsupported value is rejected when the agent is created:
For live calls, stay at
none or low. Each step up adds reasoning tokens before the first spoken word, which lands directly in time-to-first-token. See Latency.
Writing prompts for voice
Prompts for voice agents differ from chat prompts:- Use imperative sentences: “Keep all responses under 3 sentences.”
- Specify spoken format: “Never use bullet points or markdown — speak in complete sentences.”
- Define handling for off-topic questions: “If asked something outside your scope, say: ‘I can only help with appointment scheduling today.’”
- Include the welcome message in the prompt or agent config, not as part of the system prompt instructions.
Function calling
All GPT-5 and GPT-4.1 models support function calling. In Bolna, functions are defined in the Tools Tab and called automatically by the LLM during conversation.gpt-5.4, gpt-5.5 and gpt-5.6 run through OpenAI’s Responses API automatically, because function calling combined with reasoning_effort is not accepted on chat completions for those models. You don’t need to configure anything for this.
See Custom Function Calls for configuration.
FAQ
Which GPT model should I use?
Which GPT model should I use?
Use
gpt-5.4-mini for most agents: lowest time-to-first-token and significantly lower cost per call. Step up to gpt-5.6-terra or gpt-5.6-luna for newest-generation quality at moderate cost, or gpt-5.6-sol / gpt-5.5 for the most demanding tasks (complex financial/medical, long multi-step tool chains) where quality is the top priority. Avoid gpt-5.5-pro for live voice; its latency is too high for real-time calls.Can I lower temperature to make a GPT-5 agent more consistent?
Can I lower temperature to make a GPT-5 agent more consistent?
No. GPT-5-series models accept only
temperature: 1, and anything else fails agent creation with a 400. Use the prompt to constrain behaviour instead: state the exact wording, the sentence limit, and the fallback line for off-topic questions. On the previous-generation GPT-4.1 models a lower temperature still applies.How do I reduce latency?
How do I reduce latency?
Keep
reasoning_effort at none (or minimal on gpt-5/gpt-5-mini/gpt-5-nano), lower max_tokens, use gpt-5.4-mini instead of the larger models, and write shorter system prompts (large prompts increase prefill time). Reasoning effort is usually the biggest single lever. See Latency for a full breakdown.Can I use my own OpenAI API key?
Can I use my own OpenAI API key?
Yes. Connect your OpenAI account at platform.bolna.ai/auth/openai. Costs will be charged to your OpenAI account, not Bolna’s platform wallet (for the LLM component).
Related
- LLM Tab — configure LLM in the dashboard
- Anthropic Claude — alternative LLM
- Azure OpenAI — OpenAI models with enterprise data residency
- Prompting Guide — write effective prompts for voice
- Custom Function Calls — add tools to your agent

