Skip to main content
Voice agents fail in ways text agents do not: a prompt that reads perfectly can produce an agent that talks over people, mishears digits, or refuses to hang up. Test on audio, not on paper.

The four stages

1

Test in the browser

Fastest loop for prompt changes — no telephony, no cost per attempt. Use the Web Call SDK or the test call in the dashboard. Good for wording; useless for audio quality.
2

Call yourself on a real phone

Everything that matters — carrier audio, endpointing under real latency, background noise, DTMF — only appears on a real call. Do this on the same telephony provider you will use in production.
3

Run the awkward-case list

See below. These are the cases that generate support tickets in week one.
4

Dry-run a small batch

Send a batch of 20–50 calls to internal numbers or consenting testers. Batch behavior — scheduling, guardrails, retries, concurrency — is not exercised by single calls.

The awkward-case list

Keep the list as a fixed script and re-run it after every prompt or model change. Regression, not first-run behavior, is what breaks live agents.

What to read after each test call

Open the execution in Call history and check four things:
  1. Transcript — did the transcriber hear correctly? Mistakes here explain most odd replies.
  2. Recording — how did the pacing actually feel? Read time is not talk time.
  3. Latency metrics — where did the time go, per turn — Read latency metrics for a call.
  4. Extracted data — did the fields your business depends on come out populated and correctly typed?

Testing graph agents and workflows

Graph agents have their own testing, validation and debugging tools — use them to verify every transition is reachable before making any calls. Validation catches structural faults; only real calls catch conversational ones.

Next steps

Production readiness checklist

Everything else before launch

Failures and retries at scale

What happens when calls fail

Plan concurrency

Before the first big campaign