What an AI voice agent actually costs to run
Three meters run at once and vendors bundle them differently, which is why per-minute comparisons rarely compare the same thing. How to price a ninety-second call.
Ask three vendors what a call costs and you will get three per-minute numbers that are not measuring the same thing. Underneath, three meters run at once, and vendors bundle them differently depending on which one they would rather you did not look at.
The three meters
- 1. Telephony
- What the carrier charges to carry the call. Depends on the destination, whether it is mobile or landline, and which of the three kinds of line you connected — see SIP versus cloud telephony versus a carrier number. This is the meter you have the most control over, because it is the one you can renegotiate.
- 2. Speech
- Turning audio into text and text back into audio. Priced by audio duration. Better voices cost more, which is why voice quality is a plan tier almost everywhere including here.
- 3. The model
- Priced by tokens, not minutes. A caller who says "yes" costs almost nothing; one who explains their situation for ninety seconds costs meaningfully more, and a long system prompt is re-processed on every turn.
Why per-minute quotes mislead
Only the first two meters scale with time. The third scales with how much is said, and it is affected by things a quote never mentions:
- Prompt length. A 2,000-word prompt is paid for on every turn of every call.
- Knowledge retrieval. Each retrieved document is more input.
- Model choice. Often a several-fold difference for the same conversation.
- Interruptions. Barge-in means partial turns processed and discarded.
- Language. Some languages tokenise far less efficiently than English.
So the honest question is not "what is your per-minute rate". It is: what does a ninety-second call cost all-in, on the model I would actually use, in the language my customers speak? Ask for that number and watch how long it takes to arrive.
The costs nobody quotes
- Failed and abandoned calls. Someone who hangs up after four seconds still cost you a connection.
- Testing. Every prompt change should be tested against real calls, and those calls are billed.
- Retries. An outbound campaign's cost is attempts, not contacts. If your sequence tries four times, budget four calls — and pick the number deliberately.
- The human on the other end of the handover. An agent that hands over 40% of calls has not removed 40% of the work.
Compare it against the right baseline
The comparison is rarely "agent versus human". It is usually:
- Versus nothing. After-hours and overflow calls that currently ring out. The baseline is a lost enquiry, and almost any cost beats it.
- Versus a person's marginal hour. Not their salary divided by hours — the value of what they stop doing to make the call.
- Versus not following up at all. Which is what most teams actually do by the third attempt.
How to keep it down without making it worse
- Shorten the prompt. Move facts to the knowledge base, where they are retrieved only when relevant rather than paid for every turn.
- Cap the conversation. A call that has gone twelve turns is not going to be rescued by turn thirteen. Hand over.
- Match the model to the job. A hours-and-directions agent does not need your most capable model.
- Fix the retry schedule. Most campaigns are configured to try more times than the data supports.
- Watch the carrier line. At volume this is usually the biggest of the three meters, and the only one with a human you can negotiate with.
In Convarza, calling minutes and AI messages are included in the plan — 600 minutes a month on Starter through 20,000 on Scaler — and your carrier account stays yours, so you can see both halves rather than one blended number. Plans.