Pricing models for voice agents
In one sentence
Pricing models for voice agents are the ways a vendor charges, by the minute, by the conversation, by the outcome or by seat, and the model chosen shapes vendor and customer incentives more than the rate does.
Not to be confused with Unit economics of voice agents.
Definition
A pricing model is how a voice agent vendor charges you: per minute, per conversation, per resolved outcome, or per seat.
The model matters more than the rate. It decides whether the vendor is rewarded for the same thing you want or for the opposite of it.
Voice carries a real variable cost for every interaction, which sets it apart from most software and makes flat-rate pricing risky. The models in common use each solve for that cost differently, and each one lines up vendor and customer incentives in a different way.
Per minute
- The dominant model, inherited straight from telephony.
- It tracks the real cost driver, since speech recognition, model inference and synthesis all scale with how long a conversation runs.
- The perverse incentive is real. A vendor billing by the minute earns more when conversations drag on, which is the opposite of what a customer wants, and sophisticated buyers notice.
- Published examples: Vapi around USD 0.05 per minute, Retell around USD 0.07, Bland around USD 0.09. In India, Bolna prices from around Rs 7 per minute falling to around Rs 3 at volume.
Per conversation
- A fixed charge per session, whatever its length.
- The incentive aligns better, because the vendor now benefits from resolving things efficiently.
- It moves duration risk onto the vendor, so the price has to assume a distribution and long conversations eat the margin.
- It needs a clear definition of what counts as one conversation, which is less obvious than it sounds.
Per resolution or outcome
- Charges only for a successfully resolved interaction, or for a defined outcome such as a booking or a qualified lead.
- The strongest incentive alignment, and the most attractive to buyers.
- It lives or dies on an agreed definition of resolution. A vendor and a customer disagreeing about what counts as resolved is a contract dispute, not a product problem.
- It reaches a different budget. Outcome pricing draws on performance marketing or operations budgets rather than software budgets, which are usually far larger.
Per seat
- Charges per human user, borrowed from SaaS convention.
- Poorly matched to voice agents, where value scales with conversation volume and not with how many people log in.
- It shows up sensibly as a hybrid: seat pricing for the dashboard and management layer, usage pricing for the agent itself.
What a buyer should actually compare
- The all-in cost per resolved interaction, not the headline rate. A lower per-minute rate on a system that takes longer to resolve is the more expensive one.
- Whether telephony, speech recognition, the model and synthesis are included or passed through separately.
- Overage rates, which is where the quoted price and the actual bill part company.
- Minute pools and committed volume sit alongside all of this: prepaid blocks, often tiered, that give the vendor predictable revenue and the customer a cap that bounds the cost of a traffic spike.
Common misconception
That per-minute pricing is simply the norm and therefore neutral. It is inherited from telephony and it misaligns incentives. The direction of travel is towards conversation and outcome-based models, and a vendor that moves first on this has a genuine commercial argument.
Why it matters commercially
The pricing model is a positioning decision, not a rate card. Outcome pricing in particular changes which budget a purchase comes from, and a performance budget is typically an order of magnitude larger than a software budget for the same buyer.
Where AsqVox fits
AsqVox runs in the browser rather than over a phone line, so its cost stack carries no telephony charge, the cost line that per-minute telephony pricing was built to cover. Pricing structure itself is a commercial decision rather than a product feature.
Visual
Four models, four sets of incentives
The model shapes behavior. The rate is just a number.
Statistics
Every figure carries its source and year. Vendor numbers are labelled as vendor numbers, and where no reliable figure exists this page says so rather than borrowing one.
Published platform rates start around USD 0.05 per minute for Vapi, USD 0.07 for Retell and USD 0.09 for Bland.
USD 0.05 to 0.09 per minvendor claimVendor published pricing (Vapi, Retell, Bland), 2026 - Base platform rates only. Vendor pricing changes often, so treat these as indicative and check the current pricing page before quoting them.
All-in production cost, once speech recognition, the model, synthesis and telephony are stacked, lands around USD 0.11 to USD 0.33 per minute.
USD 0.11 to 0.33 per minindustry rangeIndustry-reported production cost range, 2026 - The floor any pricing model has to clear. A headline per-minute rate below this is the platform layer only, not the delivered cost.
In India, Bolna prices from around Rs 7 per minute falling to around Rs 3 per minute at volume.
Rs 7 to Rs 3 per minvendor claimBolna published pricing, 2026 - Published by the vendor, following a USD 6.3 million seed round led by General Catalyst in January 2026. Quote it as a volume-tiered range, not a single rate.
A human-handled call costs roughly USD 7 to USD 12 against roughly USD 0.40 for an agent-handled call.
USD 7 to 12 vs USD 0.40industry rangeWidely cited industry range, 2025 - Not an audited figure, and it varies substantially by geography and call complexity. Treat the ratio as directional rather than the absolute numbers as precise.
There is no published data on the distribution of pricing models across the voice agent market, so any claim about which model is becoming dominant is impressionistic.
-no reliable figureThe direction of travel towards outcome pricing is an argument, not a measurement. State it as a trend rather than a market share.
There is no benchmark for typical conversation duration in website voice contexts, which is the input a buyer needs to turn a per-minute rate into a per-conversation cost.
-no reliable figureThis is why a per-minute headline is hard to compare across vendors. Measure duration on your own traffic to make the conversion.
Vendor pricing changes frequently, so any specific rate needs a date attached and a link to the current pricing page, or it is already going stale.
-no reliable figureA caveat rather than a figure. A rate quoted with no date is a number with an unknown shelf life.
Examples
In practice
A buyer compares two vendors on headline per-minute rate and picks the cheaper one. Measured over a quarter, the cheaper agent takes materially longer to resolve the same queries, and the total cost per resolved interaction comes out higher. The correct comparison was never the rate. It was cost per resolution, which means running both systems on the same query set.
The everyday version
Most voice agent companies charge by the minute, which is how phone companies have always billed. The odd thing about that is it means they earn more when the conversation drags on, which is not what you want. Some charge per conversation instead, and a few charge only when something actually gets resolved. That last one lines up best with your interests, and it is the hardest to pin down.
Usage
Who says it
- Vendors and buyers, in commercial negotiation.
- Finance teams, in cost modeling.
- Investors, when they are assessing revenue quality and gross margin.
Where it turns up
- In an RFP, next to commercial terms, volume commitments, overage rates, what is included against passed through, and price review mechanisms.
- Outcome-based arrangements carry a definitions clause specifying resolution, which is the most important paragraph in that contract.
Common misuse
- Comparing headline per-minute rates without normalizing for conversation duration.
- Treating per-minute pricing as neutral rather than as incentive-misaligned.
- Offering outcome pricing with no rigorous, agreed definition of resolution, which produces disputes instead of alignment.
Questions people ask
How do voice agents charge for usage?
Four models are in common use: per minute, per conversation, per resolved outcome, and per seat. Per minute is the dominant one, inherited from telephony. The model matters more than the rate, because it decides whether the vendor is rewarded for the same thing you want, a fast resolution, or for the opposite, a long call.
How much does a voice agent cost per minute?
Published platform rates start around USD 0.05 per minute for Vapi, USD 0.07 for Retell and USD 0.09 for Bland. Those are base rates only. Once speech recognition, the model, synthesis and telephony are stacked, all-in production cost lands around USD 0.11 to USD 0.33 per minute, so any pricing model has to clear that floor.
What is outcome-based pricing for voice agents?
It charges only for a successfully resolved interaction, or for a defined outcome such as a booking or a qualified lead. It gives the strongest incentive alignment and is the most attractive to buyers, but it depends entirely on an agreed definition of resolution. It also reaches a different budget, drawing on performance or operations spend rather than the software budget, which is usually far larger.
Should I pick a voice agent vendor on the lowest per-minute rate?
No, because a lower per-minute rate on a system that takes longer to resolve is the more expensive one overall. Compare the all-in cost per resolved interaction instead, check what is included against what is passed through, and read the overage rate. That requires measuring both systems on the same set of queries.
Last reviewed 31 July 2026. Written and reviewed by Dhruv Dholakia, founder of AsqVox.