Highlights
- By Peush Bery, Xtreme Gen AI
- Highlights
- First separate the billing clock from the pricing model
- The five-second-call test
- How public platforms currently count time
- Why call cost is not linear throughout the minute
- Why global prices can look expensive from India
- Why Indian managed pricing can look deceptively high or low
- The layers that create the final invoice
- Mobile number, landline number and channel capacity are not cosmetic
- Would per-second billing be fairer?
- Would per-outcome pricing align incentives better?
- A practical hybrid model for India
- How Xtreme Gen AI explains its per-minute model
- Do not compare Voice AI with one human salary
- The invoice simulation every buyer should request
- What should the board measure?
- Try the Voice AI Agent
- Conclusion

Should Voice AI Charge Per Second, Per Minute or Per Outcome?
By Peush Bery
Published: August 26, 2026
By Peush Bery, Xtreme Gen AI
An Indian buyer asks for the Voice AI price. One vendor says ₹3 per minute. Another says it bills to the second. A third offers ₹2 per connected minute but adds telephony, model usage, concurrency and implementation. A fourth proposes payment for every qualified lead. All four quotes can be honest, and all four can produce very different invoices.
The confusion begins because three different questions are being mixed together: What unit records consumption? What commercial package pays for the operating stack? And what business metric proves that the spend was worthwhile?
Xtreme Gen AI charges primarily per minute because the underlying call stack consumes resources throughout a live conversation and because a blended minute is practical for Indian operating teams. But per-minute pricing is not automatically the fairest model, just as per-second or outcome pricing is not automatically more transparent. Fairness depends on pulse rules, inclusions, quality, ownership and the denominator used to judge value.
Highlights
Per-second, per-minute and pulse billing describe measurement; they do not reveal everything included in the price.
A five-second connected call can be billed as five seconds, 15 seconds, 30 seconds or a full minute depending on the commercial and telephony rules.
Global self-serve platforms often expose several component charges, while Indian managed quotes may bundle implementation and operations into one rate plus a monthly agent fee.
Outcome pricing sounds aligned, but buyers and vendors must agree on attribution, quality, duplicates, reversals and factors outside the agent’s control.
India should currently judge Voice AI more as a coverage and efficiency layer than as a simplistic replacement for human salary.
First separate the billing clock from the pricing model
A billing clock answers how duration is counted. A pricing model answers what the customer pays for. These are related but not identical. A vendor can aggregate exact seconds and still present an effective per-minute rate at month-end. Another can use a 30-second pulse, meaning every connected call is rounded to blocks of 30 seconds. A managed provider can use a minute as the commercial unit while negotiating the pulse for enterprise volume.
This distinction matters because “₹3 per minute” is incomplete without the pulse. The same rate applied to one long call and hundreds of short connected calls produces different effective economics.
The five-second-call test
Assume 1,000 calls connect for only five seconds because the customer says hello and disconnects, an answering system picks up, or the number reaches an inconclusive short call. The exact connected time is 5,000 seconds, or 83.33 minutes.
This does not prove that exact-second billing is always cheapest. The exact-second platform may have a higher platform fee, international telephony, premium voice, LLM and add-ons. The full-minute vendor may include Indian telephony, a maintained agent, reporting, QA and support. The test simply shows why pulse must be visible before rates are compared.
How public platforms currently count time
https://www.retellai.com/pricing. Retell AI’s pricing page says each call is tracked to the nearest second and total accumulated minutes are charged at the end of the billing cycle, without rounding up each call. It also says connected silence is billable because the speech-recognition layer remains active. This is a clean example of second-level measurement presented through per-minute component rates.
Bolna’s public pricing pagedescribes a pilot plan billed in 30-second pulses. Bolna is an India-oriented Voice AI platform offering pay-as-you-go, pilot and enterprise options, so its commercial model can differ by volume and agreement.
Telephony can bring its own clock. Tata Tele Business Services, for example, publicly shows a 30-second pulse for one toll-free plan and 60-second pulses for others. Plivo’s India voice pricing states a 30-second pulse. These examples are not universal rate cards for Voice AI; they show that the carrier layer may already round consumption before AI pricing is added.
Why call cost is not linear throughout the minute
A live Voice AI call does not consume every component in a perfectly flat way. Telephony usually follows connected duration. Streaming speech recognition listens while the line is active. Voice generation depends on how much the agent speaks. LLM cost depends on prompt size, conversation history, tool definitions, number of model requests and selected model. Tool calls may happen only at specific moments. Post-call summaries and evaluations happen after the conversation.
Vapi’s component-cost documentation explains this clearly: transcription is estimated by audio minutes, voice by spoken characters and the model by input and output tokens. It notes that long calls can increase model cost because growing conversation history is repeatedly sent, while prompt caching can reduce effective input cost.
A 60-second call with a short greeting and immediate opt-out is therefore not economically identical to a 60-second call containing long answers, several reasoning turns, knowledge retrieval and a booking tool call. A blended minute averages those differences across traffic. It is simple for budgeting, but the vendor carries mix risk.
Why global prices can look expensive from India
Indian buyers often convert a US-dollar rate into rupees and conclude that the international platform is expensive despite using familiar components such as OpenAI, Gemini, Deepgram, ElevenLabs or Cartesia. The observation can be valid, but the comparison needs context.
International developer platforms price orchestration, infrastructure, support, reliability, margin and product development on top of model providers. They may also assume US telephony economics or require custom telephony for India. Retell currently publishes a pay-as-you-go range of $0.07–$0.31 per minute and a component calculator. The visible stack includes voice infrastructure, TTS, LLM, telephony, optional add-ons, phone numbers and paid concurrency above the included allowance.
Bland AI’s billing documentation currently lists plan-based connected-minute rates of $0.14, $0.12 and $0.11, with monthly subscription fees on higher tiers and additional transfer-time rates. The numbers may change, so the important lesson is structural: headline usage can sit beside plan, transfer and scale economics.
Vapi positions itself as a developer-focused modular platform. Its documentation says provider costs for STT, LLM and TTS are charged at cost and a Vapi fee sits on top; buyers can also bring provider keys. This can be transparent and powerful, but it makes the customer responsible for assembling and operating the economic stack.
A global platform still needs sustainable gross margin to maintain realtime infrastructure, orchestration, monitoring, security, product development and support. “Same underlying APIs” does not mean “same delivered product,” just as buying ingredients is not the same as operating a restaurant. The legitimate buyer question is whether the added margin creates enough capability and reduced ownership for the use case.
Why Indian managed pricing can look deceptively high or low
Many Indian buyers are not purchasing only an API. They expect the vendor to discover the workflow, build the agent, maintain prompts, connect CRM, configure retries, provide numbers, monitor calls, produce reporting and keep changing the system after launch. A quote that bundles these services into ₹3–₹7 per minute is not directly comparable with a self-serve platform fee.
The reverse problem also exists. An aggressively low Indian rate may use a lower-cost STT, TTS or LLM, apply a larger billing pulse, limit language quality, exclude telephony, charge separately for setup, restrict concurrency or leave the customer to build and maintain the agent. Lower quality is not automatically wrong. It can be sensible for a simple reminder, confirmation or survey where the downside of repetition is small. It is less sensible for complex qualification, mixed-language objections or high-value customer calls.
The price must fund a sustainable service. A vendor that cannot afford call QA, prompt maintenance, carrier escalation, support and engineering changes may deliver an attractive pilot and a weak six-month operation. Margin is not evidence of overcharging; unexplained margin without delivered ownership is the concern.
The layers that create the final invoice
A low opening rate is not necessarily misleading; it may be only one row. The mistake is stopping the comparison before every row is normalised.
Mobile number, landline number and channel capacity are not cosmetic
Indian calling workflows also need a number and a channel strategy. A mobile-format number may be commercially or operationally different from a virtual landline or toll-free number. Incoming callbacks, SIP routing, transfers, caller-name programmes, Truecaller presence and network-level whitelisting can create separate costs and dependencies.
Concurrency is the number of simultaneous active calls the system can support. Ten thousand minutes spread through a month is different from ten thousand minutes concentrated into a two-hour admissions or payment campaign. Some platforms include a base concurrency allowance and charge for more; Indian telephony providers may package channels separately. The quote should specify both monthly volume and peak capacity.
Xtreme Gen AI can provide mobile or landline calling options, telephony support, Truecaller assistance and telecom-operator whitelisted branded-number support where available and approved. These should be scoped separately from the intelligence layer because number availability, carrier processes and commercial terms can change.
Would per-second billing be fairer?
Per-second aggregation is attractive for short-call-heavy campaigns because it reduces rounding. It is also easy to audit when call records show exact connected duration. But it may create false precision if AI, telephony and managed-service costs are bundled. The buyer sees an exact clock while the vendor still carries fixed costs, non-linear model usage, development and support.
A fair per-second contract should state when charging begins, whether voicemail and silence count, how transfers work, whether every component uses the same duration, and whether monthly minimums or platform fees apply. Precision without definitions is theatre.
Would per-outcome pricing align incentives better?
Outcome pricing can align the vendor with the business when the outcome is objective and largely controlled by the workflow: a confirmed appointment, completed verification, successful callback, qualified lead meeting a written rule or resolved service request. It can be powerful for mature, high-volume processes with reliable data.
It becomes difficult when “outcome” means revenue or conversion. A Voice AI Agent may identify a qualified student, but counsellor quality, programme fit, fee, brand, competition and financing decide admission. It may book home sample collection, but serviceability and phlebotomist availability decide fulfilment. The vendor should not be paid for weak outcomes, but it cannot guarantee factors it does not control.
Outcome contracts also invite arguments about duplicates, existing leads, cancellations, attribution windows, customer fraud, reversals and CRM quality. Vendors may price in this risk, reject difficult leads or optimise for the measured label rather than customer value. The best outcome model needs precise eligibility, independent event data, quality thresholds and shared attribution.
A practical hybrid model for India
For many Indian deployments, the most honest model is hybrid: a modest monthly managed-agent fee for building and maintaining the workflow, transparent usage for connected calling, separately disclosed telephony or number scope, and performance reporting against agreed outcomes. Enterprise volume can negotiate pulse, minimum commitment, concurrency and service levels.
The usage fee recognises that real resources are consumed even when the customer does not convert. The managed fee pays for the team that designs, tests, monitors and changes the agent. The outcome scorecard keeps both parties focused on value without pretending the vendor controls the entire funnel.
How Xtreme Gen AI explains its per-minute model
Xtreme Gen AI primarily charges per minute because telephony, streaming speech, orchestration and model resources operate during the connected call, while a blended rate is straightforward for monthly planning. Depending on volume and enterprise scope, billing pulse and commercial structure can be customised. A monthly managed-agent fee can cover agent development and ongoing maintenance; telephony, number type, channel capacity, integrations and special requirements are scoped transparently.
The important claim is not that every second costs exactly the same. It does not. The per-minute price is an average commercial unit across a variable stack. Xtreme Gen AI then owns the managed layer: prompt and tool logic, bulk or API-triggered calls, retries, customer-requested callbacks, CRM dispositions, reporting, transcripts, summaries, WhatsApp memory and ongoing QA.
That model should be judged against total ownership. A buyer with a strong engineering and Voice AI operations team may prefer Bolna, Vapi or Retell as platform-led routes. A buyer wanting the workflow built and maintained may prefer a managed model. Neither route wins through billing vocabulary alone.
Do not compare Voice AI with one human salary
In India, a human calling team can appear inexpensive when the comparison uses salary alone. The real human operating cost includes recruitment, training, managers, seats, devices, telephony, QA, attrition, idle time, leave, peak capacity and inconsistent follow-up. Yet humans also perform judgment, negotiation, empathy and exception handling that a Voice AI Agent may not replace.
Voice AI should currently be evaluated primarily as an efficiency and coverage layer: faster first response, more disciplined callbacks, extended hours, consistent qualification, cleaner CRM data, recovery of missed demand and better preparation for human closers. The goal is often to make the same human team handle better conversations, not to claim that every automated minute replaces a human minute.
The invoice simulation every buyer should request
Give each vendor the same scenario: monthly lead count, expected attempts, connection rate, duration distribution, number of five-second calls, voicemail rate, callbacks, transfers, language mix, peak concurrency, number requirement, average agent speech, model quality, integrations and QA scope. Ask for a month-one and steady-state invoice.
Then force three cases: a normal month, a campaign spike and a poor-connectivity month with many short calls. Require the vendor to show pulse rounding, every component, monthly minimums, implementation, support and taxes separately. This prevents a low headline from winning because the expensive assumptions were left blank.
What should the board measure?
Finance should track total monthly cost and effective cost per connected minute. Operations should track cost per reliable outcome: qualified lead, completed callback, booked appointment, resolved query, successful reminder or context-rich human handoff. Product and QA should track error severity, tool success, latency and CRM accuracy.
These measures answer different questions. Per minute explains consumption. Per outcome explains efficiency. Revenue influenced explains business contribution. One number cannot replace all three.
Try the Voice AI Agent
To experience the Voice AI Agent directly, call from your mobile. Time the call if you like, but also test whether it understands the request, creates the right next action and leaves useful context. A cheap minute that produces no dependable outcome is not cheap.
Conclusion
Voice AI can be charged by the second, a billing pulse, a connected minute, a subscription, a managed-agent fee or an outcome. Each model can be fair when its boundaries are explicit. Each can become misleading when the buyer sees only the first number.
For India, the immediate discipline is to normalise the entire stack and evaluate Voice AI as an efficiency system. Ask how duration is rounded, what is bundled, which quality level is configured, who builds and maintains the agent, how concurrency and numbers are priced, and what reliable outcome the workflow creates. The right commercial model is the one both parties can audit, sustain and improve.