HomeFeaturesUse CasesBlogsDocs

Highlights

  • By Peush Bery, Xtreme Gen AI
  • Highlights
  • First ask: what exactly costs ₹2?
  • A worked example: how the production stack grows
  • The layers that disappear from headline comparisons
  • Why international self-serve rates confuse Indian buyers
  • The cheapest stack can be the right stack
  • Self-serve versus managed: include the missing payroll
  • Rebuild every quote into one comparison sheet
  • Reconcile the pilot invoice before approving scale
  • Research references
  • Try the Voice AI Agent
  • Conclusion
How a ₹2 Voice AI Call Becomes ₹6
A practical breakdown of the platform, telephony, model, tool and managed-service costs hidden behind Voice AI minute rates.

The ₹2 Voice AI Call That Becomes ₹6 in Production

By Peush Bery

Published: September 5, 2026

By Peush Bery, Xtreme Gen AI

A buyer sees “Voice AI from ₹2 per minute” and builds a budget around it. The pilot begins. Telephony is separate. The preferred voice costs more. The faster model is extra. Transfers, phone numbers, concurrency and campaign support sit elsewhere. By production, the effective cost is closer to ₹6. Was the first price dishonest?

Sometimes it was incomplete marketing. Sometimes it was a legitimate base layer quoted without the buyer asking what the layer included. Voice AI is not one API. It is a live system combining a phone network, audio transport, speech recognition, reasoning, speech generation, business tools, observability and people who keep the workflow working.

This article is not an allegation that every ₹2 quote becomes ₹6. It is a buyer’s method for reconstructing the deployed price before signing. The figures below are illustrative because provider rates, exchange rates, commitments, models, languages and Indian carrier agreements change.

Highlights

A headline minute can describe only the orchestration layer, not the complete deployed agent.

Telephony, STT, LLM, TTS, tools, transfers, number rental, concurrency and support may follow different meters.

Self-serve vendors can be transparent while still producing multiple invoices and internal engineering cost.

Managed pricing is often higher because implementation, QA, retries, reporting and ongoing ownership are part of the product.

Compare the effective cost per reliable outcome at your expected call mix, not the cheapest visible rate.

First ask: what exactly costs ₹2?

The phrase “₹2 per minute” is meaningless until the unit is defined. It might cover platform orchestration only. It might include a standard STT, model and TTS bundle but exclude telephony. It might assume large committed volume, pooled seconds, one language, no transfers and customer-owned API keys. It might be a launch promotion rather than an enterprise quote.

Vapi’s documentation is unusually useful for understanding the platform model: it describes a platform fee and provider expenses billed at cost or through buyer-owned keys. Bolna similarly documents separate Voice AI components and telephony. Bland publishes a more bundled minute rate. These are different packaging strategies, not directly comparable price stickers.

A worked example: how the production stack grows

Illustrative layer: Platform or orchestration Example cost per connected minute: ₹2.00 Why it changes: Base workflow runtime and vendor packaging

Illustrative layer: Telephony and number infrastructure Example cost per connected minute: ₹0.70 Why it changes: Destination, carrier, mobile or landline route and pulse

Illustrative layer: STT, LLM and TTS upgrade Example cost per connected minute: ₹1.10 Why it changes: Language, model tier, voice quality and usage shape

Illustrative layer: Tools, retrieval and post-call processing Example cost per connected minute: ₹0.45 Why it changes: CRM calls, knowledge lookup, summaries and analytics

Illustrative layer: Campaign, QA, monitoring and managed ownership Example cost per connected minute: ₹1.25 Why it changes: People and systems that launch and improve the workflow

Illustrative layer: Number, concurrency and support allocation Example cost per connected minute: ₹0.50 Why it changes: Capacity, phone number, SLA and account overhead

Illustrative layer: Illustrative deployed cost Example cost per connected minute: ₹6.00 Why it changes: Complete operating model rather than a base API

This example is deliberately simple. Real bills do not rise in a perfectly linear manner. A one-minute call with a long customer monologue can use more STT but less TTS. A short qualification call can trigger several CRM, calendar and WhatsApp actions. Monthly number rental and concurrency are fixed or capacity costs that become cheaper per minute only when utilisation improves.

Currency also matters. International APIs commonly bill in dollars while Indian clients pay in rupees. Exchange movement, card or payment costs and taxes can affect the domestic cost base even if the underlying model price does not change.

The layers that disappear from headline comparisons

Layer: Telephony Common meter: Connected minute, pulse or route Question to ask: Are ringing, voicemail and transfers treated differently?

Layer: STT Common meter: Audio seconds or minutes Question to ask: Which language/model and is streaming included?

Layer: LLM Common meter: Input/output tokens or bundled time Question to ask: Does prompt, history and tool output count?

Layer: TTS Common meter: Characters, seconds or bundled time Question to ask: Which voice and is caching used?

Layer: Platform Common meter: Per minute or monthly fee Question to ask: Is this charged on top of every provider?

Layer: Tools and retrieval Common meter: Request, token or compute usage Question to ask: Are CRM, search and post-call jobs included?

Layer: Phone number and concurrency Common meter: Monthly rental or reserved capacity Question to ask: How many simultaneous calls are included?

Layer: Managed operations Common meter: Setup, retainer or bundled margin Question to ask: Who owns prompts, QA, retry logic and changes?

The mistake is adding only the variable APIs. Production also needs an operating layer: prompts versioned by use case, tool schemas, failure handling, campaign rules, consent and suppression logic, dashboards, transcript review and someone accountable when a CRM field changes.

In India, the apparent premium for managed service may partly be the cost of replacing internal product, engineering, telephony and operations work. That does not make every managed quote good. It means the comparison must put equivalent responsibility on both sides.

Why international self-serve rates confuse Indian buyers

A dollar-denominated platform price often excludes the same providers that make the call possible. The buyer may need separate accounts for telephony, STT, LLM and TTS. A published international rate can therefore be lower than the invoice total before currency conversion and internal labour.

Bland’s published rates are visibly higher than many Indian headline prices but are presented as bundled. Vapi publishes a platform fee with provider costs layered on. Bolna publishes platform and component mechanics. None should be converted into rupees and compared with an Indian managed quote until inclusion, routing, support and volume assumptions match.

Indian domestic pricing can be lower because of carrier arrangements, local cost structures, negotiated model rates or a leaner default stack. It can also be subsidised for pilots or tied to volume commitments. Ask whether the quote is sustainable after the pilot, because a vendor without margin cannot fund QA, support and product improvement.

The cheapest stack can be the right stack

Not every call needs premium speech and a frontier model. A reminder asking a customer to press or say yes can use a constrained flow, economical voice and limited reasoning. Paying for maximum expressiveness would be wasteful.

A fee discussion with objections, mixed-language speech, policy retrieval and CRM writes needs a different risk budget. A transcription error or unsupported promise can cost more than the saved rupees. Quality should follow consequence, not vendor prestige.

Use case: Simple reminder Reasonable cost posture: Economical constrained stack Do not compromise on: Correct identity, timing and opt-out

Use case: Lead qualification Reasonable cost posture: Balanced speech and reasoning Do not compromise on: Intent capture, disposition and handoff

Use case: Admissions counselling Reasonable cost posture: Higher language and knowledge quality Do not compromise on: Fee accuracy, context and human transfer

Use case: Diagnostic booking Reasonable cost posture: Reliable tools and confirmation Do not compromise on: Patient details, slot and serviceability

Use case: Collections or regulated workflow Reasonable cost posture: Risk-led premium controls Do not compromise on: Policy, consent, logging and escalation

Self-serve versus managed: include the missing payroll

Self-serve can be the best economic choice for a company with a capable AI product owner, engineers, prompt and conversation design, telephony knowledge, QA capacity and time to improve the agent. The software invoice is only part of its cost; the internal team is the other part.

Bolna is a Voice AI platform or self-serve/platform-led option. Vapi offers programmable infrastructure for teams assembling providers and logic. Bland publishes a bundled platform approach. ConvoZen is a conversational AI and customer-engagement platform that buyers may evaluate for wider contact-centre intelligence.

Xtreme Gen AI is a managed Voice AI Agent company. It owns implementation, prompt and tool logic, retries, CRM/API workflows, WhatsApp memory, QA, reporting and ongoing changes. The managed price must fund that responsibility. The buyer should still demand measurable outcomes and cost transparency; managed does not mean unquestioned.

Rebuild every quote into one comparison sheet

Give every shortlisted vendor the same 30-day scenario: call attempts, expected answer rate, connected minutes, average call length, voicemail percentage, languages, transfer rate, concurrency peak, tool calls, numbers, WhatsApp follow-ups, QA volume and required changes. Ask for low, expected and high cases.

Then calculate total vendor invoices, internal labour, setup, minimum commitments, taxes and the cost of failed outcomes. Divide by useful outcomes such as verified qualifications, completed bookings, paid renewals or accurately resolved queries. This prevents a low base rate from winning a comparison it did not actually price.

Reconcile the pilot invoice before approving scale

At the end of the pilot, export call-detail records and group them by outcome: not connected, voicemail, short human answer, completed conversation, transfer and failure. Reconcile each group against telephony time, platform time and provider usage. Differences should have a documented reason rather than being dismissed as “AI cost.”

Compare the quote’s assumed answer rate, duration, language mix, model tier and tool volume with what actually happened. If the ₹2 base was accurate but production traffic required premium multilingual speech and more CRM actions, the change is explainable. If unexplained fees appeared, the commercial model was incomplete.

Scale only after agreeing which variables can move automatically, which require approval and which are included in the managed rate. A sustainable contract gives the vendor room to operate while giving the buyer an audit trail and budget guardrails.

Research references

Vapi: platform and provider pricing structure

Vapi: how provider cost routing works

Bolna: platform, model and telephony billing mechanics

Bland AI: published bundled pricing

Twilio: example of voice feature charges beyond base calling

Try the Voice AI Agent

To experience the Voice AI Agent directly, call +91 22 6595 2901 from your mobile. While listening, ask whether the agent is only speaking cheaply or actually producing a useful business outcome.

Conclusion

A ₹2 price can be real and still not be the production price. A ₹6 managed price can be justified and still require scrutiny. The only responsible comparison is the complete stack, complete ownership model and complete workload at your expected volume.

Do not punish vendors for making a sustainable margin. Punish opacity. Ask what runs, who owns it, how it is metered and what happens when it fails. The number that matters is not the smallest rate on a pricing page; it is the cost of producing a reliable business outcome repeatedly.

Frequently Asked Questions

1. What is normally excluded from a low Voice AI per-minute price in India?

A low advertised rate may exclude telephony, premium STT or TTS, LLM tokens, phone-number rental, concurrency, transfers, tool calls, retrieval, post-call analytics, setup, QA and ongoing workflow changes. Exclusions differ by vendor. Ask for a component list and a 30-day invoice simulation using your expected call mix.

2. How can a company verify whether a ₹2 Voice AI quote will remain ₹2 in production?

Provide expected attempts, connected minutes, average duration, voicemail rate, languages, models, voice, transfers, concurrency, tool usage, number type, support and change frequency. Require low, expected and high cost scenarios. Also include internal engineers and operations staff for self-serve options. The resulting total divided by reliable outcomes is the comparable figure.

3. Why can managed Voice AI cost more than a self-serve platform?

Managed service can include use-case design, prompt and tool logic, telephony setup, CRM/API integration, retry policy, WhatsApp continuity, QA, monitoring, reporting and ongoing optimisation. A self-serve platform transfers much of that work to the buyer. The premium is justified only when responsibilities and performance are explicit and measurable.

4. When is a lower-quality Voice AI stack the financially correct choice?

It can be correct for narrow, low-risk workflows such as reminders, simple confirmations or routing where limited vocabulary and deterministic rules are sufficient. Higher-quality speech, reasoning and verification are more important when calls involve objections, mixed languages, fees, healthcare details, transactions or tool-driven commitments.

5. What metric should replace headline Voice AI cost per minute?

Use cost per reliable outcome supported by operational measures: human-answer rate, task completion, verified tool success, disposition accuracy, transfer rate, repeat-call rate and human corrections. Keep per-minute cost as an invoice control, but do not use it alone to select a production system.