Every month a new vendor will offer you an "AI voice agent" or a "GenAI chatbot" that promises to handle your customers in Hindi, Tamil and English. Unless you understand the few moving parts inside these systems, you cannot judge a demo, compare quotes or predict where the bot will fail with your real customers. This lesson gives you that working model, so the rest of the programme builds on solid ground.
What you need to know
The pipeline. Every voice bot follows the same chain: the caller speaks, speech-to-text (also called ASR, automatic speech recognition) turns audio into words, an understanding layer works out what the person wants, a decision layer applies your business rules and fetches data, a response layer composes the reply, and text-to-speech (TTS) speaks it back. A chatbot is the same chain without the first and last steps. When a bot fails, it fails at one of these stages, and each stage has a different fix.
Speech-to-text and word error rate. ASR accuracy is measured by word error rate (WER):
- WER = (substitutions + deletions + insertions) ÷ number of words actually spoken
- If a caller says 40 words and the transcript has 3 wrong words, 1 missing word and 2 extra words, WER = (3 + 1 + 2) ÷ 40 = 15%.
Indian conditions push WER up: phone audio is lower quality than a studio microphone, callers switch between Hindi and English in one sentence (code-mixing), and they call from shops, buses and factory floors. Names, addresses, pincodes, order numbers and amounts are the hardest items to capture by voice, which is why good designs collect them another way (keypad digits, or a link sent by SMS or WhatsApp).
Intents and entities. The understanding layer answers two questions. The intent is the goal: "track my order", "book a home visit", "what are your timings". Entities (or slots) are the details needed to act: order number, date, locality, product name. Older bots used a trained classifier: you give 15–30 example phrases per intent and the system learns to map new phrases to the closest intent. It is predictable, but it breaks on anything you did not anticipate.
Large language models (LLMs). Modern assistants use a language model that predicts fluent text. It understands messy phrasing far better than an intent classifier and can reply naturally in many languages. Its weakness is that it can produce confident, wrong answers ("hallucination"), such as inventing a discount or a delivery date. The cure is grounding: the system first searches your approved documents (price list, policies, FAQs), passes the relevant passages to the model and instructs it to answer only from them, or to say it does not know and hand over to a person. You will build that knowledge base in lesson 07.
Rules versus AI. A useful decision rule: let AI understand the customer, but let fixed rules decide anything involving money, commitments or compliance. Refund eligibility, credit limits, prices and legal disclosures should come from your system or a written rule, never from the model's imagination.
The decision layer and systems of record. A bot that only chats is an FAQ page with a voice. Real value comes when it can read your order system, lab system, ERP or CRM: "Your order 4521 was dispatched yesterday by Delhivery." That requires an integration (lesson 15) and customer verification before revealing any personal data.
Text-to-speech. Neural TTS voices now sound natural in Indian English, Hindi and many regional languages. Check three things in any demo: how it pronounces your brand and product names, how it reads amounts (₹12,450 as "twelve thousand four hundred fifty rupees" rather than digit by digit) and how it reads dates and phone numbers. Most platforms let you fix pronunciation with a custom lexicon.
Latency. On a phone call, every stage adds delay. A common rule of thumb is that silences longer than about a second start to feel awkward, and callers begin talking over the bot. Ask vendors for the typical response delay on a real phone line, not on a laptop demo.
Cost drivers. You usually pay per minute (telephony plus ASR plus TTS), per message (chat channels such as the WhatsApp Business Platform), per unit of language-model usage (often called tokens) and a platform subscription. Knowing the pipeline tells you which meter is running at each stage.
Step-by-step method
- Pick five real customer conversations from last week: three calls (listen to recordings or write them from memory) and two chats.
- For each, write the customer's exact words in one column and the intent in plain English in the next.
- List the entities the conversation needed: order number, name, date, product, locality.
- Mark where each entity lives today: in someone's head, a register, Excel, Tally, your order system.
- For each conversation, label the hardest stage: hearing (ASR), understanding (intent), knowing (data), deciding (rule) or speaking (reply).
- Note which parts must be rule-based (price, refund, credit) and which can be AI-understood.
- Estimate a rough cost per automated conversation using the per-minute or per-message assumptions a vendor gives you.
- Write one line per conversation: "Automate now", "Automate after data fix" or "Keep human".
Worked example
Worked example
A diagnostic lab in Pune with three collection centres receives about 350 calls a day. The manager samples 50 calls and tags them. For this example assume the split is: report status 40%, home sample collection booking 25%, test prices 15%, timings and directions 10%, other 10%.
Pipeline analysis:
- Timings and directions (10%): simple intent, no personal data, answer from a fixed text. Hardest stage: none. Automate now.
- Test prices (15%): test names are tricky for ASR ("HbA1c", "lipid profile", "thyroid profile"). The lab adds these terms to the ASR custom vocabulary and reads prices from its rate list, not from the model. Automate after vocabulary work.
- Report status (40%): needs verification (registered mobile plus patient ID) and a lookup in the lab software. Hardest stage: decision layer, because an integration is needed. Automate after integration.
- Home collection booking (25%): addresses are hard to capture by voice. Design choice: the voice bot books the slot, then sends a link by SMS or WhatsApp for the customer to type the address.
Rough cost check, for this example assume telephony ₹0.60 per minute, speech services ₹1.20 per minute and language-model usage ₹0.40 per minute, so ₹2.20 per minute. With an average automated call of 2.5 minutes, cost per automated call = 2.5 × ₹2.20 = ₹5.50.
Current cost: 4 executives at ₹22,000 per month = ₹88,000. Calls per month = 350 × 26 working days = 9,100. If about 70% of their time goes on calls, staff cost per call = (₹88,000 × 0.70) ÷ 9,100 = ₹6.77.
Result: the saving per call is modest; the real gain is answering every call at 7 a.m. peak and freeing executives for report queries that need care. The manager decides to start with timings and prices, and plan the report-status integration next.
Apply it
Template / checklist
Conversation anatomy card (one per conversation):
- Customer's words: ____
- Intent: ____
- Entities needed: __ / / __
- Where each entity lives: head / register / Excel / Tally / order system / other ____
- Hardest stage: hearing / understanding / knowing / deciding / speaking
- Must be rule-based? yes / no — which part: ____
- Needs identity check? yes / no
- Estimated cost per automated conversation: ₹____
- Decision: automate now / after data fix / keep human
Common mistakes
- Judging a voice bot from a laptop demo in a quiet room instead of a real phone call from a noisy shop in your customers' language mix.
- Letting a language model quote prices, discounts or delivery dates from its own "knowledge" instead of your price list or order system.
- Trying to capture addresses, email IDs or long order numbers by voice when a keypad or a typed link would be far more reliable.
- Ignoring code-mixed speech: testing only pure English or pure Hindi when real callers say "mera order kab deliver hoga".
- Comparing vendors only on subscription price and forgetting per-minute, per-message and model-usage meters.
- Building a talking FAQ with no connection to your data, then wondering why customers still call your staff.
Apply it
20-minute action task
Complete five conversation anatomy cards from last week's real calls and chats. Output: a one-page table with intent, entities, hardest stage and your automate / fix / keep-human decision for each.
Ask the AI Business Tutor
- "I run a [type of business] in [city] with [number] calls and [number] chats a day, mostly in [languages]. Here are five real customer conversations: [paste]. For each, identify the intent, the entities needed, the hardest pipeline stage (speech-to-text, understanding, data, rule or reply) and whether I should automate now, after a data fix, or keep it human. Explain your reasoning."