Building a Hinglish AI chatbot: lessons for India

Published 4 min read

A Hinglish AI chatbot has to do something most chatbots are never tested on: follow a person who switches between Hindi and English in the middle of a sentence, types Hindi in Roman letters, and expects a reply in the same easy mix. For a large share of Indian phone users this is simply how they write. A bot that answers “Kal ka plan kya hai?” with three paragraphs of formal English feels like it was not listening.

We build consumer apps for people in India. QuesMo is a social app for Indian friend groups, with an AI friend called Mo inside it. Many of our product videos are in Hinglish. Dukaanly, our local-commerce app in development, works in English, Hindi and Punjabi. These are the lessons we apply when a product needs to talk the way its users do.

Match the register as well as the language

The most important rule is to reply in the register the person used. That covers three separate choices.

  • Language mix. If the message is mostly Hindi with English nouns, the reply should be too.
  • Script. Roman Hindi gets Roman Hindi back. Devanagari gets Devanagari.
  • Tone. “Yaar” and “bhai” signal a casual chat. A reply in textbook Hindi breaks the mood as badly as one in formal English.

Write this into the system prompt as a clear instruction with a few short examples, rather than a vague “respond in the user’s language”. Models follow concrete examples far more reliably than abstract rules. Also say what to do when the message is ambiguous, for example a one-word “ok”. A sensible default is to continue in whatever register the conversation has used so far.

Handle Romanised Hindi carefully

Romanised Hindi has no fixed spelling. “Kya”, “kia” and “kyaa” all mean the same thing, and “main” can be the Hindi word for “I” or the English word. People also shorten freely: “h” for “hai”, “k” for “ok”.

Large models cope with most of this. Smaller and cheaper models cope less well. A recent benchmark, Indi-RomCoM, found that LLMs consistently underperform on Romanised code-mixed instructions, and that performance falls further as the mix gets denser. That matches what anyone testing these products sees: the everyday case is fine, and the heavily mixed, loosely spelled message is where replies go wrong.

Two practical consequences:

  • Do not normalise spelling before the model sees it. Converting everything to Devanagari or “correct” spelling loses the tone and adds a step that can fail.
  • Be careful with routing. If you route requests across models to save cost, check that the cheaper model handles code-mixed input before sending Hinglish traffic its way.

Test with the messages people really send

An eval set written in clean English will tell you nothing about a Hinglish product. Build test cases from the messages real users send, with their spelling and their mix left as they are.

Useful categories to cover:

  1. Pure English, pure Roman Hindi and Devanagari, so you can check script matching.
  2. Heavy code-mixing within one sentence.
  3. Slang and abbreviations.
  4. Messages where “main”, “to” or “hi” could be read as either language.
  5. Emotional messages. People often switch into their first language when upset, and the reply has to follow.

Grade each reply on two things separately: did it understand the message, and did it answer in the right register. A reply can be correct and still fail the second check. Our general approach to building eval sets is in LLM evals for small teams.

Watch the cost of Devanagari

Tokenizers are the part of a language model that splits text into pieces. Many of them were trained mostly on English and Latin script, and they often split Devanagari into more tokens than the same meaning written in English or Roman Hindi. More tokens means higher cost and slower replies.

This affects product design:

  • Measure before you price. Count tokens for real Hindi and Hinglish conversations with your provider’s tools, not an English estimate.
  • Keep long-lived context compact. Memory summaries and system prompts can stay in English even when replies are in Hindi. The model reads English instructions well, and it keeps the cached part of the prompt stable. We explain why that matters for cost in our routing note.
  • Set reply length in the prompt. Short replies help everywhere, and they matter more when each word costs more.

Voice adds another layer

Voice input in Hinglish is harder than text. Speech recognition has to decide which language each word is in, and names, places and brand words trip it up. If your product takes voice, test with speakers from different regions, and show the transcript so people can correct it. Expect accents and background noise from real life, not a studio.

Be honest about what the AI is

This applies in every language, but it matters more when the AI sounds like a friend. When a bot talks in relaxed Hinglish, people treat it like a person faster. Mo is always labelled as an AI in QuesMo, and we think every companion-style product should do the same. We cover this in more detail in how to design a responsible AI companion app.

The takeaway

Hinglish is not an edge case to handle later. For many Indian users it is the default. Match the register, leave their spelling alone, test with real messages, measure the token cost, and label the AI clearly.

If you are building a product for Indian users and want help with the conversational layer, get in touch.

FAQ

Can current LLMs understand Hinglish?
The large models understand everyday Hinglish well. Quality drops with heavy code-mixing, with Romanised spelling that varies a lot, and on smaller models, so test on the kind of messages your users actually send.
Should a Hinglish chatbot reply in Roman script or Devanagari?
Reply in the script the user wrote in. Most casual Hinglish on phones is typed in Roman script, and a sudden switch to Devanagari feels formal and is slower to read for many people.
Does Hindi cost more tokens than English?
Often, yes. Many tokenizers split Devanagari into more tokens than the same meaning in English or Romanised Hindi. Measure it on your own prompts with your provider's token counter before you set prices or limits.

Work with usNeed something like this built? Say hello

Related posts