Voice AI Localisation
Code-Switching Breaks Voice AI: Building Voicebots for Southeast Asia
Callers in Malaysia, Singapore and much of Southeast Asia switch languages mid-sentence — and most voice agents are built to pick one language per call. Here is why that breaks, and how to design a voicebot that mirrors the caller instead.
Code-Switching Breaks Voice AI: Building Voicebots for Southeast Asia
Short answer: Code-switching is mixing two or more languages inside a single utterance — routine in Malaysia, Singapore, the Philippines and much of Southeast Asia. Most voice agents assign one language per call, so a mixed sentence gets mistranscribed, misclassified, or mispronounced. Building for the region means detecting and mirroring the caller's mix, not asking them to choose.
Key takeaways
A Malaysian caller may open in English, give an address in Malay, and use a Hokkien or Tamil word for a family relationship — all in one sentence.
Single-language pipelines fail in three separate places: language identification, transcription, and speech synthesis.
Forcing "press 1 for English, 2 for Bahasa" at the start does not solve this. It moves the failure later into the call.
The fix is architectural: multilingual-first recognition or native audio models, code-switched test sets from real recordings, pronunciation lexicons for local names and places, and a persona allowed to mirror the caller mid-call.
What is code-switching, and why does it matter for voicebots?
Code-switching is the use of two or more languages within one conversation or one sentence by a fluent speaker. It is not broken language and it is not a niche behaviour — it is the ordinary register of daily speech across much of Southeast Asia, from Malaysian English through Singlish to Taglish.
In practice, a single caller sentence can carry:
an English frame and Malay content words ("Can you check my tempahan for this Saturday?")
Malay grammar with English nouns ("Saya nak reschedule appointment saya")
discourse particles that carry real meaning — lah, kan, meh, ah — which soften, confirm or express doubt
a Hokkien, Cantonese or Tamil kinship term inside an otherwise English sentence
honorifics that must be preserved: Puan, Encik, Tuan, Datuk
Voice agents built on the assumption of one language per call mishandle every one of these.
Where exactly do single-language pipelines break?
Three distinct failure points, which is why teams often fix one and stay broken.
1. Language identification. Most stacks run a language ID step, then lock the recognition model for the rest of the utterance or the rest of the call. A mixed sentence forces a binary decision on input that is genuinely both. Once locked, every subsequent turn inherits the wrong assumption.
2. Transcription. A model tuned for one language transcribes foreign-language words by finding the nearest-sounding native words. The output is fluent, confident and wrong — which is worse than a low-confidence result, because it passes downstream validation and reaches your intent classifier looking legitimate.
3. Speech synthesis. A voice locked to English pronounces Jalan Tun Razak as if it were English. A voice locked to Malay mangles English product names. Both cost credibility instantly. Callers forgive a machine for being a machine; they do not forgive it for mispronouncing their own street.
Downstream, intent classifiers trained on monolingual data compound all three, because the training distribution never contained a sentence shaped like the one that just arrived.
Why doesn't a language menu fix it?
Because language choice is not a property of the caller. It is a property of each sentence.
A caller who selects "English" at the menu will still say their address in Malay, because that is the only form the address exists in. A caller who selects Bahasa Malaysia will still say "reschedule," "appointment," "invoice" and "follow up" in English, because those are the words used for those concepts in daily speech.
Menus also add friction to the exact moment where abandonment is highest — the first fifteen seconds — and they push the failure later into the call, where it costs more.
Monolingual vs multilingual voice agent design
Design decision | Single-language agent | Code-switch aware agent |
|---|---|---|
Language selection | Menu or one bot per language | Detected continuously, mirrored per turn |
Recognition | Language-locked model | Multilingual-first or native audio model |
Mid-call switching | Not supported, or restarts the flow | Supported without losing state |
Local names and places | Mispronounced | Handled by pronunciation lexicon |
Test data | Clean monolingual scripts | Real recordings, mixed-language, per market |
Maintenance | N bots to keep in sync | One flow, localised surface |
How do you build a voice agent that handles code-switching?
Choose recognition that does not force a language lock. Either multilingual-first models or a speech-to-speech architecture that never reduces the audio to a single-language transcript.
Build your test set from real calls, not translations. Take recordings, transcribe them exactly as spoken — mixed, with particles intact — and treat that as ground truth. Clean monolingual test sets are how code-switch failures reach production undetected.
Maintain a pronunciation lexicon per market. Street prefixes (Jalan, Lorong, Taman, Persiaran), city names, honorifics, brand names, and the way local numbers are actually said aloud.
Separate quantities from labels in every language. An amount is a natural quantity: "one thousand ringgit," "seribu ringgit." A reference code is a label, read character by character. This rule holds in every language you support, and it is the detail most drafts get wrong.
Let the persona mirror, not translate. Instruct the agent to answer in the language mix the caller used, and to follow a mid-call switch without restarting the flow or re-asking questions already answered.
Localise the script, do not translate it. A closing question translated word for word from English into Malay often lands as too direct or oddly formal. Rewrite for the register, then test aloud.
Localise the compliance lines too. AI disclosure, recording notice and opt-out must be natural and unambiguous in every language you deploy — a stiff literal translation of "I am an AI assistant" reads as evasive rather than transparent.
Our data: [Insert your own word error rate or intent accuracy comparison on monolingual versus code-switched audio from your own call recordings. Regional performance data barely exists publicly, which makes yours disproportionately citable.]
What does good look like on a live call?
The caller opens in English, the agent answers in English. The caller gives an address in Malay; the agent repeats it back in Malay, pronounced correctly, and continues in English. The caller uses lah to soften a refusal; the agent treats it as a soft no rather than a hard one, and does not repeat the closing question it already used. Asked whether it is a real person, the agent says plainly that it is an AI assistant — in the language the caller just used — and continues warmly.
None of that is exotic. It is simply the behaviour of anyone who actually works a phone line in Kuala Lumpur.
Frequently asked questions
Can current speech recognition handle code-switching? Multilingual models handle it considerably better than language-locked ones, and speech-to-speech models better still, because they never compress the audio into a single-language transcript. Accuracy remains lower than clean monolingual speech, so measure it on your own recordings rather than assuming.
Should I build one voicebot per language or one multilingual agent? One agent for markets where callers genuinely mix languages, because separate bots cannot handle a mid-sentence switch and drift apart in maintenance. Separate agents make sense only where markets are cleanly monolingual and scripts differ substantially.
How do I handle particles like "lah" and "kan"? Treat them as meaning, not noise. They mark softening, confirmation-seeking and doubt, and stripping them changes how a caller's answer should be interpreted. Keep them in transcripts and in your test set ground truth.
Which languages matter most for Southeast Asian voice AI? It depends on your customer base, but Malaysian deployments commonly need Bahasa Malaysia and English at minimum, frequently Mandarin and Tamil, and always the mixed register between them rather than each in isolation.
Does code-switching affect text-to-speech quality? Significantly. A voice locked to one language mispronounces everything from the other, including street names and brand names. Multilingual voices plus a per-market pronunciation lexicon are the practical fix.
Frequently asked questions
Can current speech recognition handle code-switching?
Multilingual models handle it better than language-locked ones, and speech-to-speech models better still because they never compress audio into a single-language transcript. Accuracy is still lower than clean monolingual speech, so measure on your own recordings.
Should I build one voicebot per language or one multilingual agent?
One agent where callers genuinely mix languages, because separate bots cannot handle a mid-sentence switch and drift apart in maintenance. Separate agents suit cleanly monolingual markets with substantially different scripts.
How should a voicebot handle particles like lah and kan?
Treat them as meaning rather than noise. They mark softening, confirmation-seeking and doubt, and removing them changes how a caller's answer should be interpreted. Keep them in transcripts and test set ground truth.
Which languages matter most for Southeast Asian voice AI?
It depends on the customer base. Malaysian deployments commonly need Bahasa Malaysia and English at minimum, often Mandarin and Tamil, and always the mixed register between them rather than each language in isolation.
Does code-switching affect text-to-speech quality?
Yes, significantly. A voice locked to one language mispronounces words from the other, including street and brand names. Multilingual voices plus a per-market pronunciation lexicon are the practical fix.
See it on one of your own call flows
Bring one workflow. We will configure it, validate it with controlled test calls, and show you the whole system around it.