Voice AI Localisation

Code-Switching Breaks Voice AI: Building Voicebots for Southeast Asia

Callers in Malaysia, Singapore and much of Southeast Asia switch languages mid-sentence — and most voice agents are built to pick one language per call. Here is why that breaks, and how to design a voicebot that mirrors the caller instead.

· 6 min read
Illustration of a single speech bubble containing two interleaving colours, representing a caller switching between languages mid-sentence

Code-Switching Breaks Voice AI: Building Voicebots for Southeast Asia

Short answer: Code-switching is mixing two or more languages inside a single utterance — routine in Malaysia, Singapore, the Philippines and much of Southeast Asia. Most voice agents assign one language per call, so a mixed sentence gets mistranscribed, misclassified, or mispronounced. Building for the region means detecting and mirroring the caller's mix, not asking them to choose.

Key takeaways


What is code-switching, and why does it matter for voicebots?

Code-switching is the use of two or more languages within one conversation or one sentence by a fluent speaker. It is not broken language and it is not a niche behaviour — it is the ordinary register of daily speech across much of Southeast Asia, from Malaysian English through Singlish to Taglish.

In practice, a single caller sentence can carry:

Voice agents built on the assumption of one language per call mishandle every one of these.

Where exactly do single-language pipelines break?

Three distinct failure points, which is why teams often fix one and stay broken.

1. Language identification. Most stacks run a language ID step, then lock the recognition model for the rest of the utterance or the rest of the call. A mixed sentence forces a binary decision on input that is genuinely both. Once locked, every subsequent turn inherits the wrong assumption.

2. Transcription. A model tuned for one language transcribes foreign-language words by finding the nearest-sounding native words. The output is fluent, confident and wrong — which is worse than a low-confidence result, because it passes downstream validation and reaches your intent classifier looking legitimate.

3. Speech synthesis. A voice locked to English pronounces Jalan Tun Razak as if it were English. A voice locked to Malay mangles English product names. Both cost credibility instantly. Callers forgive a machine for being a machine; they do not forgive it for mispronouncing their own street.

Downstream, intent classifiers trained on monolingual data compound all three, because the training distribution never contained a sentence shaped like the one that just arrived.

Why doesn't a language menu fix it?

Because language choice is not a property of the caller. It is a property of each sentence.

A caller who selects "English" at the menu will still say their address in Malay, because that is the only form the address exists in. A caller who selects Bahasa Malaysia will still say "reschedule," "appointment," "invoice" and "follow up" in English, because those are the words used for those concepts in daily speech.

Menus also add friction to the exact moment where abandonment is highest — the first fifteen seconds — and they push the failure later into the call, where it costs more.

Monolingual vs multilingual voice agent design

Design decision

Single-language agent

Code-switch aware agent

Language selection

Menu or one bot per language

Detected continuously, mirrored per turn

Recognition

Language-locked model

Multilingual-first or native audio model

Mid-call switching

Not supported, or restarts the flow

Supported without losing state

Local names and places

Mispronounced

Handled by pronunciation lexicon

Test data

Clean monolingual scripts

Real recordings, mixed-language, per market

Maintenance

N bots to keep in sync

One flow, localised surface

How do you build a voice agent that handles code-switching?

  1. Choose recognition that does not force a language lock. Either multilingual-first models or a speech-to-speech architecture that never reduces the audio to a single-language transcript.

  2. Build your test set from real calls, not translations. Take recordings, transcribe them exactly as spoken — mixed, with particles intact — and treat that as ground truth. Clean monolingual test sets are how code-switch failures reach production undetected.

  3. Maintain a pronunciation lexicon per market. Street prefixes (Jalan, Lorong, Taman, Persiaran), city names, honorifics, brand names, and the way local numbers are actually said aloud.

  4. Separate quantities from labels in every language. An amount is a natural quantity: "one thousand ringgit," "seribu ringgit." A reference code is a label, read character by character. This rule holds in every language you support, and it is the detail most drafts get wrong.

  5. Let the persona mirror, not translate. Instruct the agent to answer in the language mix the caller used, and to follow a mid-call switch without restarting the flow or re-asking questions already answered.

  6. Localise the script, do not translate it. A closing question translated word for word from English into Malay often lands as too direct or oddly formal. Rewrite for the register, then test aloud.

  7. Localise the compliance lines too. AI disclosure, recording notice and opt-out must be natural and unambiguous in every language you deploy — a stiff literal translation of "I am an AI assistant" reads as evasive rather than transparent.

Our data: [Insert your own word error rate or intent accuracy comparison on monolingual versus code-switched audio from your own call recordings. Regional performance data barely exists publicly, which makes yours disproportionately citable.]

What does good look like on a live call?

The caller opens in English, the agent answers in English. The caller gives an address in Malay; the agent repeats it back in Malay, pronounced correctly, and continues in English. The caller uses lah to soften a refusal; the agent treats it as a soft no rather than a hard one, and does not repeat the closing question it already used. Asked whether it is a real person, the agent says plainly that it is an AI assistant — in the language the caller just used — and continues warmly.

None of that is exotic. It is simply the behaviour of anyone who actually works a phone line in Kuala Lumpur.

Frequently asked questions

Can current speech recognition handle code-switching? Multilingual models handle it considerably better than language-locked ones, and speech-to-speech models better still, because they never compress the audio into a single-language transcript. Accuracy remains lower than clean monolingual speech, so measure it on your own recordings rather than assuming.

Should I build one voicebot per language or one multilingual agent? One agent for markets where callers genuinely mix languages, because separate bots cannot handle a mid-sentence switch and drift apart in maintenance. Separate agents make sense only where markets are cleanly monolingual and scripts differ substantially.

How do I handle particles like "lah" and "kan"? Treat them as meaning, not noise. They mark softening, confirmation-seeking and doubt, and stripping them changes how a caller's answer should be interpreted. Keep them in transcripts and in your test set ground truth.

Which languages matter most for Southeast Asian voice AI? It depends on your customer base, but Malaysian deployments commonly need Bahasa Malaysia and English at minimum, frequently Mandarin and Tamil, and always the mixed register between them rather than each in isolation.

Does code-switching affect text-to-speech quality? Significantly. A voice locked to one language mispronounces everything from the other, including street names and brand names. Multilingual voices plus a per-market pronunciation lexicon are the practical fix.

Frequently asked questions

Can current speech recognition handle code-switching?

Multilingual models handle it better than language-locked ones, and speech-to-speech models better still because they never compress audio into a single-language transcript. Accuracy is still lower than clean monolingual speech, so measure on your own recordings.

Should I build one voicebot per language or one multilingual agent?

One agent where callers genuinely mix languages, because separate bots cannot handle a mid-sentence switch and drift apart in maintenance. Separate agents suit cleanly monolingual markets with substantially different scripts.

How should a voicebot handle particles like lah and kan?

Treat them as meaning rather than noise. They mark softening, confirmation-seeking and doubt, and removing them changes how a caller's answer should be interpreted. Keep them in transcripts and test set ground truth.

Which languages matter most for Southeast Asian voice AI?

It depends on the customer base. Malaysian deployments commonly need Bahasa Malaysia and English at minimum, often Mandarin and Tamil, and always the mixed register between them rather than each language in isolation.

Does code-switching affect text-to-speech quality?

Yes, significantly. A voice locked to one language mispronounces words from the other, including street and brand names. Multilingual voices plus a per-market pronunciation lexicon are the practical fix.

See it on one of your own call flows

Bring one workflow. We will configure it, validate it with controlled test calls, and show you the whole system around it.