Skip to content

transcription

Why Burmese-English Code-Switching Breaks Transcription Tools

Most transcription tools fail on Burmese meetings where English terms appear mid-sentence. Here is why it happens, how to measure it, and what helps.

Loka Team9 min read
Simple editorial illustration of a speech waveform splitting into Burmese script characters on one side and Latin letters on the other, joined in a single sentence line

Burmese English code switching transcription fails because most speech recognition systems assume one language per sentence, or even one language per recording. In a real Myanmar meeting, English terms appear inside Burmese grammar all the time, and the model has to identify the language, the sounds and the word boundaries at once. Most tools do at least one of those badly.

This post explains the specific mechanics of the failure, what published results show, and how to test and handle bilingual meetings in practice.

What Burmese-English code-switching looks like in a real meeting

Code-switching is the normal way many Myanmar professionals talk about work, not a sign of sloppy speech. The English words are usually the technical or institutional ones: budget, deadline, sprint, KPI, proposal, donor.

There are two kinds:

  • Inter-sentential switching changes language between sentences. One sentence is fully Burmese, the next fully English. This is the easier case, because each sentence can be treated as one language.
  • Intra-sentential switching changes language inside a sentence. For example: "နက်ဖြန် sprint review လုပ်မယ်" (we'll do the sprint review tomorrow). The English phrase sits in the middle of a Burmese clause.

Intra-sentential switching is the hard one, and it is what most office, NGO and tech meetings in Myanmar actually contain. Advice from many transcription guides is to "speak one language" or "avoid switching." That is not realistic for a team where the project documents are in English and the discussion is in Burmese. The tool has to adapt to the people, not the other way around.

Why language identification fails mid-sentence

Many systems pick a language first and then transcribe, so a switch in the middle of a sentence has nowhere to go. One guide to meeting tools notes that many commercial products ask you to pre-select a single dominant meeting language, and treat automatic detection as a file-level classification rather than something that tracks switches within a sentence.

The cost shows up at the seams. A write-up on multilingual meeting transcription reports that traditional monolingual models can see word error rate at language-switching boundaries spike by 30 to 50 percentage points compared with monolingual speech. That figure is not specific to Burmese, but it describes the mechanism: the model is confident in one language, the audio changes, and the errors cluster exactly where the switch happens.

For a meeting, that is the worst place to be wrong. The switched words are often the names, systems, numbers and terms that people search for later.

Why Burmese makes the problem harder than Spanish-English

Burmese differs from English in script, word boundaries, tone and sentence order, so a model cannot treat a Burmese-English switch like a Spanish-English one. Most code-switching advice is written for European language pairs that share a Latin alphabet and spaces between words. Burmese shares neither.

Script and tone

Burmese is spoken natively by over 33 million people and uses an abugida script descended from Mon, where pitch and tone shifts can change a word's meaning entirely. Its subject-object-verb word order also diverges from English. So an inserted English noun or verb phrase lands inside a grammatical frame that English models have never seen it in.

No spaces between words

Burmese is written without spaces separating words. The practical consequence for transcription is that the system has to decide where words begin and end before it can even output text. When an English term is inserted, the boundary between Burmese particles and the English word is another decision the model can get wrong. This also affects how you score output, which we cover below.

Pronunciation borrowing

Speakers often pronounce English words with Burmese sound patterns, and the pronunciation varies by speaker. "Budget" in a Yangon meeting may not sound like "budget" in a model trained mostly on English from native speakers. The acoustic model sees something that is neither clean Burmese nor clean English, and may map it to the wrong language or to a similar-sounding Burmese word.

Why Whisper and other large global models struggle

Large multilingual models can over-commit to whichever language they judge dominant, which researchers call matrix language bias. In the matrix language, the grammar frame of the sentence is set. The other language gets embedded in it. A model that decides the frame is Burmese may force an English term into Burmese-looking output, and a model that decides the frame is English may skip or garble the Burmese.

A paper on matrix language effects in colloquial code-switched speech found that large global models such as Whisper Large-v3 can fail to outperform smaller, locally adapted architectures, because of this bias toward a dominant language. Size and broad language coverage do not fix it. A model that has seen a lot of Burmese and a lot of English separately has not necessarily seen them mixed the way your team mixes them.

That is why a tool can look fine on a clean Burmese podcast and then produce a messy transcript from a product meeting. The test audio and the real audio are different tasks.

If you want a broader walkthrough of getting usable Burmese output, including audio quality and file preparation, see our guide on how to transcribe Burmese audio accurately.

How to measure Burmese English code switching transcription

Standard word error rate (WER) is a poor fit for Burmese on its own, so you need a second measure. Because Burmese has no spaces between words, whitespace-tokenized WER is unsuitable, and character error rate (CER) or specialized segmentation is needed for fair evaluation. If two tools segment the same sentence differently, their WER can differ even when the text is equally readable.

What published results look like

Research on Myanmar-English code-switched speech gives a sense of the range. The MEASR dataset is about 10 hours of real Myanmar-English intra-sentential code-switching audio, recorded at 16 kHz, mono, 16-bit, from online IT lectures given by 10 lecturers. Reported results on that kind of data:

SystemReported WERSource
GMM-HMM baselines31.31% to 31.95%Real-Time Transcription and Translation presentation
Deep neural network (DNN)26.69%Same presentation
DNN (p-norm) with dual pronunciation lexicons22.23%Myanmar-English Code-Switched ASR System

Read these as indicators, not promises. The setups are not necessarily identical, the audio is lecture speech from 10 speakers, and your meetings will have overlapping talk, accents, bad microphones and different vocabulary. Even the best figure in the table is about one error in every four or five words, which shows how hard the task is.

How to test on your own audio

  1. Pick 5 to 10 minutes of a real meeting with natural switching, not a scripted sample.
  2. Write a reference transcript by hand for that clip, or correct one carefully.
  3. Score Burmese with CER, or with a consistent segmentation applied to both texts.
  4. Separately list every English term spoken (names, acronyms, numbers, tools) and mark whether each one appears correctly.
  5. Run each tool on the same clip, same audio file, same language setting.
  6. Compare the English-term hit rate as well as the overall error rate. A tool with a decent overall score that drops half the English terms may still be unusable for notes.

What helps: how specialized systems handle switching

Systems that do better on Burmese-English tend to be built around the mixed speech itself rather than adapted from a monolingual model. The most concrete example in the research is pronunciation modeling. The system that reached 22.23% WER combined two pronunciation lexicons, MyanmarDict for Burmese and the CMU dictionary for English, with a DNN approach on Myanmar-English IT lecture speech.

The idea is simple. If the system has pronunciation entries for both languages, an English word spoken in a Burmese meeting has a candidate to match, instead of being forced into the nearest Burmese sound sequence. The training data matters too: the improvement from the baseline GMM-HMM models to the DNN, on the figures above, came with data that actually contained code-switching.

For buyers, this translates to a few questions to ask any vendor:

  • Is mixed Burmese-English speech part of how the model was built, or is Burmese one language pack among many?
  • Does language detection work within a recording, or only per file?
  • Can you see, and correct, the transcript afterwards?

Teams comparing options for mixed-language calls may also want our comparison of Otter alternatives for teams that switch languages mid-meeting.

Practical steps for bilingual Myanmar teams

You cannot remove switching from your meetings, but you can make transcripts more reliable by controlling the things around the model.

  1. Record with a decent microphone, close to the speaker. Pronunciation-borrowed English words are already ambiguous. Noise makes them worse.
  2. Set the recording language to the main language of the meeting. Then check the output, since the language setting is only a starting point.
  3. Say key English terms clearly once. Names of projects, donors, systems and numbers are worth stating cleanly at least once per meeting, even if people shorten them afterward.
  4. Review the first transcript on a new kind of meeting. A new team, topic or room can change accuracy. A five-minute spot check against the audio tells you how much to trust the rest.
  5. Keep a running glossary. List the English terms your team uses and check that they appear correctly in transcripts. This also makes your tests repeatable.
  6. Match the output to the record you need. If the meeting produces formal minutes, our guide to Myanmar meeting minutes and a bilingual template covers what to capture beyond the raw transcript.

Where Loka Note fits

Loka Note is built Burmese-first in Yangon, and it handles mixed Burmese-English code-switching natively rather than treating it as an edge case. It uses automatic language detection and supports mixed-language speech within a single meeting. The language you pick for a recording is checked against the audio, so a meeting recorded under the wrong language setting is transcribed in the language actually spoken.

You can record in the browser with no bot joining the call, or upload audio and video files, and get a transcript, a summary with decisions and action items, and "Ask Loka" to query past meetings in your own language. Pricing is pay-as-you-go minute top-ups starting at $1.99 (60 minutes), or Unlimited at $19.99 a month, and every account has the full feature set. Customer audio and transcripts are never used to train AI models; details are on our security page.

We would still suggest the test above: run a clip of your own meeting through any tool you are considering, including ours.

Try it on a real bilingual meeting: create a free account at Loka Note

Frequently asked questions

Why does Whisper struggle with mixed Burmese and English audio?

Large global models can over-commit to one dominant language, a pattern researchers call matrix language bias. When the model decides a sentence is Burmese, English terms inside it tend to get forced into Burmese-sounding output, and the reverse can also happen. Research on matrix language effects found that a large model like Whisper Large-v3 can fail to beat smaller, locally adapted models on this kind of speech.

Can AI transcription tools automatically detect language switches within a single sentence?

Some can, but many cannot. A number of commercial meeting tools treat automatic language detection as a file-level choice, so they pick one dominant language for the whole recording. Mid-sentence switching needs a system built for it, and you should test any tool on your own audio before relying on it.

What is the difference between inter-sentential and intra-sentential code-switching in Burmese?

Inter-sentential switching changes language between sentences: one full sentence in Burmese, the next in English. Intra-sentential switching changes language inside a sentence, such as a Burmese sentence with an English phrase like 'sprint review' in the middle. The second type is far harder for speech recognition and is what most Myanmar workplace meetings contain.

How should teams evaluate and benchmark code-switched speech recognition tools?

Use a short sample of your own meetings, not a vendor demo. Score Burmese with character error rate or a consistent word segmentation, because Burmese has no spaces between words. Also check English terms separately, since names, numbers and technical terms are usually what readers need most.

Sources

  1. 1Real-Time Transcription and Translation fornaivo.org
  2. 2Myanmar-English Code-Switched Automatic Speech Recognition Systemexa.ai
  3. 3Myanmar-English Code-Switching Speech Dataset: Measrresearchgate.net
  4. 4Matrix Language Effects on Colloquial Code-Switched Speech: A Diagnostic Analysis of Structural Bottlenecks in Neural ASR and TTS | Proceedings of the 2026 11th International Conference on Intelligent Information Technologydl.acm.org
  5. 5Turn Burmese audiospeakai.co
  6. 6Multilingual Meeting Transcription: What Works and What Doesn't | MinuteKeep Blog | 現場コンパスgenbacompass.com
  7. 7Mixed-Language Meeting Transcription and Code-Switchingtranscription.atter-ai.com
  8. 8ncwn/speech-to-textgithub.com

Keep reading