
To transcribe Burmese audio accurately, start with clean mono audio, use an engine that was trained or fine-tuned on Burmese, keep the output in standard Unicode, and have a Burmese speaker review the result. Generic speech models still make many errors on Burmese, so the workflow around the model matters as much as the model.
This guide covers why Burmese is hard for automatic speech recognition (ASR), how to prepare files, how to choose an engine, how to measure accuracy without fooling yourself, and how to avoid the encoding problems that corrupt transcripts and subtitles.
Why generic models struggle with Burmese
Burmese is a low-resource language for speech AI: there is far less labeled audio to train on than for English, Spanish or Mandarin. That is the main reason general-purpose models fall short.
One public Burmese transcription page reports that off-the-shelf Whisper models such as large-v3-turbo show a word error rate above 25% for Burmese. It places the language in a tier where output is useful for gist or search but not ready to publish. If you have ever pasted a Whisper transcript into a report and found it unusable, this is why.
A few things compound the data problem:
- Tones and pitch. Burmese is tonal, so pitch distinctions change meaning. Transcription guides note that acoustic model quality and recording clarity directly affect how intelligible the transcript is. A muffled phone recording loses exactly the cues the model needs.
- Domain vocabulary. Medical, legal, and technical speech use terms that rarely appear in general training data. The myMediCon study on Burmese medical conversations shows that Burmese speech-to-text needs specialized domain corpora and segmentation routines to be evaluated and trained properly.
- Code-switching. Real Myanmar meetings mix Burmese and English freely. A model that expects one language per file will stumble when someone says "deadline" or "budget" mid-sentence.
The script problem: word boundaries, stacked consonants, and Zawgyi
Burmese text has no spaces marking word boundaries, its characters combine into stacked clusters, and two incompatible encodings (Unicode and Zawgyi) are still in circulation. Each of these can break a transcript even when the model heard the audio correctly.
No word boundaries
Burmese writing does not use spaces to delimit words. Tools built for English split text on whitespace to count words, align subtitles and chunk text for search. On Burmese, that logic either fails or produces arbitrary results. This is why syllable-level or character-level handling is the safer approach, and we come back to it in the accuracy section.
Stacked consonants and normalization
Burmese is an abugida: consonants carry an inherent vowel, and marks and stacked consonants (such as မြ) combine into a single visual unit. The same visible text can be stored as different sequences of code points. A source on Burmese transcription notes that strict Unicode NFC normalization is needed to prevent encoding mismatches, including against legacy encodings like Zawgyi.
In practice, two transcripts that look identical on screen can fail an exact-match search or a diff, because the underlying characters are ordered differently. Normalize text to NFC before you store it, search it, or compare it with a reference.
Zawgyi versus standard Unicode
Zawgyi is an older, non-standard font encoding that reuses Unicode code points with different meanings. Text written in Zawgyi looks right only if the reader has a Zawgyi font. Open it in a Unicode environment and you get broken, scrambled characters.
Normalization does not fix this. Converting Zawgyi to Unicode needs a dedicated converter. Treat it as two separate steps:
- Make sure the transcript your engine produces is standard Unicode.
- If you also work with older Zawgyi documents (glossaries, reference transcripts), convert them to Unicode first, then normalize to NFC.
Test the final file on a phone and a laptop that you do not control. If characters look wrong anywhere, you have an encoding problem, not a recognition problem.
How to transcribe Burmese audio step by step
The steps below take a raw recording to a reviewed transcript. Most of the accuracy you can control comes from the first three.
- Record clearly at the source. Put the microphone close to the speaker. Where possible, record one speaker per channel or one microphone per person. Room echo and overlapping voices hurt every ASR model, and tonal languages suffer more because pitch detail gets smeared.
- Keep the original file. Never overwrite it. Make a working copy for processing so you can go back if a filter damages the speech.
- Convert the working copy to mono, 16 kHz. Most speech models work with 16 kHz mono audio internally. Converting yourself gives you control over the result. Use a lossless format such as WAV for the working copy if file size allows.
- Reduce noise lightly. Remove constant hum or hiss and trim long silences. Do not apply aggressive noise suppression, which can strip the consonant detail and pitch contours the model needs.
- Set the language correctly. If your tool lets you choose, choose Burmese. If the recording mixes Burmese and English, use an engine that handles mixed speech rather than forcing English.
- Run the transcription. For long recordings, split at natural pauses, not at fixed time marks that cut through a word.
- Normalize and check encoding. Output should be standard Unicode, normalized to NFC. Open it on a second device.
- Review with a Burmese speaker. Focus on names, numbers, dates, and domain terms, where errors do the most damage.
Choosing an engine: off-the-shelf Whisper or a Burmese-tuned model
Fine-tuning on Burmese data can cut errors dramatically compared with a zero-shot multilingual model, but the output of any engine still needs review for anything you will publish or act on.
The clearest public numbers come from an open-source project. In one Burmese ASR repository, baseline zero-shot multilingual models produced error rates approaching 100% WER and about 88% CER, and fine-tuning on domain-specific data brought those down to roughly 33% WER and 13% CER. That is one project on its own dataset, so do not treat it as a universal benchmark. It does show the size of the gap that Burmese-specific training can close.
| Option | What to expect | Best for | Main risk |
|---|---|---|---|
| Off-the-shelf Whisper (e.g. large-v3-turbo) | Above 25% WER; gist-level | Searching recordings, rough triage | Errors in names, numbers, and terms go unnoticed |
| Zero-shot multilingual model, no tuning | Can approach ~100% WER on some data | Not recommended for Burmese | Output may be close to unusable |
| Model fine-tuned on Burmese, ideally in your domain | Far lower error; ~33% WER and ~13% CER in one open project | Meetings, interviews, research | Domain mismatch (e.g. medical terms) still causes errors |
| Any ASR plus human review | Slower, but publication-ready | Quotes, legal, subtitles, reports | Reviewer time and cost |
Whichever you pick, run a short sample of your own audio first. A model that does well on clean read speech can do badly on a noisy phone call with regional accents.
Measuring accuracy: use CER, not only WER
Character Error Rate (CER) is the most dependable accuracy measure for Burmese, because Word Error Rate depends on spaces that Burmese does not use to mark words. Since Burmese has no whitespace word boundaries, a whitespace-based WER can swing widely depending on how someone chose to space the text. The same source points to CER or syllable-based segmentation as the primary metrics.
To check a tool on your own audio:
- Pick 5 to 10 minutes of representative audio, including your noisiest and most technical parts.
- Have a native speaker produce a careful reference transcript in standard Unicode.
- Normalize both the reference and the machine output to NFC.
- Strip or standardize spaces and punctuation the same way in both.
- Compute CER (character-level edit distance divided by reference length). If you want a word-style score, segment both texts into syllables first and compute the error rate on syllables.
- Read the errors, not just the number. Ten wrong filler words matter less than one wrong amount of money.
Compare tools on the same sample, with the same normalization. Numbers from different datasets, or computed with different segmentation, are not comparable.
Review workflow, subtitles, and summaries
The most reliable workflow is hybrid: let ASR produce a draft quickly, then have a Burmese speaker correct it. That is much faster than typing from scratch and much safer than trusting raw output.
A practical review pass
- Mark low-confidence sections if your tool shows them, and review those first.
- Check proper names, place names, numbers, and dates against another source.
- Keep a shared glossary of terms and spellings, in standard Unicode, so reviewers stay consistent.
- Note dialect or accent problems. Speech from outside the training data's usual accents tends to produce more errors, so flag them for closer review.
Subtitles in SRT or VTT
ASR tools can produce timestamped text, and that is what subtitle formats need. The work is in making the text readable on screen:
- Save the file as UTF-8 so Burmese characters survive.
- Keep lines short, and break at phrase or syllable boundaries. Never split a stacked cluster or a character from its combining marks across lines.
- Do not rely on spaces to decide where to wrap, since Burmese does not use them between words.
- Play the file in the actual player or platform. Some fonts and players render Burmese combining marks badly.
- Correct the text before you time it. Fixing errors afterward means touching every cue.
Meeting summaries
For meetings, a clean transcript is often only the input. What people use is the decisions and action items. Summaries inherit transcript errors, so check them against the transcript for names and numbers before sending them out.
Where Loka Note fits
Loka Note is an AI meeting-notes tool built Burmese-first in Yangon. You can record a meeting in the browser (no bot joins the call) or upload audio and video files. It produces a transcript, a summary with decisions and action items, and a to-do list from the action items.
A few details relevant to the problems above:
- It handles Burmese natively, including mixed Burmese-English speech in the same meeting.
- The language you pick is checked against the audio, so a recording set to the wrong language is transcribed in the one actually spoken.
- It also supports English, Thai, Vietnamese, Chinese, Yoruba, and Hausa.
- Customer audio and transcripts are used only for transcription and summarization and are never used to train AI models. Details are on our security page.
- Pricing is pay-as-you-go minute top-ups, starting at $1.99 for 60 minutes, or $19.99 per month for Unlimited. Every account has the full feature set.
We have not put a published error rate for Loka Note in this article. Test any tool, ours included, on a sample of your own audio using the CER method above. You can see the Burmese-language page at lokanote.com/mm.
Frequently asked questions
Why does Whisper make so many mistakes on Burmese audio?
Burmese is a low-resource language, so general multilingual models have seen far less training data for it than for English. One public tool page reports Whisper large-v3-turbo above 25% word error rate on Burmese, which it describes as fine for gist or search but not for publication. The script, which has no spaces between words, and the tonal sounds add further difficulty.
How do I get Burmese transcripts in standard Unicode instead of Zawgyi?
Use an engine that outputs standard Unicode and normalize the text to NFC before storing or exporting it. If you receive legacy Zawgyi text, run it through a Zawgyi-to-Unicode converter first, because normalization alone does not convert between the two encodings. Then check the result in the app where it will be read.
Can AI automatically add SRT or VTT subtitles to Burmese video?
Speech recognition tools can produce timestamped text, and that is what subtitle formats need. The catch is accuracy: errors in the text appear on screen, and Burmese line breaks need care because there are no word spaces. Always review the text and test the file in a player before publishing.
Should I measure Burmese accuracy with WER or CER?
Use Character Error Rate as your main measure, or segment the text into syllables before computing WER. Standard WER splits text on spaces, and Burmese does not use spaces to mark word boundaries, so the score becomes unreliable. CER avoids that problem.
How should I prepare Burmese audio before transcribing it?
Keep the original recording, then make a working copy as clean mono audio at 16 kHz. Remove long silences and obvious noise gently, and avoid heavy filtering that distorts speech. Clear speech matters more than anything else, because pitch differences change meaning in Burmese.
Sources
- 1Transcribe Burmese Audio Free | Free.aifree.ai
- 2Turn Burmese audiospeakai.co
- 3Burmese Speech to Text | Speechyou - Speechyouspeechyou.com
- 4ncwn/speech-to-textgithub.com
- 5Hein-HtetSan/myanmar-asrgithub.com
- 6myMediCon: End-to-End Burmese Automatic Speech Recognition for Medical Conversationsaclanthology.org


