Skip to content

transcription

Burmese Speech to Text in 2026: What Accuracy Really Looks Like

Off-the-shelf Whisper returned empty transcripts on 28 of 30 clean Burmese clips in one test. Here is what fine-tuned models achieve, why CER matters, and how to test before you commit.

Loka Team9 min read
Simple editorial illustration of a sound waveform turning into Myanmar script characters, with a few characters shown as empty boxes

Burmese speech to text works well only when the model has been trained or fine-tuned on Burmese audio. Out-of-the-box Whisper Large v3 returned completely empty transcripts on 28 of 30 clean, single-speaker Burmese clips in one published test. Fine-tuned and specialized models do much better, and the way you measure accuracy decides whether you can trust any number you are shown.

This post covers what the published results say, why Burmese is hard for speech recognition, why character error rate (CER) is the right metric, and how to test a system on your own audio before you build on it.

Burmese speech to text with Whisper: what testing shows

In one published test, Whisper Large v3 returned empty output for about 93% of clean Burmese clips, an effective character error rate of 100%. That is the zero-shot result, with no fine-tuning and no Burmese-specific adaptation.

The details are worth reading closely. The test on inferenceapis.com used clean, single-speaker audio, the easiest case, and 28 of 30 clips came back with no text at all. Yet the same model identified the Burmese language code ('my') correctly on every unprompted clip.

That combination is the trap. If you check only that the system "detects Burmese," your integration test passes. The transcript is still empty.

Empty output is not the only failure mode. A study on Burmese medical speech (myMediWhisper) reports that vanilla zero-shot Whisper models score word error rates well above 100%, and in places over 200%, before fine-tuning. A WER above 100% means the system produces more wrong words than the reference contains, which happens when a model hallucinates extra text.

What this means in practice:

  • A vendor listing Burmese as a supported language tells you very little about accuracy. AssemblyAI, for example, has a dedicated Burmese page. Ask any vendor for CER on audio that resembles yours.
  • Never accept "language detected" as evidence of "speech transcribed."
  • Always run a pilot with your own recordings.

Why Burmese is hard for speech recognition

Burmese combines tonal syllables, no word spacing, split text encodings and little open training data. Each one hurts accuracy on its own, and together they explain why general-purpose models struggle.

Tone and syllable timing

Burmese is tonal and syllable-timed, so an acoustic model must pick up pitch variation that changes word meaning. WIZ.AI's industry write-up notes this and pairs it with the shortage of open-source training data. Less data and more phonetic subtlety is a bad mix for a model trained mostly on other languages.

No spaces between words

Burmese script does not mark word boundaries with spaces. Speech recognition systems usually need a tokenizer or segmentation step to decide where words end. That choice affects both training and scoring, which is why the next section matters.

The Zawgyi and Unicode problem

Myanmar text has historically been written in two incompatible encodings: legacy Zawgyi and standard Unicode. A speech-to-text evaluation repository points out that this fragmentation creates mismatches when you compare a transcript to a reference, and that text must be strictly normalized before scoring.

The same hazard shows up outside the lab. If your CRM, database or search index expects one encoding and receives the other, the text can display as garbage or fail to match searches. Decide on Unicode as your single internal format and check it at every hand-off.

Real speech is messier than test clips

Real calls and meetings include interruptions, gaps and background noise. Researchers have studied this directly, for example in work on how missing speech segments and interruptions affect Burmese ASR. A system that scores well on clean read speech can still struggle on a noisy conference line.

Why character error rate beats word error rate for Burmese

Use CER for Burmese because WER depends on word boundaries, and Burmese does not have visible ones. The ncwn/speech-to-text project states that whitespace-delimited WER is unsuitable for Burmese and that CER is the standard reliable metric.

Here is why this matters when you compare numbers:

  • WER counts wrong words. If two systems segment the same Burmese sentence differently, they can get different WER on identical output.
  • CER counts wrong characters, so segmentation choices matter much less.
  • A transcript with one wrong syllable looks nearly perfect under CER and can look badly broken under a poorly segmented WER.

You will still see WER in Burmese papers. The myMediWhisper study reports it, and so does the myanmar-asr benchmark alongside CER. Treat WER figures as comparable only within one paper, where the segmentation was held constant. When a vendor quotes a single "accuracy" percentage, ask three things: is it CER or WER, how was the text segmented, and was it normalized to Unicode?

What fine-tuning changes: published results

Fine-tuning on Burmese audio moves error rates from unusable to workable. The published numbers come from different datasets and metrics, so read the table as a set of separate experiments, not a leaderboard.

SetupDataMetricResultSource
Whisper Large v3, zero-shot30 clean single-speaker clipsEmpty transcripts28 of 30 empty (effective CER 100%)inferenceapis.com
Whisper-Medium, zero-shotBurmese medical speechWER256.07%myMediWhisper
Whisper-Medium, full fine-tuning28-hour medical corpusWER23.44%myMediWhisper
whisper-large-v3-myanmarSame medical evaluationWER32.03%myMediWhisper
Facebook MMS-1BSame medical evaluationWER37.90%myMediWhisper
SeamlessM4T v2 Large, fine-tuned54 hours of Myanmar speechCER13.04% (best CER)myanmar-asr
DataoceanAI Dolphin (Whisper-v2), frozen encoder54 hours of Myanmar speechWER33.02% (best WER)myanmar-asr

Three takeaways:

  1. Domain data does the heavy lifting. Full fine-tuning of Whisper-Medium on 28 hours of Burmese medical speech cut WER from 256.07% to 23.44%, and it outperformed both MMS-1B and a community Whisper-Large model on that test.
  2. No single model wins on every metric. In the 54-hour benchmark, SeamlessM4T v2 Large had the best CER and Dolphin had the best WER. Which one is "best" depends on what you measure.
  3. Small, community-tuned models exist and are improving. Examples include whisper-small-burmese-v3 and v4 on Hugging Face. They are useful for experiments, but verify them on your own audio before putting them in production.

Notice also that the medical results come from a narrow domain. A model tuned on clinical dialogue has not been shown to handle sales calls, board meetings or call-center audio. Earlier work on Burmese medical conversations (myMediCon) points the same way: matched-domain data is what builds usable Burmese ASR.

Frozen encoder or full fine-tuning: how to choose

Start with the cheapest option that meets your CER target on your own audio, and move to heavier fine-tuning only if it misses. The published evidence supports both approaches, so the decision is practical.

Frozen encoder

You keep the pretrained audio encoder and train the rest. Dolphin's best WER in the 54-hour benchmark came from this kind of setup. It generally needs less compute and less data, and it is simpler to maintain. The risk is that the encoder never adapts to your acoustic conditions.

Full fine-tuning (FFT)

You update the whole model. The myMediWhisper result shows how large the gain can be: 256.07% down to 23.44% WER. It costs more compute, needs clean aligned transcripts, and you own the retraining when your domain shifts.

Hosted API or your own deployment

QuestionHosted APISelf-hosted model
Time to first resultFastSlower
Control over domain adaptationLimitedFull
Data stays on your infrastructureDepends on vendor termsYes
Engineering effortLowHigh

Some teams cannot send certain audio off-site, such as clinical, legal or HR recordings. For them, a self-hosted fine-tuned model may be the only option. For everyone else, a specialized hosted service is usually the faster route, provided you verify it with the steps below.

How to test Burmese speech to text before you commit

Run a small, honest evaluation on your own recordings. It takes a day and prevents months of rework.

  1. Collect 30 to 60 minutes of real audio. Include your normal mess: phone lines, overlapping speakers, Burmese-English code-switching, and the names and terms your business uses.
  2. Write reference transcripts by hand. Use a native speaker, and store everything in Unicode.
  3. Normalize both sides. Convert any Zawgyi to Unicode and apply the same normalization to the reference and the system output.
  4. Score with CER. Report WER only if you control segmentation and use it consistently.
  5. Check for empty and runaway outputs. Count clips that return nothing and clips where the output is much longer than the reference. These failures are invisible in averaged scores.
  6. Test the downstream path. Send the transcript through your database, search and display layers and confirm Burmese text survives intact.
  7. Repeat on a second audio condition. If you tested on office recordings, also test on phone audio.

Set a threshold before you look at results. A CER that is fine for searchable meeting archives may be too high for legal records.

Where Loka Note fits

Loka Note is a meeting-notes tool built Burmese-first, so you can try Burmese transcription without assembling an ASR pipeline yourself. We are not publishing a benchmark in this post, and you should run the test above on your own audio.

What it does:

  • Records meetings in the browser with no bot joining the call, or accepts uploaded audio and video files.
  • Transcribes Burmese natively, including the mixed Burmese-English speech common in Myanmar business meetings. It also supports English, Thai, Vietnamese, Chinese, Yoruba and Hausa.
  • Checks the language you pick against the audio, so a recording set to the wrong language is transcribed in the language actually spoken.
  • Produces a summary with decisions, action items and open questions in the meeting's language, and lets you ask questions across past meetings with Ask Loka.

Pricing has no feature tiers. Minute top-ups start at $1.99 for 60 minutes, and an Unlimited plan is $19.99 per month. A small top-up is enough to run your own recordings through it. For data handling, see our security page: customer audio and transcripts are never used to train AI models. A Burmese-language overview is on our Burmese landing page.

If you are building your own ASR system, Loka Note is not a substitute for that work. If you mainly need accurate Burmese meeting notes, it spares you from building one.

Test Burmese transcription on your own meeting audio

Frequently asked questions

Does OpenAI Whisper support Burmese speech to text?

Whisper recognizes Burmese as a language code, and in one published test it detected 'my' correctly every time. But the same test found Whisper Large v3 returned completely empty transcripts on 28 of 30 clean, single-speaker clips. Language detection working is not the same as transcription working.

Why is character error rate used instead of word error rate for Burmese?

Burmese script does not put spaces between words, so word error rate depends on how a tokenizer decides to split the text. Character error rate avoids that problem and is the more reliable metric for Burmese. Text should also be normalized to a single encoding before scoring.

Why does Burmese speech recognition return empty or hallucinated transcripts?

Burmese is tonal, so the acoustic model has to separate pitch-dependent meaning, and open-source training data is limited. Zero-shot Whisper models have been measured at word error rates above 100% on Burmese medical speech, which means the output contains more errors than the reference has words. Fine-tuning on Burmese audio is what brought those numbers down in published work.

Which models perform best for Burmese speech recognition?

It depends on the data and the metric. In one 54-hour benchmark, SeamlessM4T v2 Large had the best character error rate (13.04%) and Dolphin had the best word error rate (33.02%). In a medical corpus study, a fully fine-tuned Whisper-Medium reached 23.44% WER, ahead of MMS-1B at 37.90%.

Does the Zawgyi versus Unicode difference affect transcription?

Yes. Myanmar text exists in legacy Zawgyi encoding and in standard Unicode, and mismatches between them distort scoring and can garble text in downstream systems. Normalize to Unicode before you evaluate or store transcripts.

Sources

  1. 1Burmese speech-to-text with Whisper: measured accuracyinferenceapis.com
  2. 2ncwn/speech-to-textgithub.com
  3. 3myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASRarxiv.org
  4. 4Hein-HtetSan/myanmar-asrgithub.com
  5. 5Burmese Voice AI Readiness for Customer Engagement | WIZ.AI Industry Insightswiz.ai
  6. 6Assessing ASR Robustness for Burmese: Impacts of Missing Speech Segments and Interruptionsaclanthology.org
  7. 7myatsu/whisper-small-burmese-v4 · Hugging Facehuggingface.co
  8. 8Burmese Speech-to-Text API | AssemblyAIassemblyai.com
  9. 9myMediCon: End-to-End Burmese Automatic Speech Recognition for Medical Conversationsaclanthology.org
  10. 10myatsu/whisper-small-burmese-v3 · Hugging Facehuggingface.co

Keep reading