Skip to content

transcription

Multilingual Meeting Transcription: What Actually Works

AI can transcribe meetings in two languages, but 'multilingual' on a feature list can mean four different things. Here is what each one delivers, where it breaks, and how to work around it.

Loka Team9 min read
Simple editorial illustration of two overlapping speech bubbles in different scripts feeding into a single page of meeting notes

Yes, AI can transcribe a meeting where people speak two languages, but how well depends on what the tool means by "multilingual." Some tools need one language picked in advance, some detect the language on their own, and only some handle two languages inside a single sentence. This guide explains the differences, so you can pick the right setup for multilingual meeting transcription on a mixed-language team.

Four things "multilingual meeting transcription" can mean

Vendors use one phrase for four different capabilities, and a long language list tells you nothing about which one you are getting. Hinoter's guide separates them like this:

CapabilityWhat it meansWhat breaks
One selected language per meetingYou choose a language before recording; the whole file is transcribed in itAny other language becomes garbled text
Automatic language detectionThe tool identifies the language itselfDetection can be wrong, and may be made once, not per sentence
Multiple languages in one recordingThe transcript follows speakers as they change languageQuality varies, especially mid-sentence
Translated outputText is rendered in a language different from the one spokenTranslation errors stack on top of transcription errors

The last row is the one people confuse most. Transcription writes down what was said, in the language it was said. Translation is a second step that produces a different language. A tool can be strong at the first and only adequate at the second, so evaluate them separately.

MeetStream's overview makes a related point: multilingual transcription is not one problem but three. They are language detection, accent robustness, and code-switching in the middle of an utterance. A tool that scores well on one can fail on another, which is why a demo in clean English-and-Spanish audio says little about your Tuesday call.

Switching between sentences vs. switching inside them

Switching at sentence boundaries is much easier for tools than switching mid-sentence. If your team's habit is the second kind, which is common in Burmese-English, Thai-English or Yoruba-English business talk, you need to test for it specifically.

Sentence-level switching

One person finishes a thought in English, the next answers in Burmese. If your tool only lets you select one language per recording, ConvertAudioToText's guidance is to set the dominant language, meaning whichever fills more than 50% of the meeting. That resolves most sentence-boundary switching.

Mid-sentence code-switching

This is where models break. The same source notes that secondary-language words dropped into a sentence often degrade into phonetic approximations. The model hears an English term in a Burmese sentence and writes something that sounds right in the main language but means nothing. Product names, client names and technical terms suffer most.

We covered this failure in detail for one language pair in why Burmese-English code-switching breaks transcription tools. The pattern applies to any pair where speakers move freely between a local language and English.

The 50-50 workaround

If a meeting is split roughly evenly and your tool has no native code-switching support, ConvertAudioToText describes a two-pass approach:

  1. Run the recording once with Language A selected.
  2. Run it again with Language B selected.
  3. Compare the two transcripts and keep the better section wherever the speakers were using that language.

It is manual and slow, but the same source reports it gives higher accuracy than any single-pass setting for even splits. Use it for important recordings such as contract negotiations or board discussions, not for every stand-up.

How speaker labels behave when the language changes

Speaker labels follow the voice, not the language, so one person switching languages keeps one label. ConvertAudioToText explains that diarization is acoustic rather than linguistic. Models track pitch, formants and prosody, not vocabulary.

That is mostly good news. Your Burmese-speaking and English-speaking colleagues will not be split into extra speakers because of the language change. Two practical consequences follow:

  • If the words are wrong, the labels can still be right. You can often reconstruct who said what from context, even when a passage was transcribed poorly.
  • Labels can still fail for ordinary reasons, such as overlapping speech, similar voices, or a poor microphone. A speaker who sounds different when switching languages, for example by changing pace or register, can occasionally confuse a model, so spot-check long recordings.

Why accuracy drops for low-resource languages and accents

Non-English audio loses accuracy faster in noisy conditions because the models have seen less acoustic variety in non-English training data. That is the explanation given by MinuteKeep's write-up: noise tolerance degrades faster for non-English languages than for English.

For teams in Myanmar, Thailand, Vietnam or Nigeria, this matches what you probably see already. A call that an English-only tool handles acceptably over a laptop microphone turns messy when the language is Burmese, Hausa or Yoruba. Most published comparisons test major languages such as English, German, French, Spanish and Japanese, so they say little about the languages many of our readers work in.

Benchmarks are still useful for understanding the field. AssemblyAI reports an average cpWER of 30.17 for its Universal-3.5 Pro on multi-speaker multilingual audio, against 37.92 for Deepgram Nova-3 and 35.26 for ElevenLabs Scribe v2. Lower is better, and cpWER counts errors in who said what as well as in the words. Two cautions apply. This is a vendor's own benchmark, so treat it as one data point. And even the best figure there means a substantial share of words are still wrong on hard multi-speaker audio, before you add a low-resource language.

Names and brand terms

Proper nouns are the most visible failure. A tool that renders your company name three different ways makes the whole transcript look unreliable. Where your tool supports a custom vocabulary or glossary, use it for client names, product names and place names. If it does not, keep a short list of the terms that keep going wrong and fix them in one pass before sharing the notes. Find-and-replace on a handful of terms takes minutes and fixes most of the visible damage.

How Microsoft Teams handles multilingual transcription

Teams supports multilingual speech recognition for a fixed set of languages, and its translated captions and transcripts do not outlast the meeting. According to Microsoft's documentation, participants can choose their preferred spoken languages from English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese and Korean.

The same page states that translated captions and transcripts are only live during the meeting and do not persist afterward without tools such as Clipchamp. So if you rely on live translated captions as your record, you may have nothing to circulate afterward.

Two consequences for planning:

  • The language list covers major languages. Teams with Burmese, Thai, Vietnamese, Hausa or Yoruba speakers will not find them in that list.
  • If you need a durable record, capture the audio separately and transcribe it after the call. Our guide to transcribing Google Meet covers the same limits and fallbacks on that platform.

A workflow that works for mixed-language calls

The most reliable setup keeps transcription and summarization as separate steps, each with its own language setting. MinuteKeep describes separating the transcription language from the summary language. The AI reads the source audio in the language it was spoken and outputs a structured summary in a shared language, with no separate pre-translation step.

This avoids the worst failure, which is translating first and summarizing second, so that every mistake in the translation gets baked into the notes. It also suits teams where the working language of the call is local but the report goes to a regional office in English.

Here is a practical routine:

  1. Decide the record language. The transcript should be in the language people actually spoke. Decide separately what language the summary should be in.
  2. Pick the setting that matches the call. If your tool needs a preselected language, choose the one that covers more than half the meeting. If the split is close to even and the tool lacks native code-switching, plan the two-pass method for important recordings.
  3. Improve the audio. Because non-English accuracy drops faster with noise, a decent microphone, a quiet room and one speaker at a time help more than any setting. Ask people to say names and numbers slowly.
  4. Test on your own recording. Take 10 minutes from a real call, including the messy parts, and check how the tool handles the mixed-language stretches and the proper nouns. Do this before you pay for anything.
  5. Fix known terms in one pass. Keep a short list of names and jargon that go wrong, and correct them before the notes are shared.
  6. Review decisions and numbers by hand. Dates, amounts and owners are the details that cost you if they are wrong. A two-minute check of the action items is worth it.
  7. Store the transcript, not just the summary. If someone disputes a decision, the transcript lets you go back to the source. Our guides on writing meeting minutes and minutes templates show how to structure what comes out.

Where Loka Note fits

Loka Note is built for the case this article describes: teams that move between a local language and English in the same meeting. It records in the browser without a bot joining the call, or accepts uploaded audio and video files, and produces a transcript, a summary, decisions and action items.

What it does, according to our own product reference:

  • Transcription with automatic language detection, built Burmese-first, with mixed-language speech such as Burmese-English code-switching handled natively.
  • Seven languages for transcripts and summaries: Burmese, English, Thai, Vietnamese, Chinese (Simplified), Yoruba and Hausa.
  • A check of the chosen recording language against the audio, so a meeting recorded under the wrong language is transcribed in the one actually spoken.
  • Summaries in the meeting's language, plus Ask Loka, which answers questions across your past meetings from your own transcripts.

Pricing has no feature tiers. Minute top-ups start at $1.99 for 60 minutes and are valid for 6 months, and an optional Unlimited plan is $19.99 a month. If you are comparing tools for a mixed-language team, our posts on Granola alternatives for multilingual meetings and how to transcribe Burmese audio go into more detail. As recommended above, test any tool, ours included, on a real recording from your own team first.

Try it on a real mixed-language recording: create a free account at app.lokanote.com/signup

Frequently asked questions

Can AI transcribe a meeting where people speak two languages?

Yes, but results depend on how the tool handles language. Some tools need one language chosen per recording, some detect the language automatically, and fewer handle two languages in the same sentence. If your team switches mid-sentence, test a tool on a real recording before committing.

How do you transcribe code-switching in business meetings?

Use a tool that supports mixed-language speech natively. If yours needs a preselected language, choose the language that takes up more than half of the meeting. For a roughly 50-50 split, run the audio twice, once per language, and merge the stronger sections.

What is the difference between multilingual transcription and speech translation?

Transcription writes down what was said in the language it was spoken. Translation produces text in a different language from the one spoken. Vendors often list both under 'language support', so check which one a feature actually covers.

Why does multilingual transcription fail on non-English accents?

Speech models are trained on far less acoustic variety in non-English data, so they tolerate noise and accent variation less well outside English. Accent robustness is also a separate technical problem from language detection and code-switching, so a tool can be good at one and weak at another.

How does Microsoft Teams handle multilingual meeting transcription?

Teams lets participants choose spoken languages from a short list that includes English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese and Korean. According to Microsoft, translated captions and transcripts are live only during the meeting and do not persist afterward without additional tools.

Sources

  1. 1Multilingual Meeting Transcription: What Worksconvertaudiototext.com
  2. 2Multilingual Meeting Transcription: What Works and What Doesn't | MinuteKeep Blog | 現場コンパスgenbacompass.com
  3. 3Multilingual Meeting Transcription: A Global Team Guidehinoter.com
  4. 4Multilingual Meeting Transcription | MeetStreammeetstream.ai
  5. 5Multilingual transcription & code-switching (99+ languages)assemblyai.com
  6. 6Multilingual speech recognition in Microsoft Teamssupport.microsoft.com

Keep reading