
If you need Chinese speech to text free online, several tools will transcribe Mandarin audio without a payment, and a few do not even require an account. The catch is in the limits: free minutes, file size, where your audio goes, and whether the output is Simplified characters with usable timestamps. Below we compare seven tools on those points and explain why Mandarin trips up transcribers in ways English does not.
Why Mandarin is harder to transcribe than it looks
Mandarin transcription is harder than English because the same sound can map to many characters, tones change meaning, and written Chinese has no spaces between words. A transcriber has to solve all three at once, using context.
Tones and homophones
Mandarin has four main tones plus a neutral one, and many syllables are pronounced identically across dozens of characters. The model cannot decide from sound alone which character you meant. It has to read the surrounding sentence. ConvertSpeech notes that homophones and tonal variation call for context-aware models.
In practice, a model does well on full, natural sentences and worse on short fragments, names, product codes and isolated numbers. Those are exactly the things a meeting transcript needs to get right.
Segmentation without spaces
English transcripts get word boundaries for free. Chinese does not, so tools need a way to decide where one word ends and the next begins. STT.ai says it uses a custom tokenizer for this reason. Segmentation matters most for search, summaries and subtitles. A transcript that splits words badly is harder to scan and produces awkward subtitle line breaks.
What accuracy numbers mean
Vendors publish figures, and they come from clean audio. STT.ai estimates a word error rate under 6 percent (92 to 96 percent accuracy) on clean Mandarin audio. ElevenLabs reports 3.1 percent on the FLEURS benchmark and 5.5 percent on Common Voice for its Scribe model. These are different tests on different data, so do not rank the tools by comparing them. Treat them as a sign that clean Mandarin is well within reach, and expect worse results on a crowded call with a laggy microphone.
Chinese speech to text free online: seven tools compared
Most of these tools give you a real free allowance, but the allowances differ in kind: minutes, file length, or a daily token budget. Here is what each source states.
| Tool | Free allowance | Signup | Worth noting |
|---|---|---|---|
| STT.ai | 600 free minutes to start | Not required for the first file | Files up to 2 GB; MP3, WAV, M4A, FLAC and video |
| ElevenLabs Scribe | Free Mandarin transcription page | See the page for current terms | Character-level timestamps and speaker diarization |
| ConvertSpeech | Free browser transcription | Not required | Built for Mainland Mandarin, simplified characters |
| Subformer | Free, in-browser | Not required | Runs locally with WebAssembly; roughly 75 MB model cached |
| TTS.ai | Up to 5 minutes of meeting audio | Not stated | Zoom, Teams and Google Meet audio; files up to 500 MB |
| LiteScribe | 120 free minutes per month | No credit card | Speaker labels, AI summaries, TXT and SRT export |
| Free.ai | Anonymous daily token allocation | Not required for anonymous use | faster-whisper; about 50 tokens per minute; auto language detection; SRT export |
How to choose from the table
- One long recording, once: STT.ai's 600 minutes and 2 GB file limit are the most generous on paper. Check the tool's page for the current terms before you plan around the numbers.
- Short meeting clips: TTS.ai's 5-minute meeting allowance is enough to test, not to transcribe a full call.
- Regular monthly use: LiteScribe's 120 minutes a month, with speaker labels and SRT export, suits light recurring needs.
- Sensitive audio: Subformer's local processing is the different design here. The next section explains why that matters.
- Multiple speakers: ElevenLabs lists speaker diarization, and LiteScribe lists automatic speaker labeling.
Free tiers change often. Confirm the limit on the tool's own page on the day you use it.
Cloud upload or in-browser: what happens to your audio
With most online transcribers, your file is uploaded to a server and processed there. With an in-browser tool, the model runs on your own device and the audio stays put. For a business call, that difference matters more than a point or two of accuracy.
Cloud transcription
Cloud tools handle large files and heavy models easily, which is why they tend to lead on speed and accuracy. The tradeoff is that your audio and transcript sit on someone else's infrastructure, often shared with other customers. Before uploading a call about pricing, hiring or a contract, look for a stated retention period, a deletion option, and a clear statement on whether your data is used to train models. Our post on whether AI note takers are safe lists ten questions worth asking.
In-browser WebAssembly
Subformer runs speech-to-text inside the browser using WebAssembly and a locally cached model of about 75 MB. According to its page, it processes audio offline without uploading files to external GPUs. The tradeoffs are practical. The first load downloads the model, a small model on a weak laptop will be slower, and you are limited to the model that fits in a browser tab. For a private 20-minute call, that can be an acceptable price.
A simple rule
If the recording would be a problem in a stranger's hands, use a local, in-browser tool or a service whose retention and data-use terms you have actually read. If it is a public webinar, use whichever tool is most accurate.
Mandarin, Cantonese and regional accents: avoiding wrong-language output
Auto-detection is the weakest part of many free tools. A transcriber that decides the language from the first few seconds can mislabel Cantonese, a heavy regional accent, or a call that mixes Mandarin and English.
Mandarin vs Cantonese
Mandarin and Cantonese are spoken varieties that differ in vocabulary, grammar and pronunciation. A model trained for Mandarin will produce confusing text from Cantonese speech, often with plausible-looking characters that mean something else. Free.ai offers automatic Chinese language detection, which is convenient, but we would not rely on any auto-detect feature to separate the two. Set the language yourself where you can. If the tool only offers "Chinese," run a 30-second sample and read it before uploading the full file.
Accents within Mandarin
Putonghua, the standard Mainland form, is what most models are tuned to. Speakers from regions with strong local accents, or speakers who pronounce certain sounds differently, may see more errors. This is where a human read-through pays off, especially on names and places.
Simplified or Traditional
Output script varies by tool. ConvertSpeech states that its Mainland Mandarin option produces simplified characters. If your readers use Traditional characters, check whether the tool offers that output before processing long files. Converting afterward is possible, but some word choices differ between regions, so a script converter will not fix everything.
Mixed Mandarin and English
Teams in Southeast Asia and Nigeria often switch languages mid-sentence. Most free single-language tools struggle when a speaker drops English product names into a Mandarin sentence. Our guide to multilingual meeting transcription covers what tends to work.
How to turn Chinese audio or video into SRT subtitles
You can go from a Mandarin recording to an SRT subtitle file in about five steps with any tool that exports SRT. Free.ai and LiteScribe both list SRT export.
- Get the cleanest source file. Use the original recording, not a re-compressed copy shared through a chat app. If the file is a large video, you can extract only the audio first; our post on transcribing video to text covers that for slow connections.
- Check the limits. Compare the file size and length against the tool's free limit. STT.ai lists up to 2 GB and TTS.ai up to 500 MB on its free tier.
- Set the language to Mandarin. Choose it manually if the option exists, and confirm the script (Simplified or Traditional).
- Run a short test. Upload the first minute or two and read it. Look at names, numbers and any English terms.
- Transcribe, then export SRT. Download the subtitle file, and open it in a text editor to spot-check line breaks and timestamps before burning it into a video.
Fixing common subtitle problems
- Lines too long: Chinese subtitles read best in short lines. Split any line that runs past one screen width.
- Wrong homophones: Search the file for your company, product and people names, and correct them in one pass.
- Drifting timestamps: If timing slips late in a long file, split the audio into shorter parts and transcribe each separately.
Transcribing Zoom and Teams meetings in Mandarin
The most reliable route is to record the meeting, export the audio or video, and upload the file to a Mandarin-capable tool. This avoids relying on a meeting platform's built-in captions, which may not cover Mandarin well on every plan.
- Record the meeting from the platform or from your device, with participants' consent.
- Export the file as audio (MP3, M4A or WAV) or as video.
- Upload it to a transcriber that supports Mandarin and speaker labels.
- Read the first five minutes, then skim the rest for names, dates and numbers.
- Pull out decisions and owners. Our post on action items from meetings has a structure for this.
TTS.ai lists support for meeting audio from Zoom, Teams and Google Meet, with a 5-minute free limit, which is useful for a quick trial. For a full meeting you will need a larger allowance or a paid top-up. If you want to avoid a bot joining the call, see our notes on meeting transcription without a bot, and for Zoom summaries specifically, how to summarize a Zoom meeting.
Cleaner audio for noisy, multi-speaker meetings
Audio quality is the biggest factor you control. The accuracy figures vendors publish come from clean audio, and a noisy conference-room recording will not match them.
- One microphone per person where possible. A shared speakerphone blurs voices together, which hurts both transcription and speaker labels.
- Mute when not speaking. Keyboard noise and side conversations get transcribed as words.
- Ask people to finish sentences. Homophones are resolved by context, so clipped fragments lose accuracy.
- Say numbers and names twice, or write them in chat. Names and figures are where Mandarin homophones do the most damage.
- Avoid overlapping speech. Diarization helps with who said what, but overlapping voices still produce garbled lines.
- Record in a quiet room with soft surfaces. Echo is hard for any model to undo afterward.
After transcription, do a two-pass check: first for names and numbers, then for meaning. Our interview transcription guide describes a two-pass approach that works for meetings too.
Where Loka Note fits
Loka Note is not a free-forever tool, but it covers Chinese (Simplified) along with Burmese, English, Thai, Vietnamese, Yoruba and Hausa. It records meetings in the browser without a bot joining the call, or you can upload audio and video files. It then produces a transcript, a summary, decisions and action items. The language you pick is checked against the audio, so a meeting recorded under the wrong language is transcribed in the one actually spoken.
There are no plan tiers. Minutes are sold as one-time top-ups starting at $1.99 for 60 minutes, valid for six months, and an optional Unlimited plan costs $19.99 a month. If your team switches between languages, Loka handles mixed-language speech in a single meeting. Customer audio and transcripts are never used to train AI models, and you can delete any meeting at any time; the details are on our security page. Loka does not output Traditional characters, so for that use case check another tool first.
Try a Mandarin recording at Loka Note
Frequently asked questions
Is Chinese speech to text free online without signing up?
Yes, for some tools. STT.ai offers 600 free minutes and does not require signup for the first file. ConvertSpeech and Subformer both advertise free browser transcription without signup. Longer or repeated use usually needs an account.
How do free AI transcription tools handle Mandarin tones and homophones?
They rely on context-aware models rather than sound alone, because many Mandarin words share the same pronunciation. The model picks the character sequence that makes the most sense in the sentence. That is why clean audio and full sentences matter more than slow, isolated words.
Can free online transcribers distinguish between Mandarin and Cantonese?
Do not assume so. Many tools label everything as Chinese, and auto-detection can pick the wrong variety. Set the language manually where the tool allows it, and test a 30-second clip before committing a long file.
Does Chinese speech to text output Simplified or Traditional characters?
It depends on the tool. ConvertSpeech states that its Mainland Mandarin option produces simplified characters. Loka Note lists Chinese (Simplified). If you need Traditional characters, check the tool's output options before uploading anything long.
What is the best way to transcribe Zoom or Teams meetings conducted in Mandarin?
Record the meeting, export the audio or video file, and upload it to a Mandarin-capable transcriber. Ask each person to join from their own device and mute when not speaking. Check the first few minutes of the transcript for names and numbers before trusting the rest.
Sources
- 1Chinese (Mandarin) Speech to Text - Free Online Transcription | STT.aistt.ai
- 2Free Mandarin Chinese Speech to Text Transcriptionelevenlabs.io
- 3Chinese (Mandarin) Speech to Text — Free Transcription | ConvertSpeechconvertspeech.com
- 4Free Mandarin Speech to Text - In-Browser, No Signup | Subformersubformer.com
- 5Free AI Speech to Text Online - Transcribe Audio & Video | TTS.aitts.ai
- 6Chinese Transcription: 120 Minutes Free - LiteScribelitescribe.ai
- 7Transcribe Chinese Audio Free | Free.aifree.ai


