Mandarin transcription tool

Chinese Audio to Text Converter — Free Online Mandarin Transcription

Turn Chinese audio to text right on this page. Upload a Mandarin recording — MP3, WAV, M4A, AAC, OGG, or OPUS — or record from your microphone, and NVIDIA's Nemotron 3.5 ASR model returns clean, punctuated Chinese text in seconds. Free for 3 transcriptions a day with no account and nothing to install. Files run up to 5 minutes and 20 MB on the free tier; Pro extends that to 300 audio minutes a month.

Drop audio hereor choose a file from your device
or
Pro accessHave a license key?
Get Pro for $9/month ↗
Transcript

Your transcript will appear here with punctuation and capitalization.

How to convert Chinese audio to text online

The whole flow happens in your browser and takes under a minute for a typical clip.

  1. Upload your audio file in the transcriber above, or press the microphone button to record directly in the browser.
  2. Set the language to Chinese (Mandarin), or leave auto-detection on if you are not sure what the recording contains.
  3. Pick a speed mode. Best accuracy is the right default for Mandarin; the faster modes trade a little precision for speed.
  4. Click transcribe. The model processes the audio and prints punctuated Chinese text, usually well under real time.
  5. Copy the transcript with one click, or download it as a .txt file.

Mandarin audio to text: language details

The Chinese locale on Nemotron ASR is zh-CN — Mandarin as spoken in mainland China. Mandarin is one of the model's 40 supported locales, and automatic punctuation and sentence formatting are built into the output, so you get readable text rather than a raw character stream.

Auto-detection covers Chinese too. If a recording drifts between Mandarin and English — common in tech talks, business calls, and study-abroad interviews — the model handles mid-recording language switching instead of forcing every word into a single locale.

On script: zh-CN follows mainland writing conventions, so expect Simplified Chinese characters in your transcript. If you need Traditional Chinese for Taiwan or Hong Kong readers, run the finished text through a free Simplified-to-Traditional converter such as OpenCC; script conversion is mechanical and takes seconds.

One honest limit up front: zh-CN means Mandarin. Cantonese, Hokkien, Shanghainese, and other Chinese languages are not supported locales, and accuracy on them will be poor.

WeChat voice messages, lectures, and other Chinese recordings

Most Chinese audio to text requests fall into a few buckets. Here is what works, with the honest caveats.

  • WeChat voice messages: WeChat stores voice notes as .silk or .amr files, formats this converter does not accept directly. Export the message and convert it to MP3 or M4A with a converter app first — or simply play it out loud and capture it with the microphone recorder on this page.
  • Lectures and classes: turn a Mandarin lecture clip into text you can search, annotate, or paste into a translator. Free clips run up to 5 minutes; Pro covers full-length recordings.
  • Interviews and meetings: you get one clean block of punctuated text without speaker labels. Nemotron ASR does not do speaker diarization, so add names manually if you need a labeled Q&A.
  • Voice memos and language practice: record Mandarin dictation straight from the microphone and the text appears as soon as you stop recording.
  • Podcasts and long audio: split episodes into 5-minute chunks on the free tier, or upgrade to Pro for 300 minutes a month.

Chinese audio to text: supported formats and free limits

The converter accepts the common audio formats phones and recorder apps produce. Video files are not supported — extract the audio track first.

FeatureFreePro
PriceFree, no account needed$9/month, Creem license key
Allowance3 transcriptions per day300 audio minutes per month
Clip lengthUp to 5 minutes per fileLonger recordings within your minutes
File sizeUp to 20 MBUp to 20 MB per file
FormatsMP3, WAV, M4A, AAC, OGG, OPUS, or mic recordingSame formats
OutputPunctuated text — copy or .txt downloadSame output
DevicesOne browser connectionUp to 3 browsers

Which speed mode to use for Chinese speech

Nemotron 3.5 ASR is a 0.6B cache-aware FastConformer-RNNT streaming model, and the four modes map to how much audio context it reads per step.

  • Best accuracy (1.12s chunks): the default for Mandarin. Tones and homophones benefit from longer context, so give the model as much as it can get.
  • Balanced (0.56s chunks): a sensible middle ground for clear, close-mic recordings.
  • Faster (0.32s chunks): when you want the transcript quickly and the speech is clean studio-quality audio.
  • Lowest latency (0.08s chunks): built for responsiveness rather than maximum precision; rarely the right pick for tonal languages.

What you get — and what you do not

Plenty of transcription sites bury their limits below the fold. Here is the plain list, so you can decide in ten seconds whether this tool fits your job.

  • You get plain Chinese text with punctuation, ready to copy or download as a .txt file.
  • No timestamps and no SRT or VTT export — this converter is built for transcripts, not subtitle files.
  • No speaker diarization: multi-speaker audio comes back as one continuous text block.
  • No translation: the transcript stays in Chinese. Paste it into DeepL or Google Translate for an English version — two accurate steps beat one blurry one.
  • Recordings transcribe after you press stop, not as a live caption overlay while you speak.
  • This is an independent site running NVIDIA's open Nemotron model — it is not affiliated with NVIDIA.

Frequently asked questions

Does the transcript come out in Simplified or Traditional Chinese?

The zh-CN locale follows mainland Mandarin conventions, so expect Simplified characters in the output. If you need Traditional Chinese, run the transcript through a free converter such as OpenCC — Simplified-to-Traditional conversion is mechanical and takes seconds.

Can it translate Chinese audio to English text?

No. This is a transcription tool: Chinese speech in, Chinese text out. For an English version, copy the finished transcript into a translator like DeepL or Google Translate. Transcribing first and translating second usually beats one-step tools on accuracy.

Can I upload WeChat voice messages directly?

Not directly. WeChat saves voice notes as .silk or .amr files, which are not supported upload formats. Convert the message to MP3 or M4A first, or play it out loud and capture it with the in-browser microphone recorder on this page.

Does it handle Cantonese or other Chinese dialects?

No. The supported Chinese locale is zh-CN, which is Mandarin. Cantonese, Hokkien, and Shanghainese are not in the model's 40-locale list, and results on them will be unreliable.

Is this Chinese audio to text converter really free?

Yes — 3 transcriptions per day, up to 5 minutes and 20 MB per file, with no account and no sign-up. Pro is $9/month for 300 audio minutes when you need longer recordings.

Does it add speaker labels or timestamps?

No. The output is one block of punctuated text without speaker diarization, timestamps, or SRT export. It is built for clean, readable transcripts rather than subtitle files.

What if my recording mixes Mandarin and English?

Leave auto-detection on. The model handles mid-recording language switching, so code-switched sentences — common in business meetings and tech talks — come through without being forced into a single language.