Automatic detection across English, Chinese, Spanish, Japanese and more.
Every voice, on the record
Nemotron ASRfree speech to text online.
Upload a recording or speak into your microphone. Nemotron ASR returns clean, punctuated text in seconds — in 40 languages, without choosing the language first.
现在切换到中文,会议九点半开始。
EN-USThe final code is forty seven.
Transcription desk
Free speech to text: drop in audio, take out text.
Three free transcriptions per day, up to five minutes each. MP3, WAV, M4A, AAC, OGG or OPUS.
Your transcript will appear here with punctuation and capitalization.
Cache-aware recognition processes new audio without repeating old work.
Punctuation and capitalization are produced by the model, not patched on later.
What is Nemotron ASR?
Nemotron ASR is a free online speech to text tool built on NVIDIA's Nemotron 3.5 ASR model — a compact, multilingual speech-recognition model designed for streaming transcription. This site wraps that model in a browser tool anyone can use: no code, no GPU, no account. Drop in a lecture recording, a voice memo, an interview or a meeting, and get readable, punctuated text back in seconds. It is an independent project and is not affiliated with NVIDIA; the model itself is documented openly on Hugging Face.
How to transcribe audio with Nemotron ASR
- Add your audio. Drag an MP3, WAV, M4A, AAC, OGG or OPUS file into the transcription desk above — up to 20 MB and five minutes — or press record and speak into your microphone.
- Pick language and mode. Auto-detect handles 40 languages, including recordings that switch language mid-sentence. Choose Best accuracy for interviews or Lowest latency for quick notes.
- Transcribe and take the text. The transcript appears with punctuation and capitalization already in place. Copy it to your clipboard or download it as a .txt file.
Nemotron ASR questions, answered
Is Nemotron ASR really free?
Yes. You get 3 free transcriptions per day — up to 5 minutes of audio each — with no account, no credit card and no watermark. New Pro purchases are currently paused; existing license keys still work.
Which languages does the speech to text support?
40 locales, including English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, Russian, Hindi, Arabic and most European languages. Leave the language on auto-detect and the model identifies it for you, even when a recording mixes languages.
Which audio formats can I upload?
MP3, WAV, M4A, AAC, OGG and OPUS files up to 20 MB and 5 minutes. You can also record directly from your microphone in the browser.
How accurate is the transcription?
It runs NVIDIA's Nemotron 3.5 ASR model, which produces punctuation and capitalization natively rather than patching them on afterwards. Accuracy depends on audio quality — a close microphone and low background noise matter more than anything else. Four speed/accuracy modes let you trade latency for precision.
Does it label different speakers?
No. Speaker diarization is a separate task and this tool does not pretend to do it. You get one continuous, punctuated transcript.
Can I export subtitles (SRT)?
Not yet. The output is plain text you can copy or download as .txt. SRT needs verified timestamps, which are on the roadmap — see the audio to SRT guide for an honest breakdown.
Is my audio stored?
Audio is uploaded over HTTPS for processing and transcription runs server-side; keys never reach your browser. See the privacy policy for exactly what is and is not retained.
Simple usage, visible limits
Start free. Existing Pro licenses remain supported.
The free tool remains available with three runs per day. New Pro purchases are paused, while existing license holders can keep activating and using their paid allowance. See the plan availability page.
$0forever
- 3 transcriptions per day
- Up to 5 minutes per file
- 40 language locales
- Plain-text copy and download
$9per month
- No daily transcription count
- Up to 5 minutes per file
- License works in up to 3 browsers
- Cancel through the Creem customer portal
Already have a Creem license key? Enter it on the activation page. No new checkout is available.
Built for spoken work
Meetings, interviews, voice notes, field recordings.
- 01
Record quicklyCapture a thought without typing it twice.
- 02
Work across languagesLeave language detection on auto for mixed teams.
- 03
Move the text anywhereCopy it or download a plain-text file.
What powers it
NVIDIA Nemotron 3.5 ASR: a compact, multilingual streaming model.
It uses a 0.6B-parameter cache-aware FastConformer-RNNT architecture and supports configurable speed/accuracy settings. Read the full breakdown in our Nemotron 3.5 ASR model guide or see how it stacks up in Nemotron ASR vs Whisper and Nemotron vs Parakeet.
Read the official model card ↗Free transcription tools for every format
The same Nemotron ASR engine powers a set of focused converters — each one explains its format quirks and transcribes right on the page: