Automatic detection across English, Chinese, Spanish, Japanese and more.
40-language speech recognition
Every voice,on the record.
Upload a recording or speak into your microphone. Get clean, punctuated text in seconds—without choosing the language first.
现在切换到中文,会议九点半开始。
EN-USThe final code is forty seven.
Transcription desk
Drop in audio. Take out text.
Three free transcriptions per day, up to five minutes each. MP3, WAV, M4A, AAC or OGG.
Your transcript will appear here with punctuation and capitalization.
Cache-aware recognition processes new audio without repeating old work.
Punctuation and capitalization are produced by the model, not patched on later.
Built for spoken work
Meetings, interviews, voice notes, field recordings.
- 01
Record quicklyCapture a thought without typing it twice.
- 02
Work across languagesLeave language detection on auto for mixed teams.
- 03
Move the text anywhereCopy it or download a plain-text file.
What powers it
Nemotron 3.5 ASR is a compact, multilingual streaming model from NVIDIA.
It uses a 0.6B-parameter cache-aware FastConformer-RNNT architecture and supports configurable speed/accuracy settings.
Read the official model card ↗