40-language speech recognition

Every voice,on the record.

Upload a recording or speak into your microphone. Get clean, punctuated text in seconds—without choosing the language first.

English中文EspañolFrançaisDeutsch日本語
Live signal09:30:14
AUTO · ZH-CN

现在切换到中文,会议九点半开始。

EN-US

The final code is forty seven.

Transcription desk

Drop in audio. Take out text.

Three free transcriptions per day, up to five minutes each. MP3, WAV, M4A, AAC or OGG.

Drop audio hereor choose a file from your device
or
Transcript

Your transcript will appear here with punctuation and capitalization.

Language40 locales

Automatic detection across English, Chinese, Spanish, Japanese and more.

LatencyStreaming-first

Cache-aware recognition processes new audio without repeating old work.

OutputReady to read

Punctuation and capitalization are produced by the model, not patched on later.

Built for spoken work

Meetings, interviews, voice notes, field recordings.

What powers it

Nemotron 3.5 ASR is a compact, multilingual streaming model from NVIDIA.

It uses a 0.6B-parameter cache-aware FastConformer-RNNT architecture and supports configurable speed/accuracy settings.

Read the official model card ↗