Transcribe audio to text free

Typing out a recording by hand takes about four times as long as the recording itself. This free tool does the first draft for you. Choose an audio file and it will transcribe audio to text in your browser, ready to edit, copy or download, with no account and no daily file limit.

Transcribe audio to text (free)

Turn a voice note, interview, podcast or lecture into editable text with timestamps. It runs in your browser, so there is no upload, no sign-up and no daily limit.

  • No upload
  • No watermark
  • No sign-up
  • Free

Choose your audio or video

Speed and data options

Models are saved in your browser after the first download so you don't pay for the data twice.

Your file never leaves this device: speech recognition (OpenAI Whisper) runs inside your browser. Only you must have the rights to the media you use. Keep this tab open while it works.

How it works

  1. Choose an audio or video file from your phone or computer. It is read by your browser only; nothing is uploaded.
  2. Pick the spoken language and a model. Tiny is fastest; Base is more accurate but needs a stronger device. The first time you use it, your browser downloads the model once (about 41 MB for Tiny or 77 MB for Base) and keeps it for next time. Use Wi-Fi for that first download if data is expensive.
  3. Whisper turns the speech into text with timings, about 30 seconds of audio at a time, and you can watch the progress.
  4. Fix any wrong words or timings in the editor, play a line to check it, and choose short Reel-style captions or longer subtitle lines.
  5. Download TXT, SRT or VTT files, or burn the captions into a video of up to 60 seconds with no watermark.

Frequently asked questions

Is there a daily limit?

No. Because transcription runs on your own device, there is no daily file limit and no account. Each file can be up to 30 minutes long.

Which audio formats work?

MP3, WAV, M4A/AAC, OGG and Opus (including WhatsApp voice notes), FLAC and WebM, plus the sound from MP4, MOV and WebM videos. Support for unusual formats depends on your browser; if a file will not open, convert it to MP3 first.

Can I get the transcript with timestamps?

Yes. Download "TXT with times" for a transcript with a timestamp on every line, plain "TXT" for clean paragraphs, or SRT/VTT subtitle files.

How accurate is it?

For clear English speech the Base model is usually very good; Tiny is faster but makes more mistakes. Background music, crosstalk and strong noise reduce accuracy. Always read the transcript before you publish or quote it.

Is my audio uploaded anywhere?

Your audio and video never leave your device. The speech recognition model (OpenAI Whisper) is downloaded to your browser and runs there, so nothing is uploaded to Centertainment or anyone else. The only downloads are the model files (from Hugging Face) and the speech engine (from the jsDelivr CDN), which your browser fetches like any other web file.

Why is it slow on my phone?

Speech recognition is heavy work and your device is doing all of it. Chrome on a recent Android phone or any laptop works best. Keep the tab open and the screen on, close other apps, and use the Tiny model. As a rough guide, a laptop transcribes a minute of audio in well under a minute with Tiny; a budget phone can take several times longer.

It works with the formats you actually have on your phone: MP3, M4A, WAV, OGG and FLAC, and the .opus files that WhatsApp uses for voice notes. You can also choose a video and the tool will transcribe its soundtrack.

Who it’s for

  • Journalists and bloggers: transcribe interviews and press conferences, then pull quotes from the timestamped version.
  • Podcasters: create show notes and searchable transcripts for every episode.
  • Students: turn recorded lectures and group discussions into notes you can search.
  • Anyone with long voice notes: read that five-minute message instead of playing it in a noisy bus.

Private by design

Most online transcription services upload your audio to their servers. Here, OpenAI’s open-source Whisper model runs on your own device, so the recording never leaves it. That makes it a good choice for sensitive interviews, but still check the text before you publish or quote it, because no speech recognition is perfect.

Getting the best result

Pick the language that’s spoken in the recording. Use Base for better accuracy or Tiny if your phone is slow or the file is long. Clear audio with one person speaking at a time gives the best result; music and crosstalk lower accuracy. When it’s done, download plain TXT, TXT with times, or SRT/VTT subtitles.

For Yoruba recordings, see Yoruba audio to text (beta). To put the words on a video, use the auto caption generator.