Need the words from a video, not the video itself? This free video to text converter listens to the soundtrack of an MP4, MOV or WebM file and writes out what’s said, with timestamps, in a few minutes. It works on the video files on your phone or computer, and nothing is uploaded.
Video to text converter
Get a written transcript from the sound in an MP4, MOV or WebM video: interviews, sermons, tutorials, skits or Reels. Your video stays on your device.
- No upload
- No watermark
- No sign-up
- Free
Choose your audio or video
Speed and data options
Models are saved in your browser after the first download so you don't pay for the data twice.
Your file never leaves this device: speech recognition (OpenAI Whisper) runs inside your browser. Only you must have the rights to the media you use. Keep this tab open while it works.
Check and edit
How it works
- Choose an audio or video file from your phone or computer. It is read by your browser only; nothing is uploaded.
- Pick the spoken language and a model. Tiny is fastest; Base is more accurate but needs a stronger device. The first time you use it, your browser downloads the model once (about 41 MB for Tiny or 77 MB for Base) and keeps it for next time. Use Wi-Fi for that first download if data is expensive.
- Whisper turns the speech into text with timings, about 30 seconds of audio at a time, and you can watch the progress.
- Fix any wrong words or timings in the editor, play a line to check it, and choose short Reel-style captions or longer subtitle lines.
- Download TXT, SRT or VTT files, or burn the captions into a video of up to 60 seconds with no watermark.
Frequently asked questions
Can I paste a YouTube or TikTok link?
No. We do not download videos from other sites, because that breaks their terms. Use a video file you own or have permission to use, for example one from your phone's gallery or your editing app.
Which video formats work?
MP4 and MOV from phones and cameras and WebM from screen recorders work in most browsers. The tool reads only the audio track. Very large files (over 300 MB) are blocked so your browser does not run out of memory.
How long can my file be?
Up to 30 minutes for transcription and up to 60 seconds for burned-in captions (exported at up to 720p). On older phones we recommend files under 10 minutes and the Tiny model. Longer files, batch jobs and 1080p are planned for Caption Studio Pro.
Can I turn the transcript into captions?
Yes. After transcribing, choose a caption length and download an SRT or VTT file, or burn the captions straight into the video if it is 60 seconds or shorter.
Is my audio uploaded anywhere?
Your audio and video never leave your device. The speech recognition model (OpenAI Whisper) is downloaded to your browser and runs there, so nothing is uploaded to Centertainment or anyone else. The only downloads are the model files (from Hugging Face) and the speech engine (from the jsDelivr CDN), which your browser fetches like any other web file.
Why is it slow on my phone?
Speech recognition is heavy work and your device is doing all of it. Chrome on a recent Android phone or any laptop works best. Keep the tab open and the screen on, close other apps, and use the Tiny model. As a rough guide, a laptop transcribes a minute of audio in well under a minute with Tiny; a budget phone can take several times longer.
It’s handy for turning interviews into articles, pulling quotes from a press conference, writing a blog post from a tutorial, making notes from a church service or lecture, or getting a script from your own skits and Reels so you can repost them as text.
How it handles video
The converter reads only the audio track. Your browser extracts the sound, the Whisper speech model transcribes it on your device, and you get editable text you can download as plain TXT, TXT with timestamps, or SRT/VTT subtitles. Free files can be up to 30 minutes long and 300 MB. If your video is bigger, trim it or export just the audio first.
We don’t download from other sites
You can’t paste a YouTube, TikTok or Instagram link here, and that’s deliberate. Downloading other people’s videos breaks those platforms’ terms, and often copyright too. Use a video you made or have permission to use.
Tips for a cleaner transcript
- Choose the language that’s spoken. For mixed English and Pidgin, English usually works best.
- Use the Base model for accuracy, or Tiny on slower phones and for long videos.
- Background music lowers accuracy. If you have the original voice recording, transcribe that instead.
- Read the result before you publish it, especially names, numbers and quotes.
Want the text back on the video as captions? Use the auto caption generator or the full Caption & Transcript Studio.
