Most free caption apps make you pay in some other way: a watermark across your Reel, a daily limit, a forced sign-up, or your audio sitting on somebody else’s server. The Centertainment Caption & Transcript Studio does it differently. The speech recognition runs inside your own browser, so your files never leave your phone or laptop, and there is no watermark to remove.
Caption & Transcript Studio
Transcribe audio or video, edit the words, export TXT, SRT or VTT, and burn bold captions into your Reels. Free, no sign-up, no watermark, and your files never leave your device.
- No upload
- No watermark
- No sign-up
- Free
Choose your audio or video
Speed and data options
Models are saved in your browser after the first download so you don't pay for the data twice.
Your file never leaves this device: speech recognition (OpenAI Whisper) runs inside your browser. Only you must have the rights to the media you use. Keep this tab open while it works.
Check and edit
Download
Burn captions into your video
Videos up to 60 seconds, saved at up to 720p with no watermark. Rendering takes about as long as the video.
How it works
- Choose an audio or video file from your phone or computer. It is read by your browser only; nothing is uploaded.
- Pick the spoken language and a model. Tiny is fastest; Base is more accurate but needs a stronger device. The first time you use it, your browser downloads the model once (about 41 MB for Tiny or 77 MB for Base) and keeps it for next time. Use Wi-Fi for that first download if data is expensive.
- Whisper turns the speech into text with timings, about 30 seconds of audio at a time, and you can watch the progress.
- Fix any wrong words or timings in the editor, play a line to check it, and choose short Reel-style captions or longer subtitle lines.
- Download TXT, SRT or VTT files, or burn the captions into a video of up to 60 seconds with no watermark.
Frequently asked questions
Is it really free with no watermark?
Yes. Transcripts, subtitle files and burned-in captions are free, with no sign-up and no watermark. You can add a small "Captions: centertainment.site" credit if you want to, but it is off by default. It costs us nothing because the work happens on your own device.
Is my audio uploaded anywhere?
Your audio and video never leave your device. The speech recognition model (OpenAI Whisper) is downloaded to your browser and runs there, so nothing is uploaded to Centertainment or anyone else. The only downloads are the model files (from Hugging Face) and the speech engine (from the jsDelivr CDN), which your browser fetches like any other web file.
How long can my file be?
Up to 30 minutes for transcription and up to 60 seconds for burned-in captions (exported at up to 720p). On older phones we recommend files under 10 minutes and the Tiny model. Longer files, batch jobs and 1080p are planned for Caption Studio Pro.
Does it work for Yoruba, Hausa, Igbo and Pidgin?
English, including Nigerian accents, works best. Yoruba and Hausa are an early beta: Whisper supports them, but its small models get many words wrong and miss tone marks, so treat the result as a draft with timings and correct it. Igbo is experimental because Whisper was not trained on Igbo. For Nigerian Pidgin, choose English. If you only need the meaning, try "Translate to English".
What is the difference between SRT and VTT?
Both are plain-text subtitle files with numbered lines and timings. SRT is the most widely accepted (YouTube, Facebook, CapCut, VN, Premiere, DaVinci Resolve). VTT is the web format used by HTML5 video players and some platforms. They hold the same captions, so download whichever your app asks for.
Does it work on iPhone?
Transcription works in recent versions of Safari, but iPhones limit how much memory a web page can use, so we start with the Base model and suggest keeping files short. If your iPhone struggles or the page reloads, switch to the faster Tiny model. Burned-in captions need video recording support in the browser; if your iPhone cannot record, download the SRT file and add it in CapCut or VN instead.
It covers the whole job in four steps:
- Transcribe: choose an audio or video file (voice notes, interviews, podcasts, skits, sermons or Reels) and the studio turns the speech into text with timings.
- Edit: fix words and timings line by line, play any line to check it, and switch between short Reel captions and longer subtitle lines in one tap.
- Export: download an SRT or VTT subtitle file for YouTube, Facebook, CapCut or VN, or a TXT transcript with or without timestamps.
- Burn in: add bold captions to a video of up to 60 seconds, with a red word-by-word highlight, a red box, classic white on black or a white outline.
Why it runs in your browser
The studio uses OpenAI’s open-source Whisper speech model, loaded into your browser with Transformers.js. The first time you use it, your browser downloads the model (about 41 MB for Tiny or 77 MB for Base) and saves it for next time. After that, everything happens on your device. That keeps your recordings private, and it means we can offer the tool free with no daily cap.
The trade-off is speed: your device does the work. A laptop or recent Android phone handles a few minutes of audio comfortably, while budget phones are slower. The studio warns you if your device looks underpowered and suggests the Tiny model. Free files can be up to 30 minutes long.
Languages
English works best, and that includes Nigerian accents and Pidgin-flavoured English. Yoruba and Hausa are available as an early beta: Whisper supports them, but its small models get many words wrong and miss tone marks, so use the result as a timed draft and correct it. Igbo is experimental because Whisper wasn’t trained on it. French, Spanish, Portuguese, Arabic, Swahili, Hindi, German and Indonesian are also available, along with a “Translate to English” option.
Need just one part of the job? Try the auto caption generator, transcribe audio to text, the SRT subtitle generator, video to text or Yoruba audio to text. Then see what your content could earn with our creator earnings calculators.
