Yoruba audio to text (beta)

Speech-to-text tools have been built mostly for English, which leaves a lot of Nigerian creators, journalists and families out. This page lets you transcribe Yoruba audio to text or try turning it into English, free, in your browser. It’s an early beta, and we want to be upfront about what that means.

Yoruba audio to text (beta)

Transcribe Yoruba speech in your browser, or translate it into English text. This is an early beta: the small models that run on a phone get many Yoruba words wrong, so use it for a first draft and timings, then correct every line.

  • No upload
  • No watermark
  • No sign-up
  • Free

Choose your audio or video

Speed and data options

Models are saved in your browser after the first download so you don't pay for the data twice.

Your file never leaves this device: speech recognition (OpenAI Whisper) runs inside your browser. Only you must have the rights to the media you use. Keep this tab open while it works.

How it works

  1. Choose your audio or video. Yoruba is already selected, and the Base model is recommended for it (Tiny is much weaker in Yoruba).
  2. Choose "Transcribe" for Yoruba text, or "Translate to English" to get an English version of what was said.
  3. The first time you use it, your browser downloads the model once (about 41 MB for Tiny or 77 MB for Base) and keeps it for next time. Use Wi-Fi for that first download if data is expensive.
  4. Correct the text in the editor. Tone marks (such as à, é, ẹ, ọ, ṣ) are often missing from the result, so add them where they matter.
  5. Download TXT, SRT or VTT, or burn captions into a short video.

Frequently asked questions

How good is it at Yoruba?

Honestly, not good yet. In our own test with a short, clearly spoken Yoruba proverb, the Base model got most words wrong and the English translation missed the meaning. OpenAI's Whisper was trained on far less Yoruba than English, and the small versions that can run in a browser are its weakest. Where it helps today is timing: it splits your recording into timed lines, so you can type the correct Yoruba into each line and export a proper SRT file.

Can it translate Yoruba to English?

You can try. Choose "Translate to English" before you start and Whisper writes English text with timings. With the small browser models the translation is often wrong, so only use it as a rough hint and never publish it unchecked.

What about Hausa and Igbo?

Hausa is in the language list as a beta, like Yoruba. Igbo is experimental because Whisper was not trained on Igbo, so results will be poor; for now, translating to English often gives a more useful result.

Why are tone marks missing?

A lot of Yoruba text online is written without tone marks, and the model learned from that. Add the marks in the editor where the meaning depends on them, for example to tell ọkọ (husband) from ọkọ̀ (vehicle).

Is my audio uploaded anywhere?

Your audio and video never leave your device. The speech recognition model (OpenAI Whisper) is downloaded to your browser and runs there, so nothing is uploaded to Centertainment or anyone else. The only downloads are the model files (from Hugging Face) and the speech engine (from the jsDelivr CDN), which your browser fetches like any other web file.

Is it really free with no watermark?

Yes. Transcripts, subtitle files and burned-in captions are free, with no sign-up and no watermark. You can add a small "Captions: centertainment.site" credit if you want to, but it is off by default. It costs us nothing because the work happens on your own device.

What works, and what doesn’t yet

The tool uses OpenAI’s open-source Whisper model, which does include Yoruba. But Whisper learned from far less Yoruba than English, and the small versions that can run on a phone are its weakest. We tested it on a short, clearly spoken Yoruba proverb, and the Base model still got most of the words wrong. Tone marks are usually missing too. So, for now, don’t expect a finished transcript.

What it does well already is the timing. It splits your recording into timed lines, and you can type the correct Yoruba into each one, add the tone marks, play any line to check it, and download a proper SRT file for YouTube, CapCut or VN. That’s much faster than timing subtitles by hand.

Transcribe or translate

  • Transcribe writes Yoruba text, for captions in Yoruba or for your records.
  • Translate to English writes English text with timings. With the small browser models the translation is often wrong, so use it only as a rough hint and never publish it unchecked.

The Base model is selected for Yoruba because it’s noticeably more accurate than Tiny. It’s a bigger one-off download (about 77 MB), so use Wi-Fi the first time.

Hausa and Igbo

Hausa is in the language list as a beta too. Igbo is marked experimental: Whisper wasn’t trained on Igbo, so results will be poor, and translating to English often gives a more useful result. We’ll update these pages as better open models for Nigerian languages arrive.

Your recordings stay private because everything runs on your device. When you’re done, download TXT, SRT or VTT, or add captions to a video with the auto caption generator. For English recordings, use transcribe audio to text.