General-purpose models handle major European languages well because the internet is full of them. Nigerian languages are spoken by enormous populations and written online comparatively rarely, so models see little of them and perform accordingly.
Why it is not just a volume problem
- Tone. Yoruba is tonal, and diacritics carry meaning. Text stripped of them is genuinely ambiguous, and most scraped text is stripped.
- Code-switching. Real speech moves between English and a local language mid-sentence. Corpora that treat them as separate miss how people actually communicate.
- Orthographic variation. Spelling conventions differ by region and publication, so the same word appears many ways.
What actually helps
Curated, properly transcribed data beats scraped volume. A few hundred hours of correctly diacritised, consistently transcribed speech outperforms a much larger scrape. That work is slow, needs native speakers, and cannot be automated away, which is exactly why the resulting dataset holds its value.
The commercial case
Voice interfaces that work in the language a customer thinks in reach people that text-first English interfaces never will. For banking, health and agriculture in Nigeria, that is not a niche.