Artificial intelligence

Speech datasets for Nigerian languages with almost no transcribed audio

Dynaiz Research , Dynaiz Technology · · 1 min read

Speech models fail on languages where very little transcribed audio exists, which covers most of what is actually spoken here.

Modern speech recognition assumes thousands of hours of transcribed audio. For Yoruba, Igbo and Hausa the public pool is small, and for smaller languages it is close to nothing.

The idea we want to test is whether community recorded audio with light, imperfect transcription is enough to fine tune an existing multilingual model to a useful level. Not perfect, just useful enough for a voice interface that people would actually choose over typing.

Open questions we cannot answer alone: how much audio is the real floor, how much transcription error a fine tune tolerates before it degrades, and how to collect this in a way that is fair to the people who contribute their voice.

0 contributions

Every contribution is reviewed before it appears.

No one has built on this yet. Add the first contribution below.

Build on this idea

Add what you have tried, point to prior work, or challenge the assumption. Your email stays private and is only used if we need to reach you about what you posted.

Never shown publicly.

Reviewed before it goes live, usually within a day.