Artificial intelligence
Speech datasets for Nigerian languages with almost no transcribed audio
Speech models fail on languages where very little transcribed audio exists, which covers most of what is actually spoken here.
Modern speech recognition assumes thousands of hours of transcribed audio. For Yoruba, Igbo and Hausa the public pool is small, and for smaller languages it is close to nothing.
The idea we want to test is whether community recorded audio with light, imperfect transcription is enough to fine tune an existing multilingual model to a useful level. Not perfect, just useful enough for a voice interface that people would actually choose over typing.
Open questions we cannot answer alone: how much audio is the real floor, how much transcription error a fine tune tolerates before it degrades, and how to collect this in a way that is fair to the people who contribute their voice.
The idea we want to test is whether community recorded audio with light, imperfect transcription is enough to fine tune an existing multilingual model to a useful level. Not perfect, just useful enough for a voice interface that people would actually choose over typing.
Open questions we cannot answer alone: how much audio is the real floor, how much transcription error a fine tune tolerates before it degrades, and how to collect this in a way that is fair to the people who contribute their voice.
0 contributions
Every contribution is reviewed before it appears.
No one has built on this yet. Add the first contribution below.
Build on this idea
Add what you have tried, point to prior work, or challenge the assumption. Your email stays private and is only used if we need to reach you about what you posted.