Building machine-readable datasets and translation models for Nigerian languages including Igbo and Nigerian Pidgin, where data scarcity remains a major barrier to equitable AI.
Nigeria is one of the most populous and linguistically diverse countries in Africa, with over 500 languages spoken. Yet the vast majority of NLP research focuses on high-resource languages like English and Mandarin, leaving Nigerian language speakers severely underserved by modern AI tools.
Our work in this area focuses on creating high-quality, machine-readable datasets for low-resource Nigerian languages — particularly Igbo and Nigerian Pidgin — and developing neural machine translation (NMT) models that can bridge these languages with English and other widely-spoken languages.
We survey the landscape of machine translation research on Nigerian languages, identify gaps in existing datasets and methodologies, and propose future directions for growing the community of researchers working on African language technologies.