Finding current resources
Last reviewed: 2026-07-07.
The Before You Start pages tell you which African-language NLP resources exist and what our editorial opinion is on each. Those pages age. Datasets get superseded, models get deprecated, new work drops at every AfricaNLP workshop. This chapter is the small companion that stays honest for years: it names the primary sources — the organisations, archives, workshops, and search patterns — where the truth lives, and points you at them.
The rule that makes this chapter survive without a maintainer: we name sources, not contents. A Hugging Face organisation page updates itself; a Zenodo community indexes new uploads automatically; a workshop's proceedings arrive on schedule every year. We do not try to list what is on those pages today. We point you at them and trust the sources to be current.
When to use this chapter
- The Before You Start page for your task was last reviewed more than six months ago, and you need to know what has shipped since. The dates at the top of each Before You Start page are the tripwire.
- Your task is not covered by a Before You Start page — you need to find your own primary sources.
- You are scoping a new project and want the current state of the ecosystem, not a snapshot.
Where the corpora and models actually live
The Hugging Face organisations to trust
These are the organisations whose model cards and dataset cards are the primary source of truth for African-language NLP resources. Bookmark them.
- huggingface.co/masakhane — the community organisation. MasakhaNER 1 and 2, AfriSenti, LAFAND-MT, AfriQA, plus per-language fine-tunes and community-curated corpora. Start here for any African-language dataset or model search.
- huggingface.co/hausanlp — the Hausa-focused community organisation. Hausa language models, sentiment and hate-speech corpora, and per-task fine-tunes centred on Hausa specifically. Start here when your project is Hausa-first.
- huggingface.co/ethionlp — the Ethiopian-language community organisation. Amharic, Tigrinya, Afaan Oromo, and other Ethiopic-family resources; the Ge'ez-script-first counterpart to what Masakhane maintains for the broader continent. Start here when your project is Ethiopian-language-first.
- huggingface.co/sunbird — Sunbird AI. Luganda ASR, MT, TTS, and integrated speech-and-text pipelines. The working reference for a small team shipping African-language NLP in production.
- huggingface.co/Davlan — David Adelani's organisation. AfroXLMR base and large variants (including 76-language), per-language fine-tunes across NER, sentiment, MT.
- huggingface.co/castorini — Waterloo's IR + multilingual research. AfriBERTa, mDPR, multilingual retrieval and QA baselines.
- huggingface.co/facebook — Meta AI research releases. MMS (ASR + TTS, 1000+ languages), NLLB-200 (translation, 200 languages), SeamlessM4T, XLS-R.
- huggingface.co/google — Google research releases. FLEURS evaluation benchmark, Gemma multilingual variants, mT5.
- huggingface.co/CohereForAI — Cohere For AI research. Aya multilingual instruction-tuned models with meaningful African-language coverage.
- huggingface.co/openai — Whisper family (ASR). Coverage claims are marketing-broad; verify per language.