ወደ ዋናው ይዘት ዝለል

Named Entity Recognition

Last reviewed: 2026-07-07.

Named Entity Recognition (NER) is the most well-served African-language NLP task. If your project is NER for a language covered by MasakhaNER 2, you probably do not need to build a new corpus — you need to extend an existing one, or fine-tune on top of a released model. This page is here to make that decision fast.

What already exists

The definitive resource is MasakhaNER 2 (Adelani et al., 2022), a 20-language named entity recognition benchmark for African languages, released with data, annotation guidelines, baselines, and evaluation scripts. It supersedes but does not replace MasakhaNER 1 (Adelani et al., 2021), which covers 10 languages and is still worth reading for its retrospective on what changed between versions.

Datasets — per language

LanguageFamilyCorpusSplitsLicenseNotes
Amharic (amh)SemiticMasakhaNER 2train/dev/testCC BY-NC 4.0Ge'ez script; MasakhaNER 1 also covers it
Bambara (bam)MandeMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Ewe (ewe)GbeMasakhaNER 2train/dev/testCC BY-NC 4.0Tone marks; check preprocessing
Fon (fon)GbeMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Ghomala (bbj)GrassfieldsMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Hausa (hau)ChadicMasakhaNER 2train/dev/testCC BY-NC 4.0Ajami variant not covered; Latin only
Igbo (ibo)Volta-NigerMasakhaNER 2train/dev/testCC BY-NC 4.0Diacritics matter for evaluation
Kinyarwanda (kin)BantuMasakhaNER 2train/dev/testCC BY-NC 4.0High-agglutinative — see morphology notes below
Luganda (lug)BantuMasakhaNER 2train/dev/testCC BY-NC 4.0
Luo (luo)NiloticMasakhaNER 2train/dev/testCC BY-NC 4.0
Mossi (mos)GurMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Chichewa/Nyanja (nya)BantuMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Chishona (sna)BantuMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Kiswahili (swa)BantuMasakhaNER 2train/dev/testCC BY-NC 4.0Largest split in the corpus
Setswana (tsn)BantuMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Twi (twi)TanoMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Wolof (wol)SenegambianMasakhaNER 2train/dev/testCC BY-NC 4.0
isiXhosa (xho)BantuMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2
Yoruba (yor)Volta-NigerMasakhaNER 2train/dev/testCC BY-NC 4.0Tone diacritics load-bearing
isiZulu (zul)BantuMasakhaNER 2train/dev/testCC BY-NC 4.0New in v2

The corpus is available on the Hugging Face Hub and the GitHub repository has the annotation guidelines and evaluation scripts.

Editorial opinion. MasakhaNER 2 is the state of the art for African-language NER. If your language is in that table, use it. If your language is not, read the MasakhaNER papers for the annotation guidelines and workforce model before designing your own corpus — they are the reference for how to do this well.

Models — what has been trained on this data

  • AfroXLMR and AfroXLMR-76L — multilingual encoder base models fine-tuned or adapted on African-language data, widely used as the NER fine-tuning starting point.
  • AfriBERTa (Ogueji et al., 2021) — a purpose-built multilingual model pretrained on African-language corpora.
  • SERENGETI (Adebara et al., ACL Findings 2023) — UBC-NLP's massively multilingual foundation model for 517 African languages and language varieties, pretrained on 42GB of African-language text. Explicitly designed as a competitor / complement to AfroXLMR and AfriBERTa. Worth measuring against AfroXLMR-large on your target language before locking in a base — for languages with sparse pretraining data in the XLMR lineage, SERENGETI's coverage can matter.
  • AfriqueLLM (Yu et al. 2026, ACL 2026 Main) — 2026 continued-pretraining recipe applied to Llama 3.1, Gemma 3, Qwen 3 for 20 African languages on 26B tokens. Nine variants (AfriqueQwen 4B/8B/14B, AfriqueGemma 4B/12B, AfriqueLlama 8B). CC BY 4.0. Larger decoder-only alternative to AfroXLMR when compute permits an LLM-based NER pipeline; worth measuring against AfroXLMR fine-tuning on your target language.
  • Per-language fine-tunes for many of the MasakhaNER 2 languages exist on the Hugging Face Hub under masakhane/.
  • XLM-RoBERTa large (Conneau et al., 2020) is a defensible baseline if you want to compare against a general multilingual encoder — the MasakhaNER 2 paper reports its numbers alongside the specialised models.

Editorial opinion. For a new project on a MasakhaNER 2 language, the shortest path to a good NER model is: fine-tune AfroXLMR-large on the MasakhaNER 2 split for your language. That is the recipe MasakhaNER 2 itself used and the numbers in the paper are the honest floor.

Fork or start fresh?

Is your language covered by MasakhaNER 2?
├── Yes — is the label set (PER, ORG, LOC, DATE) enough for your task?
│ ├── Yes → Use MasakhaNER 2. Fine-tune AfroXLMR-large. Done.
│ └── No, you need extra entity types (MONEY, PRODUCT, CULTURE-specific)
│ → Extend MasakhaNER 2 by annotating extra types on their splits.
│ Reuse their guidelines. Do NOT re-annotate PER/ORG/LOC — merge.
└── No — is your language related to any of the 20 covered?
├── Yes (same family / same script)
│ → Start with cross-lingual transfer from the closest MasakhaNER 2
│ language. See the [cross-language transfer](../cross-language-transfer/index.md)
│ chapter for family-by-family pivot guidance. Only then design a
│ small validation corpus in your target language.
└── No — the language is genuinely uncovered.
→ You need to build a corpus. Read the
[long-tail language onboarding](../long-tail-language/index.md)
chapter first — it lays out the realistic milestones and the
Step-0 community, orthography, and IP decisions that determine
whether the project succeeds. Then chapters 2, 3, and 4 of this
playbook (Data Collection, Annotation Design, Data Quality),
and MasakhaNER 2's annotation guidelines before you write your
own. Aim for 5-10k sentences to start, following the MasakhaNER
2 partitioning.

What it will actually cost you

These are order-of-magnitude estimates drawn from the MasakhaNER papers and the participatory workflow they describe. They will not be exact for your project. They are here so nobody submits a proposal claiming NER for a new language is a two-week task.

  • Fine-tuning on an existing language (MasakhaNER 2 covered). One to two person-weeks, one GPU-day. Most of that is evaluation, not training.
  • Extending MasakhaNER 2 with new entity types on covered languages. Four to eight person-weeks per language, depending on how many new types. Includes writing the guideline addendum, training annotators, and adjudication.
  • Building a new NER corpus for an uncovered language, aiming for 5–10k sentences. Three to six months elapsed, one to three person-months of annotator work, plus one person-month of lead annotator effort for guidelines and adjudication. This assumes a participatory setup with two to four native-speaker annotators and one senior annotator.
  • Achieving MasakhaNER-2-comparable inter-annotator agreement on a new language. Budget one full round of annotator recalibration after the first 500 sentences — do not treat the first pass as production.

Known limitations to watch for

  • Diacritics and tone marks. Yoruba, Igbo, and several other languages carry semantic information in diacritics. Losing them in preprocessing changes entity boundaries and inflates apparent error rates. Verify your tokenisation preserves them end-to-end.
  • Ajami and non-Latin scripts. MasakhaNER 2 is Latin-script only. Hausa Ajami, Arabic-script Wolof, and Ge'ez script for Ethiopic languages are not covered. Cross-script transfer is a research problem, not a solved one.
  • Code-switching. Real African-language text mixes languages within a sentence (Hausa-English, Swahili-English, French-Wolof). MasakhaNER 2 splits are monolingual by construction. If your target text is code-switched, the corpus numbers are optimistic.
  • Named entity classes are culturally specific. "Organisation" in a MasakhaNER-2 sentence may not map to the same concept a Nigerian government agency uses. Read the annotation guidelines before assuming the label set fits your use case.

For the mechanics of fine-tuning an encoder model on a NER dataset, use the Hugging Face token-classification tutorial. It is maintained, versioned, and covers everything from datasets loading to evaluation with seqeval. Point your team there rather than writing your own training loop. Your job is the corpus and the evaluation choices, not the boilerplate.

Further reading

Additional references
Contributor
@abumafrim

Join the discussion

Spotted an error, have a question, or want to share what worked on a real project? Sign in with GitHub to add your voice — every thread lives in the open, powered by GitHub Discussions.

Loading discussion…

Thanks to our Contributors

The Playbook is built by a growing community of researchers, students, and language experts. If you've contributed code, content, or review — thank you.

SUPPORTED BY

Masakhane African Languages HubBayero University, KanoBahir Dar UniversityHausaNLPEthioNLP