ወደ ዋናው ይዘት ዝለል

From the archive

AfricaNLP Progress

Week 3524 – 30 Aug 202612 papers

Languages
JournalEast African Journal of Information Technology

A Machine-Learning Model for Phishing Detection in Swahili Messages: A Case of Tanzania

Rehema Abdallah Njame, Gustaph Sanga, I. Tende

Phishing conducted in Swahili has become a persistent threat to the millions of Tanzanians who depend on mobile-money services, yet the detection tools in common use are built for English and transfer poorly to a language whose morphology, register, and transa

arXiv

TabuLM: Morphology-Aware Tabular Pre-training for Low-Resource Languages

Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre

We present TabuLM, the first language model pre-trained on Kinyarwanda tabular data. Kinyarwanda is a morphologically rich Bantu language spoken by over 12 million people in Rwanda, yet lacks any dedicated tabular representation learning resource. TabuLM exten

View on arXivPDFS2
arXiv

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

Mo El-Haj

Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 detoxified

View on arXivPDFS2
arXiv

The Geometry of Low-Resource Language Representations

Francois Meyer, Jan Buys

The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities. In this paper, we characterise this gap through the lens of representational geometry. Comparing

View on arXivPDFS2

How this works. Every Monday we crawl arXiv and Semantic Scholar for papers that name African languages, countries, or communities, score them, and ask Claude to summarise the ones that are substantively about African languages. A maintainer reviews the list before it is published. Missed a paper? Tell us, or get a monthly roundup from the newsletter.