Artificial Intelligence (AI) has emerged as one of the most transformative technologies of the twenty-first century, rapidly reshaping academic fields and research methodologies. In African historiography, where oral traditions, written document, archaeologica…
Background: Morphological awareness (MA) is widely recognised as contributing to successful reading and reading comprehension. Despite this, research examining the nature of morphological errors among second-language (L2) English learners remains limited, part…
Phishing conducted in Swahili has become a persistent threat to the millions of Tanzanians who depend on mobile-money services, yet the detection tools in common use are built for English and transfer poorly to a language whose morphology, register, and transa…
Social power plays a fundamental role in shaping human interaction, yet computational studies of power remain limited to narrow linguistic and cultural settings. Existing datasets further lack the demographic and relational depth needed for robust cross-cultur…
Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre
We present KinyaEmbed, the first dedicated sentence embedding model for Kinyarwanda, a morphologically rich Bantu language spoken by over 12 million people in Rwanda. Existing multilingual embedding models such as LaBSE, mE5-large, and OpenAI text-embedding-3-…
Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre
We present TabuLM, the first language model pre-trained on Kinyarwanda tabular data. Kinyarwanda is a morphologically rich Bantu language spoken by over 12 million people in Rwanda, yet lacks any dedicated tabular representation learning resource. TabuLM exten…
Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speec…
Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 detoxified…
We measure tablet-2, a production long-term memory engine for language models, on the text benchmarks the field already uses and on cross-lingual retrieval of photographs stored with no text at all. Its retrieval path contains no lexical matching, no keyword s…
The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities. In this paper, we characterise this gap through the lens of representational geometry. Comparing …