<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
  <channel>
    <title>AfricaNLP Progress</title>
    <link>https://waraka.org/progress</link>
    <description>Weekly digest of new research on African languages, curated by the Waraka community.</description>
    <language>en</language>
    <item>
      <title>INVESTIGATING ASPECTS OF PREFIXATION IN ENGLISH AND IGBO LANGUAGES</title>
      <link>https://doi.org/10.30574/wjarr.2026.31.2.2125</link>
      <guid isPermaLink="false">doi:10.30574/wjarr.2026.31.2.2125</guid>
      <pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W36</category>
      <description>Prefixation is a morphological process which studies how morphemes are affixed in the beginning of their host words for either grammatical functions or for lexical expansion and derivation. This paper investigates aspects of the grammar of prefixation in English and Igbo Languages. It adopts a comparative theoretical framework, so as to identify the morphological similarities and difference in bot</description>
    </item>
    <item>
      <title>Where Do Multilingual Vision-Language Encoders Fail on Low-Resource Languages?</title>
      <link>https://arxiv.org/abs/2608.30725</link>
      <guid isPermaLink="false">arxiv:2608.30725</guid>
      <pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W36</category>
      <description>Recent multilingual vision--language encoders cover hundreds of languages in a single model, yet on two state-of-the-art instances retrieval on low-resource languages (LRL; e.g. Swahili) trails high-resource ones (HRL; e.g. English) by $30^+$\,pp. We ask where in the trained encoder this gap is located. Prior modality-gap and cross-lingual subspace work suggests a linear language direction at the </description>
    </item>
    <item>
      <title>Harnessing Artificial Intelligence (AI) in African Historiography: Augmenting Historical Sources</title>
      <link>https://doi.org/10.36349/zamijoh.2026.v04i02.019</link>
      <guid isPermaLink="false">doi:10.36349/zamijoh.2026.v04i02.019</guid>
      <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>Artificial Intelligence (AI) has emerged as one of the most transformative technologies of the twenty-first century, rapidly reshaping academic fields and research methodologies. In African historiography, where oral traditions, written document, archaeological materials and historical linguistic form the core foundations of historical sources. Artificial intelligence offers unprecedented opportun</description>
    </item>
    <item>
      <title>Identifying morphological errors in the English of Deaf learners in Lesotho</title>
      <link>https://doi.org/10.4102/rw.v17i1.652</link>
      <guid isPermaLink="false">doi:10.4102/rw.v17i1.652</guid>
      <pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>Background: Morphological awareness (MA) is widely recognised as contributing to successful reading and reading comprehension. Despite this, research examining the nature of morphological errors among second-language (L2) English learners remains limited, particularly among Deaf learners.
Objectives: This study identified and classified English MA errors made by Deaf learners in Lesotho and determ</description>
    </item>
    <item>
      <title>A Machine-Learning Model for Phishing Detection in Swahili Messages: A Case of Tanzania</title>
      <link>https://doi.org/10.37284/eajit.9.2.5655</link>
      <guid isPermaLink="false">doi:10.37284/eajit.9.2.5655</guid>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>Phishing conducted in Swahili has become a persistent threat to the millions of Tanzanians who depend on mobile-money services, yet the detection tools in common use are built for English and transfer poorly to a language whose morphology, register, and transactional vocabulary differ sharply from it. This study makes three contributions. It establishes that classical machine learning, given featu</description>
    </item>
    <item>
      <title>The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues</title>
      <link>https://arxiv.org/abs/2608.28144</link>
      <guid isPermaLink="false">arxiv:2608.28144</guid>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>Social power plays a fundamental role in shaping human interaction, yet computational studies of power remain limited to narrow linguistic and cultural settings. Existing datasets further lack the demographic and relational depth needed for robust cross-cultural analysis. To address this gap, we introduce a theoretically grounded framework for studying social power in naturalistic multilingual dia</description>
    </item>
    <item>
      <title>The effect of orthography on learning to read in Xitsonga</title>
      <link>https://doi.org/10.2989/16073614.2026.2691523</link>
      <guid isPermaLink="false">doi:10.2989/16073614.2026.2691523</guid>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description></description>
    </item>
    <item>
      <title>Orthography and spelling challenges faced by Xitsonga home language learners in essay writing</title>
      <link>https://doi.org/10.2989/16073614.2026.2705950</link>
      <guid isPermaLink="false">doi:10.2989/16073614.2026.2705950</guid>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description></description>
    </item>
    <item>
      <title>KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum Training</title>
      <link>https://arxiv.org/abs/2608.26941</link>
      <guid isPermaLink="false">arxiv:2608.26941</guid>
      <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>We present KinyaEmbed, the first dedicated sentence embedding model for Kinyarwanda, a morphologically rich Bantu language spoken by over 12 million people in Rwanda. Existing multilingual embedding models such as LaBSE, mE5-large, and OpenAI text-embedding-3-large perform poorly on Kinyarwanda due to severe under-representation in their pre-training corpora. KinyaEmbed is built on KinyaBERT-large</description>
    </item>
    <item>
      <title>TabuLM: Morphology-Aware Tabular Pre-training for Low-Resource Languages</title>
      <link>https://arxiv.org/abs/2608.26923</link>
      <guid isPermaLink="false">arxiv:2608.26923</guid>
      <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>We present TabuLM, the first language model pre-trained on Kinyarwanda tabular data. Kinyarwanda is a morphologically rich Bantu language spoken by over 12 million people in Rwanda, yet lacks any dedicated tabular representation learning resource. TabuLM extends KinyaBERT-large, a two-tier morphological transformer, with additive row, column, and cell-type embeddings and a learned table-structure </description>
    </item>
    <item>
      <title>AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition</title>
      <link>https://arxiv.org/abs/2608.26434</link>
      <guid isPermaLink="false">arxiv:2608.26434</guid>
      <pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, released with switch-level English span tags, perutterance Code-Mixing Index (CMI), </description>
    </item>
    <item>
      <title>AraDetox: A Multi-Dialect Arabic Detoxification Dataset</title>
      <link>https://arxiv.org/abs/2608.22894</link>
      <guid isPermaLink="false">arxiv:2608.22894</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 detoxified rewrites generated using GPT-5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. The generated outp</description>
    </item>
    <item>
      <title>Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matching</title>
      <link>https://arxiv.org/abs/2608.23920</link>
      <guid isPermaLink="false">arxiv:2608.23920</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>We measure tablet-2, a production long-term memory engine for language models, on the text benchmarks the field already uses and on cross-lingual retrieval of photographs stored with no text at all. Its retrieval path contains no lexical matching, no keyword scoring, and no language model of its own. On LongMemEval-S (500 questions) it scores 95.7% [93.4, 97.1]; on BEAM-1M (700 questions, 2.21M st</description>
    </item>
    <item>
      <title>The Geometry of Low-Resource Language Representations</title>
      <link>https://arxiv.org/abs/2608.23358</link>
      <guid isPermaLink="false">arxiv:2608.23358</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W35</category>
      <description>The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities. In this paper, we characterise this gap through the lens of representational geometry. Comparing the geometric properties of hidden representations across 30 languages reveals that LLM geometry is systematically related to language data </description>
    </item>
    <item>
      <title>Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator</title>
      <link>https://arxiv.org/abs/2608.22432</link>
      <guid isPermaLink="false">arxiv:2608.22432</guid>
      <pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W34</category>
      <description>Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic, Chinese, Hindi, Japanese, Spanish, Turkish, and Swahili, and 7 of 15 backbone pairs show statistically significant pairwise rank reversal. We treat this as a measurement problem. The multi</description>
    </item>
    <item>
      <title>A Distilbert Case-Based Deep Learning Model for Part-of-Speech Tagging for Under-Resourced Kenyan Language: A Case of Dholuo</title>
      <link>https://doi.org/10.14445/22312803/ijctt-v74i7p105</link>
      <guid isPermaLink="false">doi:10.14445/22312803/ijctt-v74i7p105</guid>
      <pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W34</category>
      <description>One of the basic Natural Language Processing (NLP) tasks is Part-of-Speech (POS) tagging, which helps in various applications like sentiment analysis and information retrieval. However, creating accurate POS taggers for low-resource African languages continues to be difficult due to the scarcity of linguistic resources that are annotated. Its contribution is a deep learning method for POS tagging </description>
    </item>
    <item>
      <title>Assessing the quality and effectiveness of isiXhosa translations in public health communication materials to advance health equity in Nelson Mandela Bay</title>
      <link>https://doi.org/10.1186/s12982-026-02389-w</link>
      <guid isPermaLink="false">doi:10.1186/s12982-026-02389-w</guid>
      <pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W34</category>
      <description>Language is a critical determinant of health equity in South Africa’s multilingual healthcare system. Despite constitutional mandates to communicate in indigenous languages such as isiXhosa, many public health materials remain inadequately translated, compromising access to essential health information. This study assesses the quality and effectiveness of isiXhosa translations of two key public he</description>
    </item>
    <item>
      <title>TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation</title>
      <link>https://arxiv.org/abs/2608.18655</link>
      <guid isPermaLink="false">arxiv:2608.18655</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W34</category>
      <description>The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a digital divide that limits AI adoption on the continent. Recent open-source LLMs systematically underperform on African language machine translation, while the lack of large-scale, high-quality, open-source parallel data has constrained the development of competitive small language models (SLMs). We in</description>
    </item>
    <item>
      <title>Enhancing dialectal Arabic aspect based sentiment analysis through a novel Algerian dialect telecommunication dataset using a unified end-to-end approach</title>
      <link>https://doi.org/10.1007/s44443-026-00703-9</link>
      <guid isPermaLink="false">doi:10.1007/s44443-026-00703-9</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W34</category>
      <description>This study introduces a novel, publicly available balanced Algerian Arabic Aspect Based Sentiment Analysis (ABSA) dataset consisting of 11,338 comments written in the Algerian dialect. The dataset follows the semantic evaluation 2016 annotation guidelines and includes both explicit and implicit aspects within the telecommunications domain. Furthermore, the study proposes a unified End-to-End (E2E)</description>
    </item>
    <item>
      <title>Structure, Association, and Decision Value: Representation-Based Difficulty Estimation for Adaptive Inference in African-Language NLI</title>
      <link>https://arxiv.org/abs/2608.19003</link>
      <guid isPermaLink="false">arxiv:2608.19003</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W34</category>
      <description>We ask whether internal representation statistics can provide useful example-level difficulty signals for adaptive inference in multilingual African NLP, and find that they cannot in this setting. Studying natural language inference across 15 African languages with frozen off-the-shelf checkpoints, we report four results. First, AfriXNLI&apos;s English configuration shares 1,047 of its 1,050 examples v</description>
    </item>
    <item>
      <title>A Morpho-Syntactic Study of Forms and Functions of Interrogative Markers in the Adamawa Fulfulde</title>
      <link>https://doi.org/10.66490/jajolls.2026.v10.i2.010</link>
      <guid isPermaLink="false">doi:10.66490/jajolls.2026.v10.i2.010</guid>
      <pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W34</category>
      <description>This study is on the forms and functions of interrogative markers in Fulfulde. It aims at providing a clear view of interrogative markers and their functions in Fulfulde sentences. Specifically, to determine: the forms, functions and the syntactic constraints of interrogative markers in Fulfulde sentences. Observation and oral interview is used in data collection. The data are presented and analys</description>
    </item>
    <item>
      <title>An AI-Based Adaptive Learning Platform for Multilingual and Low-Resource Educational Contexts: A Case Study on Nigeria</title>
      <link>https://arxiv.org/abs/2608.15738</link>
      <guid isPermaLink="false">arxiv:2608.15738</guid>
      <pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>Educational platforms in under-resourced and multilingual contexts, such as Nigeria, often struggle with limited personalisation, inadequate language support, and weak curriculum internationalisation, leading to reduced learner engagement and inclusivity. This paper presents an AI-based adaptive learning platform designed for multilingual and low-resource educational contexts, with a case study on</description>
    </item>
    <item>
      <title>Use of artificial intelligence for medication adherence assessment in patients with type 2 diabetes: an exploratory study in Morocco using ChatGPT and the validated general medication adherence scale.</title>
      <link>https://doi.org/10.24171/j.phrp.2025.0396</link>
      <guid isPermaLink="false">doi:10.24171/j.phrp.2025.0396</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>Objectives
Medication adherence remains a major challenge in the management of type 2 diabetes (T2D), especially in middle-income countries such as Morocco. With the rapid development of artificial intelligence, large language models, including ChatGPT, may offer new opportunities for clinical research through simulated patient profiles. This study examined the feasibility of using ChatGPT to gene</description>
    </item>
    <item>
      <title>Language-Specific Gaps in AI Safety Training Datasets</title>
      <link>https://arxiv.org/abs/2608.13695</link>
      <guid isPermaLink="false">arxiv:2608.13695</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage claims frequently do not survive inspection at the level of an individual language. Auditing 21 resources across 25 language slices, of which 20 count as datasets under our</description>
    </item>
    <item>
      <title>Digital Genetic Diagnosis of Malaria Using Explainable Deep Learning</title>
      <link>https://doi.org/10.54361/ajmas.269827</link>
      <guid isPermaLink="false">doi:10.54361/ajmas.269827</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>Malaria, caused by Plasmodium spp., remains a leading cause of mortality in sub-Saharan Africa, with an estimated 282 million cases and 610,000 deaths in 2024. The var gene family primarily drives the virulence of P. falciparum, encoding the polymorphic adhesin PfEMP1, which mediates cytoadherence and immune evasion while inducing a quantifiable morphological footprint on host erythrocytes, includ</description>
    </item>
    <item>
      <title>A Hybrid Spatio-Temporal Transformer-LSTM Network for Explainable Crop Yield Prediction Using Multi-Source Remote Sensing in Arid African Regions</title>
      <link>https://doi.org/10.62411/faith.3048-3719-408</link>
      <guid isPermaLink="false">doi:10.62411/faith.3048-3719-408</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>The escalating food security crisis in Sub-Saharan Africa, driven by population growth and climate volatility, necessitates robust agricultural monitoring and prediction systems. However, existing approaches often rely on historical trends and struggle to capture the nonlinear effects of erratic rainfall and thermal shocks, particularly in arid regions. Furthermore, limited geographic representati</description>
    </item>
    <item>
      <title>The Nature and Function of Code-Switching Amongst Students of Higher Institutions in Adamawa State, Nigeria</title>
      <link>https://doi.org/10.56201/ijelcs.vol.11.no5.2026.pg61.69</link>
      <guid isPermaLink="false">doi:10.56201/ijelcs.vol.11.no5.2026.pg61.69</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>This study examines the nature and function of code-switching among students in higher 
institutions of learning in Adamawa State, Nigeria, a linguistically diverse region where English, 
Hausa, Nigerian Pidgin, and indigenous languages interact in daily communication. Guided by 
two research questions concerning the predominant structural patterns of code-switching and the 
primary communicative </description>
    </item>
    <item>
      <title>Hierarchical Curriculum Transfer for Swahili QA with mT5: SQuAD Pretraining and KenSwQuAD Extractive-to-Abstractive Refinement</title>
      <link>https://doi.org/10.59543/14d9kq02</link>
      <guid isPermaLink="false">doi:10.59543/14d9kq02</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>The introduction of the Kencorpus Swahili Question Answering Dataset (KenSwQuAD) presents a compelling opportunity to advance Machine Reading Comprehension (MRC) for the Swahili language. However, the mixed composition of this dataset—comprising 67.6\% extractive and 32.4\% abstractive answers—introduces substantial hurdles for standard training pipelines. In this study, we benchmark the performan</description>
    </item>
    <item>
      <title>Humanising artificial intelligence for enhancing engagement and retention in open and distance learning in sub-Saharan Africa</title>
      <link>https://doi.org/10.17159/2520-9868/i105a03</link>
      <guid isPermaLink="false">doi:10.17159/2520-9868/i105a03</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>The integration of artificial intelligence (AI) into open and distance learning (ODeL) is reshaping engagement and retention in sub-Saharan Africa and offering new opportunities while also presenting significant constraints. In this conceptual paper, we examine how humanising AI can be embedded into ODeL design to strengthen dialogue, agency, care, and belonging. The study was guided by Freire&apos;s n</description>
    </item>
    <item>
      <title>Phonemic Contrasts in Hausa: A Minimal Pair Analysis Across Word Positions</title>
      <link>https://doi.org/10.56201/ijelcs.vol.11.no5.2026.pg102.111</link>
      <guid isPermaLink="false">doi:10.56201/ijelcs.vol.11.no5.2026.pg102.111</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>This paper investigates phonemic contrasts in Hausa through the analysis of minimal pairs 
occurring in word-initial, medial, and final positions. The study employs a descriptive qualitative 
research design. Data were obtained from Hausa lexical items using secondary linguistic sources 
and supplemented by native-speaker intuition to ensure authenticity. The analysis is based on the 
minimal pair</description>
    </item>
    <item>
      <title>Matsayin Harshen Kamuku a Ma’aunin Ra’in Fisherman</title>
      <link>https://doi.org/10.36348/sijll.2026.v09i08.003</link>
      <guid isPermaLink="false">doi:10.36348/sijll.2026.v09i08.003</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>This research work focused on language endagerment and revitalissation. The aim of this research is study the linguistic influence of Hausa on Kamuku, a Kainji language spoken in parts of Niger and Kaduna States, Nigeria. Due to centuries of trade, intermarriage, and cultural contact, Hausa has functioned as a regional lingua franca, leading to significant lexical borrowing, phonological adaptatio</description>
    </item>
    <item>
      <title>Correlations of auditory discrimination, phonemic awareness, and literacy: Evidence from Grade 4 learners in a rural Gauteng primary school</title>
      <link>https://doi.org/10.17159/2520-9868/i105a08</link>
      <guid isPermaLink="false">doi:10.17159/2520-9868/i105a08</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>In South Africa, many Grade 4 learners transition from first-language instruction to English as the language of learning and teaching (LoLT). This often occurs without sufficient support for the development and transfer of foundational phonological and literacy skills. Additionally, there is limited research on the role of auditory discrimination in this process. In this study we examine the relat</description>
    </item>
    <item>
      <title>Artificial intelligence in microbial biotechnology for food security: current state and challenges</title>
      <link>https://doi.org/10.3389/fbrio.2026.1921669</link>
      <guid isPermaLink="false">doi:10.3389/fbrio.2026.1921669</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>The global food system is constantly being constrained by biotic and abiotic challenges, resulting in instability and insecurity, particularly in regions like sub-Saharan Africa, where agricultural productivity often remains below global averages. Microbial biotechnology includes many sustainable ways of leveraging the metabolic potential of microorganisms, such as bacteria, fungi, and viruses, to</description>
    </item>
    <item>
      <title>The Illusion of Cross-Lingual Safety in Low-Resource Languages</title>
      <link>https://arxiv.org/abs/2608.11146</link>
      <guid isPermaLink="false">arxiv:2608.11146</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pair</description>
    </item>
    <item>
      <title>Reduplication Process of Hausa Spoken in Yola and Jimeta Adamawa State Nigeria</title>
      <link>https://doi.org/10.66490/jajolls.2026.v10.i2.004</link>
      <guid isPermaLink="false">doi:10.66490/jajolls.2026.v10.i2.004</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>This research paper attempts a discussion on reduplication process of Hausa as spoken in Yola and Jimeta in Adamawa State. It documents, the Hausa speech forms as spoken by Hausa speakers in the respected areas of study. These papers also analyse Reduplication process as one of the word formation processes where part of the base or root is fully or partially repeated to create a new meaning or gra</description>
    </item>
    <item>
      <title>DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition</title>
      <link>https://arxiv.org/abs/2608.11441</link>
      <guid isPermaLink="false">arxiv:2608.11441</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-</description>
    </item>
    <item>
      <title>Sociolinguistic Analysis of the Role of Silence in Communication Among Speakers of Tigun-Njuande Dialect of Taraba State</title>
      <link>https://doi.org/10.66490/jajolls.2026.v10.i2.007</link>
      <guid isPermaLink="false">doi:10.66490/jajolls.2026.v10.i2.007</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>The study examines silence in Tigun-Njuande language and communication in Taraba State. It aims to explain the sociolinguistic functions of silence in Tigun-Njuande dialect in communication. Despite its significance to communication within the communities’ silence has received limited attention by researchers as non-natives often misunderstand it. It also shows how silence functions as a meaningfu</description>
    </item>
    <item>
      <title>Naija-Petro AI: Adapting Large Language Models (LLMs) for Oil and Gas Applications Through Domain-Specific Fine-Tuning</title>
      <link>https://doi.org/10.2118/234918-ms</link>
      <guid isPermaLink="false">doi:10.2118/234918-ms</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>
 The global oil and gas industry has been increasingly turning to artificial intelligence to solve knowledge management challenges and domain-specific large language models (LLMs) catering to the petroleum engineering industry are still scarce. This issue is especially pronounced in Nigeria – the largest producer of crude oil in Africa – where the brain drain issue and data sovereignty concerns, </description>
    </item>
    <item>
      <title>SafetyBuddy: A Multimodal LLM-Based Safety Intelligence Platform for Regulatory Compliance in Process Industries</title>
      <link>https://doi.org/10.2118/234917-ms</link>
      <guid isPermaLink="false">doi:10.2118/234917-ms</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>
 Process industries are a major source of occupational injuries and fatalities worldwide due to safety violations. Nigeria alone has recorded more than 412 fatalities in the oil and gas industry and the construction industry has a documented history of accidents associated with personal protective equipment (PPE) non-compliance. According to the World Risk Poll (2024), Sub-Saharan Africa is the h</description>
    </item>
    <item>
      <title>Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities</title>
      <link>https://arxiv.org/abs/2608.09046</link>
      <guid isPermaLink="false">arxiv:2608.09046</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W33</category>
      <description>Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a</description>
    </item>
    <item>
      <title>Dialect Gloss</title>
      <link>https://doi.org/10.34778/a48xn206</link>
      <guid isPermaLink="false">doi:10.34778/a48xn206</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W32</category>
      <description>Brief Description
This variable proposes LLM-Assisted Dialect Gloss Detection (see also Bommarito &amp; Michael, 2025; Liang et al., 2026) and refers to the automated identification, contextual interpretation, and measurement of dialectal, vernacular, or otherwise non-standard linguistic forms in digital texts (Joshi et al., 2025). The procedure combines dictionary-based detection with large language </description>
    </item>
    <item>
      <title>Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation</title>
      <link>https://arxiv.org/abs/2608.07629</link>
      <guid isPermaLink="false">arxiv:2608.07629</guid>
      <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W32</category>
      <description>Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon. When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection. We implemented an embedding initialization strategy where a language to</description>
    </item>
    <item>
      <title>Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili</title>
      <link>https://arxiv.org/abs/2608.03532</link>
      <guid isPermaLink="false">arxiv:2608.03532</guid>
      <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W32</category>
      <description>Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by submitting 4,900 symmetric English--Swahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes, yielding 19,600 completions evaluated for stereotype p</description>
    </item>
    <item>
      <title>Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation</title>
      <link>https://arxiv.org/abs/2608.02555</link>
      <guid isPermaLink="false">arxiv:2608.02555</guid>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W32</category>
      <description>Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most NLP research on Arabic texts treats it as a temporary phenomenon resulting from limited technological support for the Arabic script. In this work, we engage with Arabic speakers to collect insights on their perceptions an</description>
    </item>
    <item>
      <title>University Students&apos; Perceptions and Use of Artificial Intelligence Tools for Academic Activities and Critical Thinking: A Thematic Literature Review with Reflections from Tanzanian Higher Education</title>
      <link>https://doi.org/10.37284/eajes.9.3.5472</link>
      <guid isPermaLink="false">doi:10.37284/eajes.9.3.5472</guid>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W32</category>
      <description>Artificial intelligence (AI) tools are rapidly transforming academic practices in higher education, especially through generative platforms and AI-supported writing assistants. University students increasingly use these tools for academic writing, idea generation, summarisation, proofreading, translation, research support, presentation preparation, and clarification of complex concepts. This thema</description>
    </item>
    <item>
      <title>English Influence on the Structure of Igbo Personal Names</title>
      <link>https://doi.org/10.58578/edumalsys.v4i3.11635</link>
      <guid isPermaLink="false">doi:10.58578/edumalsys.v4i3.11635</guid>
      <pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>English–Igbo language contact increasingly shapes contemporary personal naming practices in Igboland, where English functions as both a second and an official language. This study examines the phonological and morphological structures of English-influenced Igbo personal names and investigates how these transformations reflect patterns of language contact and changing naming conventions among Igbo </description>
    </item>
    <item>
      <title>Low-Resource Machine Translation of Yoruba Text to Nigerian Pidgin Using NLLB-200 and Parameter-Efficient Fine-Tuning</title>
      <link>https://doi.org/10.7759/s44389-026-00202-y</link>
      <guid isPermaLink="false">doi:10.7759/s44389-026-00202-y</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>Machine translation for low-resource and structurally divergent languages remains a significant challenge, particularly when mapping highly tonal languages to contact languages with fluid orthographies. This study presents the development of a translation system for Yoruba to Nigerian Pidgin, which addresses a critical gap in African natural language processing. Using a custom-curated parallel cor</description>
    </item>
    <item>
      <title>AryWiki-Instruct: A high-fidelity instruction tuning dataset for Moroccan Arabic (Darija)</title>
      <link>https://doi.org/10.1016/j.dib.2026.113140</link>
      <guid isPermaLink="false">doi:10.1016/j.dib.2026.113140</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>This article presents AryWiki-Instruct, a high-fidelity instruction tuning dataset for Moroccan Arabic (Darija), comprising 46,590 Question and Answer (QA) pairs. The dataset was derived from a snapshot of the Moroccan Arabic Wikipedia (arywiki) and generated using the Gemini-2.5-Flash model via a Context Aware batch processing architecture. The data creation process involved parsing raw XML Wikip</description>
    </item>
    <item>
      <title>MORAD: A Multimodal Dataset of Authentic Emotional Expressions in Moroccan Arabic</title>
      <link>https://doi.org/10.1109/TCSS.2026.3688593</link>
      <guid isPermaLink="false">doi:10.1109/tcss.2026.3688593</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>Emotion recognition plays an important role in enhancing social communications and is a key component of improving human-machine interactions. With the rise of deep learning, emotion recognition systems are becoming more effective and accurate. However, progress is often limited by the availability of datasets, a problem affecting low-resourced languages such as the Moroccan Arabic Dialect (Darija</description>
    </item>
    <item>
      <title>Zero-Shot Cross-Lingual Transfer for Hate Speech Detection Using Multilingual Models</title>
      <link>https://doi.org/10.1109/eSmarTA70636.2026.11652104</link>
      <guid isPermaLink="false">doi:10.1109/esmarta70636.2026.11652104</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>Detecting hate speech in low-resource and unseen languages remains challenging due to limited labeled data and linguistic diversity. This paper presents a comparative study of zero-shot cross-lingual transfer for hate speech detection using two multilingual transformer models: mDeBERTa-v3 and XLM-RoBERTa. To the best of our knowledge, mDeBERTa-v3 has not been previously used by researchers for zer</description>
    </item>
    <item>
      <title>Post-stroke aphasia assessment in Arabic-speaking populations: A critical review of validated tools, structural barriers, and the diglossic validity problem.</title>
      <link>https://doi.org/10.1016/j.cortex.2026.08.005</link>
      <guid isPermaLink="false">doi:10.1016/j.cortex.2026.08.005</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>Arabic is spoken by more than 400 million people across a region experiencing a substantial increase in stroke burden. Yet, no comprehensively validated aphasia assessment battery exists that provides dialect-stratified norms across major Arabic-speaking communities. This critical narrative review evaluates Arabic aphasia assessment against explicit psychometric and linguistic criteria, identifyin</description>
    </item>
    <item>
      <title>Length and Spectral Dynamics of Rising Diphthongs in Nuer</title>
      <link>https://doi.org/10.1121/10.0045250</link>
      <guid isPermaLink="false">doi:10.1121/10.0045250</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>Nuer, a West Nilotic language spoken in South Sudan and Ethiopia, is typologically unusual in that it exhibits intraparadigmatic alternations across a three-way division of vowel length categories: short /V/, long /VV/, and overlong /VVV/. Nuer has a series of rising-sonority diphthongs, e.g., /wa/, /we/, /ja/, that participate in this three-way length distinction (e.g., /lwaak/ ∼ /lwaaak/, cow by</description>
    </item>
    <item>
      <title>Explainable CNN-BiLSTM Framework for Okra Leaf Disease Detection and Classification Using Grad-CAM and SHAP Visualization</title>
      <link>https://doi.org/10.7759/s44389-026-00237-1</link>
      <guid isPermaLink="false">doi:10.7759/s44389-026-00237-1</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>Okra (Abelmoschus esculentus) is an economically important vegetable crop in Nigeria and across sub-Saharan Africa, yet it remains underrepresented in agricultural artificial intelligence research, and its leaf diseases are difficult to distinguish by eye. This study presents an explainable CNN-BiLSTM (ECBL) framework for automated okra leaf disease classification across six categories: Alternaria</description>
    </item>
    <item>
      <title>Rethinking and formalising the state across languages: a unified computational learning theory account</title>
      <link>https://arxiv.org/abs/2608.00523</link>
      <guid isPermaLink="false">arxiv:2608.00523</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-specific morphosyntactic phenomenon. This article argues instead that the state is a systemic, context-dependent morphosyntactic mechanism that selects grammatical templates across synthetic languages. Within the Template-Based Modular Cognitive fr</description>
    </item>
    <item>
      <title>Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models</title>
      <link>https://doi.org/10.22214/ijraset.2026.84356</link>
      <guid isPermaLink="false">doi:10.22214/ijraset.2026.84356</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>The use of social media in East Africa has grown rapidly, and with it, the spread of hate speech has become a serious
concern. This problem is even more complex in online spaces where people often switch between English and Swahili within the
same sentence or conversation. Such code-switching makes it difficult for existing systems to accurately detect harmful content,
especially because there is </description>
    </item>
    <item>
      <title>An AI-driven bilingual clinical language processing framework for English-Yoruba medical text and voice translation and understanding</title>
      <link>https://doi.org/10.65752/14j44s91</link>
      <guid isPermaLink="false">doi:10.65752/14j44s91</guid>
      <pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>The research came up with a bilingual natural language processing (NLP) model of a patient-focused healthcare system that is geared towards improving healthcare communication in low-resource settings, based on both the English and Yoruba languages. The system combines multilingual transformer-based text processing, speech-to-text and text-to-speech modules, retrieval-augmented response generation,</description>
    </item>
    <item>
      <title>Code-mixing between Egyptian Arabic and English in Food Review Vlogs on Instagram</title>
      <link>https://doi.org/10.21608/jssa.2026.477846.1831</link>
      <guid isPermaLink="false">doi:10.21608/jssa.2026.477846.1831</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description></description>
    </item>
    <item>
      <title>A Yoruba Language Automatic Speech Recognition System using Deep Learning Approach</title>
      <link>https://doi.org/10.5120/ijca0dcb3a1b147c</link>
      <guid isPermaLink="false">doi:10.5120/ijca0dcb3a1b147c</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description></description>
    </item>
    <item>
      <title>Building ‌a ‌Transformer-Based ‌Neural Machine Translation System for English–Kibajuni Translation: A Low-Resource Deep Learning Approach for Indigenous Language Preservation</title>
      <link>https://doi.org/10.61250/ssmj/v1.i4.2</link>
      <guid isPermaLink="false">doi:10.61250/ssmj/v1.i4.2</guid>
      <pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>Recent progress in artificial intelligence has pushed machine translation to high levels of accuracy for widely resourced languages. Yet for many indigenous and endangered languages, comparable tools remain absent, largely because digitized linguistic data are scarce. Kibajuni, a minimally documented Bantu language spoken along the Kenyan coast, illustrates this gap. Publicly accessible English–Ki</description>
    </item>
    <item>
      <title>Precision Agriculture: A Comprehensive Review of Technologies, Architectures, Deep-Learning Pipelines, and Future Prospects</title>
      <link>https://doi.org/10.70917/ijcisim-2026-2174</link>
      <guid isPermaLink="false">doi:10.70917/ijcisim-2026-2174</guid>
      <pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate>
      <category>2026-W31</category>
      <description>Artificial Intelligence is shaking up the way we grow food. Large farms already use AI-powered tools, such as deep learning, remote sensors, and IoT networks, to improve yield prediction, early disease detection, and more judicious resource use . Large-scale operations, about 60% of them, have improved their crop yields by 15–20%, while input costs have been cut down by a quarter. Smaller and mid-</description>
    </item>
  </channel>
</rss>
