Panel 5 Wed 23 · 09:00 EN
Friederike Lüpke, Jules Mansaly
University of Helsinki · University of the Gambia
The movement to include African languages in AI systems has concentrated on the continent’s largest languages, while the vast majority of Africa’s 2,000+ languages remain extremely low-resourced. For speakers of small, locally confined languages in ‘francophone’ multilingual West Africa, this creates a triple marginalisation: they are excluded first by the global dominance of English in AI, with user experiences in French being already degraded, and further by the emergent hierarchy among African languages themselves. Wolof and Pulaar are gaining AI visibility, but these are already languages of wider communication for rural multilinguals whose multiple primary languages — Baïnounk Gubëeher, Joola Kujireray, Balant Ganja, etc. — will foreseeably never have sufficient digital data to train dedicated language models. Yet these same speakers, routinely using five or more languages in their daily lives, possess multilingual competence potentially able to overcome monolingual AI architectures. LILIEMA (Language-Independent Literacies for Inclusive Education in Multilingual Areas), a Senegalese association and literacy programme developed collaboratively since 2017, offers a promising site for exploring how existing AI tools can serve users at the margins of the digital language divide. Rather than selecting a single language, LILIEMA teaches letter-sound associations using the official alphabet for Senegalese languages across learners’ entire linguistic repertoires. Through this, language-independent AI tools do not need to “know” any specific small language to be useful; they can operate at the level of orthographic pattern, phonological structure, and pedagogical template creation. We report on LILIEMA workshops conducted in May 2026 piloting the use of generative AI to create adaptive multilingual learning materials. We explore several concrete applications: (1) transliteration between widespread informal French-based lead-language writing practices and the official orthographies of Senegalese languages, enabling LILIEMA to bridge everyday written practice with standardised literacy; (2) generation of phonologically targeted word lists containing specific letters and letter combinations matching the sequential introduction order in LILIEMA’s beginner curriculum, drawn from any language in the learner’s repertoire; (3) creation of customisable exercise templates that teachers can adapt to specific language ecologies without requiring technical expertise; (4) production of varied exercise types at calibrated difficulty levels; (5) adaption of a core set of learning materials to each village’s specific multilingual profile and thus reducing the manual labour of customisation; (6) prompt engineering as a transferable literacy skill in its own right. Formulating effective prompts for AI tools through collective oral multilingual practice is itself a form of digital literacy and positions users as active agents in the AI ecosystem. While we present on the development of sustainable AI infrastructures, we also address axes 1 and 3. We conduct ethnographic research on LILIEMA members and our own digital and AI practices transforming methodological frameworks, and through this advance epistemological considerations centring African language users, who are not mono- or bilingual but versatile multilinguals with longstanding experiences of sustaining oral multilingual practices alongside writing practices centred on focal languages of literacy. We conclude with recommendations for adapting AI tools to similar conditions of extreme digital language scarcity in thriving offline multilingual environments.
Panel 2 Mon 21 · 16:00 EN
Khaoula Stiti
Unaffiliated scholar
When researchers in African Studies reach for AI tools to document, classify, or analyse heritage, they inherit a problem that predates AI itself. This paper argues that the failures now emerging in AI-assisted heritage research are not new, they are the continuation of a pattern already visible in earlier digital heritage platforms, and that understanding this pattern is the first step toward designing against it. The argument draws on a concrete case: the P@trimonia 2.0 project (Belgium–Tunisia, 2019–2024), a North South collaboration that attempted to build a heritage participatory platform. The project failed in instructive ways. Technologically, it was built on a platform designed for a Belgian open-air museum and then assumed to transfer to a non-heritagised urban landscape in Tunis, a one-size-fits-all logic that collapsed when confronted with a radically different heritage context. Culturally, it reproduced three overlapping problems: cultural bias, where Eurocentric design frameworks displaced local Tunisian knowledge and community input; power imbalance, where funding control remained with Belgian partners and development occurred physically distant from the Tunisian context; and technological solutionism, where the digital tool was treated as a universal fix that bypassed the social and historical complexity of colonial heritage. These are not merely platform-design failures. They are methodological failures, and they map directly onto what AI systems now do at scale. AI models trained on already digitised, Western catalogued heritage data reproduce the same erasures as P@trimonia 2.0, only faster and with greater institutional authority. The infrastructural assumptions embedded in AI discourse mirror the power imbalances of North South research partnerships. And the apparent objectivity of AI classification systems is a more powerful form of the same technological solutionism that reduced the colonial heritage of Tunis to a set of material data points. From this diagnosis, the paper derives a set of practical methodological criteria for African Studies researchers working with AI tools. These include: evaluating whether training data reflects the specific heritage context under study or merely the nearest digitised approximation; auditing infrastructural dependencies before adopting AI-assisted workflows; designing community co-creation into the process from the outset rather than retrofitting it; and treating local classification systems and knowledge structures as inputs to tool design, not as content to be processed by tools designed elsewhere. The paper concludes that decolonial computing is not a theoretical framework to be applied after methodology is settled, it is the methodology, and AI cannot be responsibly used in African heritage research without it.
Panel 3 Tue 22 · 09:00 EN
Jiayu Yang, Durgesh Nandini
University of Bayreuth
Research databases in African Studies accumulate rich metadata, including subject keywords, themes, geographic and cultural references, but much of this data remains isolated, difficult to retrieve and reuse across databases. The key problem is semantic grounding: without stable identifiers, two databases describing the same cultural practice or geographic region have no way to know it, even when the underlying subjects overlap. Our work draws on 3,975 metadata records from five partner institutions across Africa, Europe, and South America, spanning disciplines from musicology to social anthropology. Our paper addresses this gap with CAREL (Context-Aware Routing for Entity Linking), a training-free, cost-aware pipeline that links research metadata keywords to Wikidata QIDs, the shared identifiers that make cross-database search and, in the longer term, knowledge graph construction possible. An LLM extraction stage first identifies concept-level entities from titles, abstracts, and keyword fields. Linking then runs through a four-layer cascade. The first three layers are rule-based, built around Cross-lingual Retrieval Consensus (CRC), a novel signal that reads the agreement of retrieval ranks across English, French, Portuguese, and German queries as a measure of linking confidence. Only genuinely ambiguous keywords reach the final layer, where a locally deployed open-source LLM reasons over the record context using live Wikidata tools. Local deployment is a deliberate choice: research metadata never leaves the institution, and the communities behind the data keep control over how it is used. The pipeline is built for the realities of African research metadata: variant spellings, transliterated terms, and culturally embedded concepts underrepresented in standard ontologies. On two manually verified benchmarks constructed from our corpus, authority-anchored subject headings and free researcher tags, CAREL reaches between 88.8% and 90.9% linking accuracy with two open-source models, against 53.6% for Wikidata's native ranking and 68.0% for OpenRefine reconciliation, the established cultural heritage workflow. Equally important, the pipeline recognises when a concept has no Wikidata entity at all, detecting these cases at an F1 of 88 rather than forcing a wrong link, and low-confidence links are flagged for human review rather than silently propagated. The evaluation also surfaces a structural asymmetry. The research materials are often in African languages; their descriptive metadata is predominantly English; and the label spaces used for grounding are European. The pipeline inherits this asymmetry rather than corrects it. Grounding in language-neutral QIDs is a partial counterweight, since a linked record becomes reachable through every language label its entities carry, and extending the search languages to Swahili and Yoruba is a concrete next step. Beyond the pipeline, we contribute two resources to the community: the benchmarks themselves, in a domain where evaluation sets are nearly absent, and a pathway from NIL detection to the creation of missing Wikidata entries, so that culturally specific African concepts enter the global knowledge infrastructure rather than remaining outside it.
Panel 6 Wed 23 · 11:00 EN
Lauren Coetzee
University of Luxembourg
African history before colonialism is not absent from the archive — it is dispersed across thousands of pages of travel writing, missionary records, and merchant accounts. Yet, African-authored sources and oral histories still remain significantly harder to locate, access, and digitise for this period — a disparity that itself reflects the archival legacies of colonial knowledge production. Across this uneven landscape, the interpretive labour required to transform narrative prose into structured, analysable evidence has remained a barrier to systematic historical inquiry. The problem is not scarcity but scale, and the absence of computational methods capable of handling the cultural and historical specificity that African sources demand. Recovering dynamic, diachronic economic practices from historical written accounts is therefore not only a historiographical problem but a methodological one: the source base needed to challenge these narratives is too large for traditional close reading, yet too historically and culturally specific for off-the-shelf natural language processing pipelines trained predominantly on contemporary, non-African data. This paper presents a methodology for doing exactly that, and questions what becomes possible for African history and digital humanities research when it is applied. Drawing on research spanning pre-colonial African trade networks, commodity currencies, and the digital analysis of European travelogues, this paper presents an LLM-assisted workflow for extracting structured economic data from the Time Traveller corpus — a large corpus of European travelogues compiled from accounts produced before 1900, capturing observations of African societies, economies, and landscapes at the moment of encounter. The methodological centrepiece is a historical codebook: a set of variable definitions, decision rules, and annotated examples developed from these sources, designed to calibrate LLM annotation to the evidentiary logic of a specific archive rather than to generalised text-processing categories. Applied to the African bead trade — the circulation of glass, shell, and metal beads as commodity currencies, status markers, and exchange media across continental networks — this workflow has produced over 27,000 coded observations capturing bead type, exchange context, geographic location, and trading partners, enabling spatial and temporal analysis of a market that conventional scholarship has left largely underresearched. Historical codebooks reframe what LLMs are asked to do: they become instruments for recovering specific categories of evidence that historians already know to look for but cannot extract at scale, moving beyond using LLMs as blunt instruments toward deploying them as precision tools calibrated to a specific historical problem. For African history, this matters enormously — it opens access to a vast body of primary source material that has shaped how the continent’s past has been narrated, and creates the conditions for scholars to interrogate and access those narratives from within the source archive itself. The paper reflects on where the method works well, where human validation and domain knowledge remain essential, and what the workflow makes possible for researchers working across African corpora and languages.
Panel 6 Wed 23 · 11:00 EN
Oreen Yousuf
Uppsala University
_Ajami_ refers to African languages written in the Arabic script. The use of Ajami began in the 10th century and spread to various sociopolitical states across Africa. Both existing handwritten and optical character recognition systems, as well as text-based NLP models perform poorly on recognition of Ajami manuscripts and analyzing digital Ajami text. This is mainly due to a lack of research and inclusion of Ajami manuscripts in NLP and Digital Humanities as a whole. I will present my work on building HTR/OCR and NLP infrastructure for Ajami manuscripts by providing historical and linguistic background of Ajami, data curation, and tasks currently being worked on.
Panel 4 Tue 22 · 11:00 EN
Rachel Maina
University of Wisconsin–Madison
AI-driven language learning platforms are expanding rapidly, yet their underlying pedagogical and technological architectures remain poorly aligned with the linguistic, infrastructural, and epistemological realities of African language contexts. Despite increased adoption, African languages remain underrepresented in both training data and pedagogical design. Existing computer-assisted language learning (CALL) tools such as Duolingo, Memrise, and earlier platforms like _Kiswahili kwa Kompyuta_ (KIKO) continue to rely predominantly on text-driven and decontextualized interaction, with limited support for sustained oral production and multimodal engagement. While many platforms incorporate audio and speech-based features, they typically prioritize standardized language forms and are optimized for technological conditions that do not consistently align with many African learning environments. Methodologically, my paper adopts a design-based research (DBR) approach, positioning the study as an initial design phase that combines comparative platform analysis with multimodal learning theory to generate context-sensitive design principles for African CALL. The paper represents the design articulation stage of DBR, focusing on problem framing and principle generation to guide future implementation and iterative testing. My analysis identifies three persistent limitations: the marginalization of dialectal variation, the constrained integration of interactive oral practice, and the isolation of language from its social and performative contexts. In response, I advance a design framework structured around four principles: audio-first and tone-aware interaction; integration of narrative and oral traditions; support for dialectal variation; and adaptability to low-resource environments. By shifting from critique to design, my paper offers a scalable model for developing CALL systems grounded in African language practices and knowledge systems.
Panel 7 Thu 24 · 11:00 EN
John Oluwafemi Daniel
University of Ibadan
Digital Humanities (DH) and Artificial Intelligence (AI) have globally ushered in new possibilities for data visualisation, archival preservation, and public engagement with cultural heritage. Yet, in the Nigerian context, these epistemological advancement are crippled by inequities in accessibility, infrastructural deficits, and a lack of genuine co-production between DH practitioners and the communities of study. Initiatives such as the Centre for Digital Humanities, University of Lagos, and Archivi.ng have made notable strides, however, their valorisation activities are heavily concentrated in urban centres, with indigenous communities with these histories disconnected from both digital resources and DH awareness. This paper critically interrogates the “town question” in Nigerian DH: how can DH research outputs be made accessible to indigenous communities who are often unaware of these materials, lack digital literacy, skills, and require information in indigenous languages? Historically, non-professional historians such as Isaac Babalola Akinyele, Chief Samuel Ojo Bada, and Nathaniel Oyerinde were deeply invested in ensuring that African history and cultures were represented authentically through the writing of history in indigenous language; a concern that contemporary African DH has yet to fully address. The uncritical transfer of European DH models to Africa has resulted in a wide gap. This paper is, therefore, grounded in reflexive, problem-oriented research stemming from an ongoing doctoral thesis on “Digital archives, public access and Ekiti Division, Ondo Province, 1914–1960.” The study uses qualitative methodology and draws on both primary sources (such as archival research, and interviews with community members and secondary sources such as books and journal articles. Ekiti Division is chosen for its unique position in Yoruba history particularly the Ekitiparapo War which ended through British intervention in 1886. Adopting a critical South–South perspective aligned with calls for sustainable African research infrastructures, this paper analyses practical obstacles and the paradoxical situation where communities have access to new media but not to formal digital archives. The paper argues that without moving DH infrastructures closer to those whose histories are being digitised, DH in Nigeria will neither be sustainable nor equitable. The study proposes, among others, the creation of historically verifiable digital platforms, such as dedicated websites or new media spaces for sharing historical data and fostering engagement.
Panel 2 Mon 21 · 16:00 EN
Iginio Gagliardone, Christine Mataranyika, Max Milella, Karabo Mohapeloa
University of the Witwatersrand
Calls to decolonise knowledge have transformed the discursive landscape of African higher education over the past decade. Movements such as #RhodesMustFall and #FeesMustFall demanded structural change to institutional cultures and curricula, and advocated for a deeper epistemological challenge, insisting that the production and adoption of knowledge in African universities be refigured around African intellectual traditions, researchers, and institutions. Yet the gap between discursive commitment and measurable practice remains poorly understood. How can one assess, empirically and at scale, whether these demands have altered the actual conduct of research — the granular, everyday decisions about which scholars to cite, which frameworks to invoke, and whose intellectual labour to acknowledge? This paper presents the methodological foundations and early findings of a collaborative project — developed under the auspices of the SA-UK Chair in the Digital Humanities at the University of the Witwatersrand and the Helsinki Institute for Social Sciences and Humanities — designed to answer this question. The project uses Electronic Theses and Dissertations (ETDs) as its primary corpus: a systematically archived, metadata-rich, and computationally tractable record of knowledge production at the postgraduate level. Drawing on PhD theses submitted at Wits University between 2021 and 2025, the project develops a methodology for identifying and quantifying the presence of African-authored and Africa-based scholarship within doctoral citation practices, disaggregated by faculty and discipline. The paper engages with three challenges that sit at the heart of this workshop’s agenda. First, it confronts a fundamental methodological problem: how to define “African knowledge” in ways that are both computationally operable and epistemologically defensible. The project tests one principal proxy — institutional affiliation — but also seeks to explore alternative routes to capture African knowledge production, including diasporic scholarship, against their technical tractability when applied to large unstructured bibliographic datasets. This requires the integration of tools such as OpenAlex, Scopus, and ORCID, and raises broader questions about how existing digital scholarly infrastructures encode — and often obscure — African academic labour. Second, the paper addresses the challenge of building research infrastructure under African conditions. The ETD corpus at Wits presents a case study in the practical obstacles of African DH research: heterogeneous document formats, inconsistent metadata schemas, and the absence of standardised citation extraction practices calibrated for African institutional contexts. The project’s iterative, pilot-based design — moving from data collection through methodology development to scalable metrics — offers a model for capacity-building that prioritises replicability and South-South transferability over high-end computational dependency. Third, and most critically, the project offers an opportunity to reflect on the encounter between computational method and decolonial inquiry. The use of AI-assisted citation analysis is not treated as a neutral technical exercise but as a methodological intervention in an ongoing political debate. Digitisation practices, bibliometric databases, and citation indices are themselves the products of particular choices about what counts as scholarship, which institutions belong to the archive, and whose research warrants discoverability. The paper reflects on how these infrastructural conditions shape what can be measured, and what cannot — and argues that any credible computational approach to African knowledge must account for the political economy of the tools it deploys.
Panel 3 Tue 22 · 09:00 EN
Sanjin Muftić
University of Cape Town Libraries
Heritage collections from the Global South face a compounding discovery problem. Sparse or inconsistent metadata limits findability, and the AI tools that could help address this gap are trained predominantly on Global North data, meaning the collections are the least well served by the tools available to describe them. This paper asks a practical question: how can AI-assisted metadata enrichment make Southern African heritage collections more findable, and what happens when we turn that question around, using these same collections to improve the tools that describe them? The empirical starting point is the Banned Persons Project, a collection of video interviews documenting individuals subjected to apartheid-era banning orders, hosted on UCT’s Ibali Digital Collections platform. Using Microsoft Copilot, available through UCT’s institutional licensing, the DSS team processed existing English-language interview transcripts to generate descriptive keywords, evaluating whether AI assistance could meaningfully improve the discoverability of video content that has historically been difficult to surface. I report on what this process revealed, where it worked, where it fell short, and what the quality of source material had to do with the quality of output. From this starting point, I map out two directions for extending the work. The first is multilingual keyword generation. English-only metadata excludes significant communities of potential users, and the communities most connected to these collections are often not English-speaking ones. Generating and translating keywords into isiXhosa, Afrikaans, or other relevant languages is technically feasible with current tools, but raises immediate questions about accuracy, cultural appropriateness, and who gets to validate what the machine produces. The second direction involves visual collections, specifically, a planned evaluation of Wise (University of Oxford), a computer vision tool for heritage object identification, applied to the Jagger Library Fire Photography Collection, with the question of whether approaches that work there can scale to older and more historically complex photographic archives. Underlying both directions is what I consider the more significant argument. When a collection is well-described and made FAIR-compliant, it does not only serve researchers working with it now, it becomes reusable data that can feed back into training and fine-tuning the models that were potentially too limited to describe it well initially. The act of enrichment is not just a service to the present; it builds capacity for the future, creating a feedback loop between heritage description and model improvement that could, over time, shift the balance of which collections AI handles well. This raises a question I argue deserves more attention. Commercial language models routinely crawl open heritage platforms, extracting content to train systems that offer little back to the communities whose histories they draw on. If Southern African institutions are going to invest in making their collections FAIR and machine-actionable, licensing frameworks should reflect that, prioritising reuse by models developed for and by African research communities, rather than simply contributing to infrastructures built elsewhere.
Panel 4 Tue 22 · 11:00 EN
Tajuddeen Gwadabe, Lydia Kila Taban
Masakhane
Africa is home to approximately 1.5 billion people, with nearly one billion speakers across roughly 50 major languages, and over 2,000 additional languages that remain significantly under-resourced in digital ecosystems. This linguistic diversity presents both a challenge and an opportunity for the development and deployment of inclusive artificial intelligence systems. The Masakhane African Languages Hub addresses this gap through a dual strategy. First, it focuses on building foundational AI infrastructure — datasets and models such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Machine Translation (MT) — to enable African languages to integrate into the broader AI ecosystem. Secondly, it develops community-centered tools, including data collection playbooks and platforms, alongside funded pilot programs, to support decentralised dataset creation and ownership for low-resource languages. This work is guided by the Hub's “4D” model: Discover, Develop, Deploy, and Deliver. Discover focuses on mapping ecosystems and identifying gaps; Develop centres on building datasets and models; Deploy involves integrating these technologies into real-world tools and systems; and Deliver emphasises translating these deployments into measurable social and economic impact. This paper asks: How can community-driven approaches and foundational AI infrastructure be combined to enable scalable, inclusive AI adoption across diverse African language contexts? To address this, we examine two applied learning initiatives within the Hub. ECHO focuses on understanding the impact of language AI on women's livelihoods on the continent, while Lingua Africa explores how AI can be integrated into key development domains such as agriculture, healthcare, and education. These initiatives serve as testbeds for understanding how the 4D model operates in practice, particularly in moving from technological capability to real-world outcomes. Through these efforts, the Hub contributes to developing methodologies for integrating AI into contextually relevant applications, while generating evidence on how language-inclusive AI can drive meaningful and scalable impact across African communities and livelihoods.
Panel 7 Thu 24 · 11:00 EN
Leonard Kibet Kirui
Moi University
For a decade, discourse on African research—from UNESCO's Open Science Recommendation to transformation strategies—has championed infrastructure, implementation gap remains: commitments rarely translate into capacity. Humanities and social sciences face three obstacles: (1) limited connectivity makes cloud-dependent workflows unreliable; (2) funding cycles prioritize pilots over maintenance; (3) scarce training data for African languages marginalizes non-English production. Partnership model reproduces dependency through Northern-controlled platforms, agendas, funding. Question: how can African humanities and social sciences build digital capacity using frugal, offline-first infrastructures, access protocols, curricula, and South-South funding instruments to circumvent dependency? The Africa Multiple Cluster (2019–2025) used action research across its five partner centres. Though successful, implementation obstacles emerged. These challenges are illustrated through a case study of Moi University's African Cluster Centre (MuACC) as the first phase ended, drawing on internal evaluations, observations, and conversations with the team. Digital Literacy Gap Researchers at MuACC had a digital literacy gap requiring capacity-building, delaying DSpace uploads. Compiling contributions was slow due to reservations: some feared appropriation or publication without consent, while others saw no benefit, feeling the project gained at their expense. This trust deficit and lack of incentives became a bottleneck for open sharing, necessitating consent protocols to clarify work status and access rights once materials are on the platform. Low Bandwidth and Unstable Power Low bandwidth and unreliable power at MuACC made hybrid participation impossible: a video-conferencing device failed due to a power fault and expired battery; another meeting halted from low bandwidth. The network switch, shared with a Technical Vocational Entrepreneurship Training (TVET) institute, caused congestion and lacked speed. Fragile connectivity blocked digital research participation. Un-Sustained ICT Infrastructure Funding shortfalls left tools unused after Africa Multiple (AM) 1.0 ended due to lapsed licenses. Unbudgeted maintenance of printers and UPS units undermined durability. Strategies must sustain equipment via locally generated funds. A mindset shift is needed toward edge computing—an edge node (computer or Synology NAS) enables local processing and storage, reducing dependency and allowing retrieval under low bandwidth. Data Curation and Stewardship Data curation was time-consuming; metadata extraction from audio, video, and physical artefacts into Excel required skills, sometimes forcing consultation with Global North curators. DSpace upgrades were difficult due to bugs, so systematic training is needed for future phases. Sustainable research in African humanities requires a frugal, offline-first model that tackles deficits—trust/incentives, infrastructure dependency, and funding asymmetry—by replacing cloud-reliant frameworks with edge computing nodes for local data stewardship, consent protocols that turn researcher reluctance into collaborative ownership, and South-South funding instruments that prioritize maintenance, training, and revenue over donor-cycle driven pilots. South-South Ecologies To realize South-South ecologies, AM should leverage network: work with South Africa Centre for Digital Language Resources (SADiLaR) on data curation training, UbuntuNet on bandwidth procurement, University of Ghana on digital preservation and Maseno on African-language Artificial Intelligence (AI). These partners move AM from donor-dependent pilot to a resilient, continent-wide research ecosystem that builds trust through local protocols, relies on regional infrastructure, and develops trained data curators.
Panel 6 Wed 23 · 11:00 EN
Bruno Allahissem, Mirjam de Bruijn, Luca Bruls, Jelena Prokic, Matthew Sung
Leiden University
The field of digital humanities and computational social sciences is necessarily interdisciplinary. In anthropology, computation is still marginal, although computational means may advance anthropological arguments and understanding of ethnographic data by verifying outside of ethnographic data collection and providing overview through distance reading. Ethnography, by contrast, enhances computation by making it possible to observe what occurs outside of “the frame”. Collaboration is undeniably valuable if scholars want to identify and contextualise culture as a system of ideas and symbols (Bail 2014). In this article we discuss our practice of combining Computational and Ethnographic methods, i.e. ‘Computational Ethnography’. We are especially interested to understand how computational ethnography reveals new layers of complexity and also overlooked biases? This presentation addresses this question based on case studies of an ongoing research project on (trans)national networks of Fulani (see nomadesahel.org). The interdisciplinary study focuses on Fulani networks in a context of increasing conflict in the Sahel, combining historical-ethnographic and computational methods to understand the ‘workings’ of networked conflict. Networks are places that imply agency, decision making and strategy. The project focuses on social media and offline engagements among various sub-ethnic groups of Fulani (three platforms, with an emphasis on Chad and Mali and locals’ connections to other Sahelian countries and diaspora communities in Europe and USA). Due to the linguistic diversity, negotiated access, and cultural nuances across the corpus the identification of patterns requires a triangulation between ethnography (in-depth interviews, participant observation), computation (Social Network Analysis, Natural Language Processing and Computer Vision), and historical analysis. Triangulation allows for a more coherent communication biography of Fulani creators, pages, and groups, whom discuss Islam, ethnicity, nomadism, gender, and conflict in (semi-)public online spaces. With computation we discern the actors on WhatsApp, Facebook, and TikTok, as well as the content the digital brokers participate in. With ethnography we interpret these findings and address the role the platform plays in people’s everyday lives. Firstly, this paper addresses the development of our approach to ‘computational ethnography’ and how it has enhanced understandings of today’s conflicts and the organisation of Fulani networks in the Sahel. Secondly, this paper critically reflects on the epistemological and ethical challenges of computational science and ethnography. The interdisciplinary approach brought a double subjectivity in our findings: 1. ethnographic positionality, 2. algorithmic positionality. Anthropologists view their research process as subjective and institutionally situated. At the same time, algorithms have effect on a researcher’s biases, as they function as a kind of filter to what you see, hear, and do; these technological features hence play a role in the demarcation of a researcher’s field site. This presentation reflects on these patterns of power that come with the unfolding of a new methodology, and that consequently inform knowledge of Fulani history and culture.
Panel 1 Mon 21 · 11:00 FR
Falimatou Pemgbou
University of Ebolowa
Ce travail est une exploration de la manière dont les émotions du personnage de Ramatoulaye, dans Une si longue lettre de Mariama Bâ, se construisent à travers les absences, les sous-entendus et les ellipses du récit. Au moment où l’intelligence artificielle promet de lire les textes littéraires à notre place, et surtout de détecter leurs émotions en quelques secondes, que devient la douleur contenue de Ramatoulaye lorsqu’un algorithme la réduit à un simple code « tristesse » ou « neutralité » ? Les non-dits, les silences, les phrases qu’elle n’achève pas, tout ce qui fait la profondeur de son personnage risque de passer inaperçu. Dans le contexte où la prolifération d’outils d’analyse émotionnelle automatisée peine à saisir les silences, les sous-entendus et la polyphonie du non-dit susceptible de rendre compte des émotions d’un personnage, dont la dignité et la douleur s’expriment davantage par ce qu’elle tait que par ce qu’elle dit, cette analyse se propose de respecter le travail que fait l’intelligence artificielle. Cependant, elle envisage de construire une méthode d’annotation qui prenne au sérieux ce qui ne se dit pas. Pour ce faire, nous nous proposons de tester dans un premier temps quelques extraits du roman avec des outils courants qu’offre l’intelligence artificielle afin de déceler les limites de l’automatisation. Sur le plan méthodologique, nous adoptons une approche mixte, articulant analyse humaine qualitative et réflexion sur l’outillage numérique. En nous inspirant des travaux d’Aline Etienne (Université Paris Nanterre, 2023) sur L’Analyse automatique des émotions dans le texte qui met en lumière, à la fois, l’annotation manuelle d’un corpus et la technique d’affinage de CamenBERT sur le corpus, nous concevrons une grille d’annotation des émotions adaptée aux non-dits, en nous appuyant sur des indices concrets (ruptures syntaxiques, ellipses, blancs typographiques, changements de focalisation) mais aussi sur notre compréhension de la situation et les inférences contextuelles liées à la situation d’énonciation et au partage culturel. Enfin, nous confronterons plusieurs lecteurs humains à cette grille pour vérifier la régularité émotionnelle des interprétations, avant d’envisager, s’il le faut, une aide ponctuelle des humanités numériques et de l’intelligence artificielle pour les tâches les plus répétitives. Cette analyse va confronter la lecture faite par un algorithme et celle faite par un humain afin d’élaborer une grille d’annotation fidèle à la complexité de Ramatoulaye. Cette grille sera enfin appliquée à des scènes clés comme celle de l’annonce du décès, la confrontation, la résilience pour restituer la complexité émotionnelle de l’héroïne.
Panel 7 Thu 24 · 11:00 FR
Mohamadou Konaté
Joseph Ki-Zerbo University
Lorsqu’un chercheur burkinabè en cybersécurité cherche des données sur les incidents cyber touchant les infrastructures critiques de son pays, il ne trouve pratiquement rien. Les bases de données de référence internationales comme KDD Cup 99, CICIDS 2017 et UNSW-NB15 se basent principalement sur des environnements réseau nord-américains ou occidentaux. Les modèles d’IA produits à partir de ces données sont souvent présentés comme universels, mais appliqués au contexte africain, ils reposent sur des hypothèses qui ne reflètent pas la réalité locale. Cette communication, basée sur une recherche doctorale sur la cybersécurité des infrastructures critiques au Burkina Faso, identifie trois défis structurels et propose des voies alternatives. Le premier défi est le manque de données de référence africaines et l’inadéquation des protocoles représentés. Une revue systématique de 89 datasets de détection d’intrusions (Goldschmidt & Chudá, 2025) confirme qu’aucun ne provient d’Afrique. Le trafic africain diffère structurellement : 73 % des connexions restaient sur 3G et moins en 2023 (Cisco, 2020), le protocole USSD non chiffré domine le mobile money, et les attaques par SIM swap (43 % des fraudes) et fraudes par agents (38 %), responsables de 1,5 milliard de dollars de pertes en 2022 (Adongo, 2025), sont absentes des trois datasets de référence. Le deuxième défi est le fossé entre les exigences infrastructurelles des frameworks d’IA et les capacités réelles des institutions africaines. Les systèmes de détection d’intrusions reposant sur le deep learning nécessitent des GPU NVIDIA, entre 16 et 32 Go de RAM, ainsi qu’une connexion Internet constante (Dietz et al., 2024) — des exigences inaccessibles pour de nombreuses universités et centres de recherche africains. La vitesse médiane fixe au Burkina Faso était de 42,46 Mbps début 2024, en baisse de 5,3 % sur un an (Ookla/DataReportal, 2024). Les universités d’Afrique de l’Ouest disposent de 10 à 100 Mbps, contre 3 Gbps recommandés aux États-Unis (Banque mondiale, 2020), avec une bande passante coûtant 50 fois plus cher qu’en Europe (Bandwidth Consortium, 2005). Le troisième défi est l’invisibilité épistémique des jeunes chercheurs du continent. Contraints de travailler sur des données étrangères et des outils hors de portée, ils produisent structurellement une science décontextualisée — rejoignant les constats de Buolamwini & Gebru (2018) sur les biais algorithmiques et de Cremer et al. (2022) sur la pénurie de données en cybersécurité. La méthodologie repose sur une revue systématique de 89 ensembles de données, une analyse comparative des protocoles et des vecteurs d’attaque entre l’Occident et l’Afrique, ainsi qu’une confrontation avec les défis réels du contexte burkinabè et africain. Trois pistes sont envisagées : la création collaborative de datasets africains, inspirée par le modèle de Masakhane pour le traitement du langage naturel ; la conception d’architectures légères (lightweight AI) entraînables hors ligne ; et un plaidoyer pour des benchmarks africains reconnus internationalement. Cette communication s’inscrit dans la continuité des travaux de Madore (2021) sur les humanités numériques au Burkina Faso et de Madore & Hiribarren (2025) concernant l’IA et les archives africaines. La décolonisation de l’IA évolue de la critique vers la construction.
Panel 2 Mon 21 · 16:00 FR
Aminata Kane, Augustin Ndione
Cheikh Anta Diop University
Cet atelier propose une réflexion critique sur la place des systèmes de connaissances africains dans les environnements numériques contemporains à partir du cas des manuscrits ajami, textes rédigés en langues africaines transcrites à l’aide de l’alphabet arabe. Longtemps marginalisés dans les récits dominants de l’histoire de l’écriture et des savoirs en Afrique, ces corpus constituent pourtant des archives vivantes témoignant de traditions intellectuelles, linguistiques et documentaires anciennes, dynamiques et translocales. À l’heure du développement des humanités numériques et de l’intelligence artificielle, la numérisation, l’annotation et le traitement automatique de ces manuscrits ouvrent des perspectives importantes en matière de conservation, d’accessibilité et de transmission des patrimoines africains. Toutefois, l’intégration des corpus ajami dans les infrastructures numériques contemporaines soulève des enjeux politiques, techniques et épistémiques majeurs. Les standards documentaires, les modèles de données, les dispositifs d’indexation ainsi que les systèmes d’intelligence artificielle mobilisés dans le traitement des patrimoines écrits ont été majoritairement conçus à partir de référentiels linguistiques et technologiques occidentaux, souvent peu adaptés aux spécificités des langues africaines et des traditions scripturales non latines. Dans ce contexte, cette communication posera une question centrale : que devient un système de connaissances africain lorsqu’il entre dans des infrastructures numériques pensées ailleurs ? À partir d’une approche située au croisement des sciences de l’information et de la communication, des études patrimoniales et des humanités numériques critiques, l’intervention analysera les risques de marginalisation numérique des corpus africains : invisibilisation des spécificités linguistiques, standardisation des formes d’écriture, dépendance aux modèles dominants d’IA ou encore réduction des dimensions sociales, mémorielles et culturelles des documents à de simples données exploitables. L’intervention défendra l’idée que les manuscrits ajami constituent un terrain privilégié pour repenser les humanités numériques africaines à partir d’une logique de justice épistémique, de souveraineté documentaire et de co-construction des infrastructures de savoir. Elle montrera également comment les pratiques d’édition numérique, d’encodage et de valorisation patrimoniale peuvent devenir des espaces de réappropriation culturelle, de médiation des savoirs et de résistance face aux asymétries numériques contemporaines.
Panel 3 Tue 22 · 09:00 FR
Frédérick Madore
University of Bayreuth
Panel 1 Mon 21 · 11:00 EN
Augustine A. Farinola
University of Alberta
This paper emerges from a childhood shaped by African folktales, told through song, gesture, and moral questioning in familial and communal settings. Drawing on this performative tradition, I examine how African folktales have moved from orally transmitted narratives to written texts, classroom materials, and now digital and algorithmic objects of study. I present “Moral Ecology of African Folktales” (digitalafricanstorytelling.com), a Digital Humanities project that builds an online archive and exploratory infrastructure for studying African folktales with and against contemporary AI-driven methods. Currently, the project includes 385 folktales from five African regions, annotated into 55 themes derived from moral, social, and cosmological concerns in the stories. These themes inform a relational database and interactive dashboard designed to visualise narrative patterns, moral motifs, and regional variations. Methodologically, the project uses computational pattern detection and thematic querying to trace how recurring figures (such as the tortoise) and moral scenarios shift as stories move through transcription, translation, recirculation, animation, and algorithmic analysis. At the same time, I argue for a cautious deployment of AI techniques so as not to erase the performative, sonic, and situated dimensions of the narratives—the voices of narrators, the music that accompanies stories, and the contexts in which they are told. Finally, the project experiments with centring African knowledge systems in digital research design by grounding the data model in African moral categories and narrative logics.
Group discussion Tue 22 · 14:00 EN
Menno van Zaanen
South African Centre for Digital Language Resources (SADiLaR)
Panel 5 Wed 23 · 09:00 EN
Benito Trollip, Sanjin Muftić
South African Centre for Digital Language Resources (SADiLaR) · University of Cape Town Libraries
Barriers to research engagement in South Africa, including awareness of available resources, institutional support, and socio-economic factors, shape how scholars in the Global South produce and share knowledge. Platforms such as the SADiLaR repository and the Ibali digital collections at the University of Cape Town (UCT) have been developed to facilitate the sharing and reuse of indigenous language resources. Yet a critical question remains: to what extent do FAIR-compliant platforms translate into meaningful preservation and reuse of indigenous language resources in Southern Africa? This paper addresses that question through a comparative case study of two initiatives that approach resource sharing from different vantage points: the South African Centre for Digital Language Resources (SADiLaR) repository and the Ibali showcase platform at UCT. Both serve the broader digital humanities research community but cater to distinctly different audiences and modes of engagement. Central to our discussion are the FAIR principles (Findability, Accessibility, Interoperability, and Reuse) which provide a framework for making indigenous language resources not only available but actively usable. FAIR applies to data, metadata, and the infrastructure supporting them, with a particular emphasis on machine-actionability, enabling computational systems to locate and reuse materials at scale. Embedding FAIR into open science practices creates conditions where indigenous languages are not only preserved but remain active within global knowledge networks. Our comparative analysis reveals how different design philosophies produce different forms of engagement. Ibali offers a visually rich, curated experience, such as browsing an early isiXhosa newspaper archive in a layout faithful to the original. SADiLaR’s repository, by contrast, makes structured datasets available for download and reuse, mainly enabling research related to natural language research. As Adam and Kaur (2021:167) state “Effective implementation and management of institutional repositories can streamline the contribution of African researchers and facilitate discovery and access to their findings globally”. This statement holds across both field-specific repositories and showcase platforms. Designing for FAIR means accommodating diverse modes of engagement without compromising accessibility or reuse, a challenge that requires teams with interdisciplinary skill sets. Making resources FAIR is not only about access; it is about enabling reuse in meaningful ways. This distinction becomes urgent in the age of AI-driven research. Autonomous systems increasingly crawl digital heritage sites, feeding indigenous language data into dominant language models. This raises ethical questions that FAIR alone cannot resolve: does this process benefit the communities whose languages are represented? Licensing frameworks become central here. Instruments such as the Nwulite Obodo Open Data License offer models for acknowledging origin and, where necessary, restricting certain uses. Reuse must reinforce cultural memory and respect for the communities behind the data, not merely serve computational convenience. We argue that even FAIR practices do more than support present research, they safeguard the capacity to conduct future research, preserving the memory of the world. The platforms we examine act as bridges between digital infrastructure and living language communities, demonstrating that sustainable, equitable digital humanities practice in Africa depends on building systems that serve both researchers and the communities whose knowledge they hold.
Keynote Thu 24 · 09:30 FR
Emmanuel Ngue Um
University of Yaoundé 1
Panel 1 Mon 21 · 11:00 EN
Hammed Olalekan Lawal
University of Bayreuth
The integration of Artificial Intelligence (AI) into African Studies offers transformative potential for cultural preservation (Olukoya et al., 2025; Ndede, 2026), yet it risks the “flattening” of indigenous epistemologies and authenticities (Moatemsu & Nigamananda, 2026). This study investigates the re-mediation of Ìtàn Ìjàpá (Yorùbá tortoise folktales) through AI-generated videos. It explores the relational dynamics between the AI-content creators and digital audiences on TikTok and Facebook. Historically, Ìjàpá has served as a pedagogical and performative tool for negotiating Yorùbá morality and identity (Taiwo, 2019). However, as creators “re-code” these oral traditions into AI-generated videos, the figure of the trickster (Ìjàpá) becomes a site of tension between what is presented through our ancestral orature versus the modern day algorithmic AI architecture. This research addresses a critical methodological gap in Digital Humanities: how to evaluate the “cultural accuracy” and “narrative authority” of AI-generated indigenous content. Adopting a Relational Digital Ethnographic framework, I move beyond a unidirectional analysis of AI “outputs.” Instead, I examine the “collaborative authority” (Kashaemei & Sousa, 2026) that emerges when AI-content producers, platform algorithms, and digital audiences interact. The study interrogates how the “uncanny” visuality of AI-generated characters triggers performative responses from Yorùbá audiences that either affirm or reject the digital “re-shelling” of their heritage. Methodologically, the study employs a three-generational interview framework (Kashaemei & Sousa, 2026). By conducting paired interviews with AI-creators and viewers across different age cohorts, alongside traditional Yorùbá cultural custodians, the research explores how “authority” is renegotiated in networked spaces. This approach is synthesised with systematic observation of platform-specific engagement patterns on Facebook and TikTok, building on established methods for analysing creator-algorithm relationships (Jerasa & Burriss, 2024). To conclude, this research is a response to the workshop’s call, moving from “description to design” focusing on the Yorùbá epistemologies utilising digital and relational methods. It argues that AI-generated folktales are not static digital artefacts but “relational events” that test the limits of Western-centric AI models. This study will also contribute to the decolonial project by questioning whether AI serves as a tool for cultural reclamation or a new frontier for indigenous displacement.
Panel 5 Wed 23 · 09:00 FR
Evelyne Amana
University of Yaoundé I
L’Afrique est un continent naturellement riche par sa diversité naturelle et culturelle. Elle se caractérise aussi par des regroupements ethniques qui utilisent des langues vernaculaires précises. Le Cameroun, appelé Afrique en miniature par sa diversité linguistique, comprend environ 300 langues vernaculaires dont les plus répandues sont : l’ewondo, le bassa, le fufulde, le ghomala ou le bamoun, etc. Avec l’avènement de la colonisation, les langues étrangères ont été imposées aux peuples africains. Le pays contexte de cette étude a adopté deux langues officielles pour l’enseignement-apprentissage et le service administratif, avec deux autres qui sont dites secondes. Dans le cadre de la recherche, deux langues dominent : le français et l’anglais. Les langues locales sont ainsi discriminées et victimes des préjugés. Le manque de données numériques et des faibles ressources en langues africaines renforcent le fossé dans le cadre de la recherche. Cette situation a suscité la prise de conscience des chercheurs qui œuvrent actuellement pour leur intégration dans les numériques à l’ère où elles sont ignorées par l’Intelligence artificielle. La question de recherche est la suivante : comment la recherche en sciences sociales doit-elle contribuer à la pérennisation des langues africaines à l’ère du numérique ? L’hypothèse de travail est ainsi formulée : par l’adaptation des algorithmes linguistiques et la traduction de proximité, la recherche en sciences sociales contribue à pérenniser les langues locales. L’objectif de l’étude est d’examiner les stratégies par lesquelles la recherche scientifique par les africains contribue à la survie des langues vernaculaires à l’ère de l’Intelligence artificielle. Les théories issues de l’ingénierie informatique permettent d’expliquer la problématique développée et de proposer des stratégies. La recherche est mixte avec recours aux entretiens et questionnaires adressés aux chercheurs des disciplines variées. L’échantillon obtenu par choix raisonné est constitué de linguistes, informaticiens et psychologues. Les résultats attendus : conception des architectures algorithmiques adaptées pour la numérisation, l’archivage et l’enregistrement multimodaux.
Keynote Mon 21 · 14:00 EN
Sarah Oberbichler
University of Luxembourg
Group discussion Thu 24 · 14:00 EN
Vincent Hiribarren
King's College London
Panel 4 Tue 22 · 11:00 FR
Eliette Ngo Tjomb Assembe
University of Yaoundé 1
La traduction automatique des langues africaines peu dotées fait face à des défis qui dépassent largement l’ingénierie linguistique. Cet article s’appuie sur le développement d’un système de traduction automatique à base de règles (RBMT) construit dans une architecture FLEx/FlexTrans/Apertium et testé sur un corpus parallèle de 124 paires de phrases alignées, pour la paire français-ewondo, une langue bantoue du Cameroun. L’étude des transferts structurels met en évidence un premier type de problèmes, tels que la concordance allitérative des classes nominales bantoues, l’expression de la possession via le connectif génitif, l’absence d’inversion sujet-verbe dans les interrogatives ainsi que le réordonnancement des constituants, qui ont tous requis la formalisation de règles computationnelles dédiées. Ces phénomènes illustrent que la modélisation d’une langue à richesse morphologique et agglutinante requiert une description contrastive préalable, laquelle ne peut être précisément assurée par les grands modèles de langage contemporains pour l’ewondo. Par ailleurs, l’analyse des transferts lexicaux et sémantiques révèle des défis encore plus sérieux. Trois types de non-correspondances ont été identifiés : les lacunes lexicales, où les concepts juridiques et administratifs français n’ont pas d’équivalent en ewondo et exigent des stratégies de compensation impliquant chacune une décision épistémologique ; les divergences aspectuelles, puisque l’ewondo encode dans sa morphologie verbale des oppositions statiques/dynamiques absentes en français ; et les constructions existentielles impersonnelles, dont la restructuration syntaxique et le choix du verbe existentiel approprié en ewondo dépendent des traits sémantiques du nom thème et résistent à une réduction à une règle générale. L’évaluation du système à l’aide des métriques BLEU et ChrF++ confirme ces observations. Le score BLEU de 30,79 % reflète une couverture solide des transferts structurels, tandis que la baisse progressive selon l’ordre des n-grammes (56,0 % au niveau unigramme contre 16,4 % au niveau 4-grammes) révèle la difficulté du système à reproduire des séquences longues où s’accumulent les non-correspondances lexicales et sémantiques. Ces résultats montrent que les non-correspondances profondes définissent une limite structurelle. Là où les deux langues ne partagent pas les mêmes catégories conceptuelles, la traduction automatique atteint une frontière que seul le jugement du linguiste-locuteur peut franchir. Cette limite est un indicateur épistémologique, cartographiant les points où les systèmes de savoir encodés en ewondo résistent à la formalisation computationnelle et, ce faisant, remettent en question les présupposés intégrés dans les infrastructures numériques conçues pour d’autres langues.