ZMO Programmatic Texts · EN
For Whom and For What Purpose? A Position Paper on Digital Humanities and AI in African Studies
Abstract
African Studies scholars have been involved in Digital Humanities (DH) projects since the 2000s. Recent advances in artificial intelligence (AI), particularly large language models (LLMs), have expanded the possibilities for textual analysis and archival research, but implementation raises difficult questions: who controls access, whose data trains these systems without consent, whose languages remain underserved, and who bears the labour and environmental costs. This position paper is the collective work of a Volkswagen Foundation-funded scoping workshop in Hanover, Germany, in February 2026, with twenty-six scholars from sixteen countries. Across four working groups – language technologies; archives and visual heritage; infrastructure, governance, and access; and epistemologies, decoloniality, and ethics – one question recurred: for whom and for what purpose is this work undertaken? Ownership and sovereignty, access, sustainability, and standardisation emerged as transversal problems, which we set in dialogue with critiques of data colonialism and extractivism. We propose situated practices: licensing that conditions openness on return to source communities; digitisation agreements as living documents; ontologies built collaboratively; design grounded in access, governance, and sovereignty; and ethics treated as a process rather than a deliverable. Many of the problems we examine are common across DH and AI: consent, standardisation that erases variation, algorithmic bias, hidden labour, and the funding of innovation over maintenance. African contexts make them sharper and more visible, and Africa-based practice has produced concrete responses. We therefore argue that African epistemologies and decolonial critique belong at the centre of both fields, and that African scholars, practitioners, and institutions belong where these technologies are designed and governed.
- Keywords
- African epistemologies, African languages, data colonialism, decolonial digital humanities (DH), digital archives, digital preservation, digital sovereignty, large language models (LLMs), natural language processing (NLP), standardisation
1 Introduction
Between 18 and 20 February 2026, twenty-six scholars from sixteen countries met in Hannover, Germany, for “Charting New Territory: Digital Humanities and AI in African Studies”, a scoping workshop funded by the Volkswagen Foundation and convened by Frédérick Madore and Vincent Hiribarren.[1] Participants included linguists, literary scholars, historians, art historians, anthropologists, archivists, AI governance specialists, Islamic Studies scholars, and digital humanities (DH) practitioners working on and from the African continent, Europe, and North America. The programme dispensed with keynote lectures and panel presentations. Plenary sessions, four thematic working groups, World Café rounds, poster presentations, and collaborative writing exercises alternated across three days, with collective drafting underway by the end of the second day.
DH has a longer history in African Studies than is often acknowledged, from the Transatlantic Slave Trade Database[2] to more recent projects such as the African Ajami Library,[3] Open Restitution Africa,[4] and Archivi.ng.[5] That history has not always been an easy one. Early efforts to digitise Africa’s documentary heritage were sharply contested: participants in the Aluka project, which digitised records of the Southern African liberation struggles, suspected it of being one more North American attempt to appropriate the continent’s patrimony (Isaacman et al. 2005), and Pickover, a South African curator, described the export of digitised heritage to the Global North as “a new form of cultural theft” (2005, 10). Nearly a decade later, a survey of the field still registered African librarians’ “frustration or resentment at domination or interference from the North” (Barringer et al. 2014, 6), a mistrust that persisted for years in countries such as South Africa.
Outside South Africa, the Centre for Digital Humanities at the University of Lagos (CEDHUL) launched a Digital Humanities summer school in 2017 to build interdisciplinary capacity for DH work in Nigeria, but even that initiative has relied on funding from the Global North. A 2024 special issue of Reviews in Digital Humanities maps these initiatives and foregrounds the continent-based projects now shaping the field (Guiliano et al. 2024). Rapid advances in artificial intelligence (AI), and in large language models (LLMs) in particular, have expanded the possibilities for textual analysis and cross-cultural research, but implementation raises difficult questions. Decisions about what gets digitised, how it is catalogued, and who controls access are not neutral; they tend to favour institutions that already have resources in place. Most AI systems still underperform on African languages, and the pace of adoption frequently outstrips attention to long-term preservation and local capacity. These earlier tensions resurfaced in the workshop’s discussion of the “digital saviour complex” (Shringarpure 2020): projects led from the Global North that reproduce colonial dynamics even as they claim to democratise access.
Three themes organised the workshop’s discussions. “Methodological Integration & Digital Preservation” addressed how to adapt AI for African languages and build sustainable preservation models. Participants documented existing computational methods and infrastructure barriers, then asked what technical standards and protocols might suit resource-constrained settings. “Equitable Collaboration” focused on partnership models that address power imbalances in North–South and South–South research relationships, and on mechanisms for sharing resources within the continent, with African epistemologies at the centre. In DH and AI design, that commitment means favouring oral and multimodal sources alongside text, building ontologies collaboratively and holding them open to revision, and treating source communities as co-owners of the digital record. “Ethical Frameworks & Digital Sovereignty” examined how AI implementation can protect community control over data and promote equitable scholarly exchange, and treated community engagement as co-creation rather than consultation.
A shared concern ran beneath all three themes without ever becoming a topic in its own right: the costs of AI itself, which fall disproportionately on the African continent. The labour that trains and moderates commercial AI systems is concentrated in countries such as Kenya, where annotators and content moderators face precarious conditions and rates of psychological distress documented well above those of comparable workers elsewhere (Perrigo 2026). The environmental burden is similarly skewed: the energy, water, and critical minerals that AI infrastructure depends on are drawn largely from the Global South, and African communities remain among those most exposed to the climate crisis it aggravates (Kyomuhendo 2025; Musa 2025). We are specialists in neither domain and cannot offer solutions, but they have not led us to reject AI. The workshop convened around these tools because they can open up research that would otherwise stay out of reach. The task, as we see it, is to weigh each application, asking whether the benefits to scholarship and to source communities justify the cost, and to accept that for some applications they will not. We highlight these costs so that the practices we propose, and the requests we make of funders, policymakers, and institutions, rest on an honest account of who pays for AI and how.
This position paper is the collective work of four working groups (Language Technologies, NLP & Corpora; The Archive: Preservation, Community Custody & Visual Heritage; Infrastructure, Governance & Access; and Epistemologies, Decoloniality & Ethical Frameworks). Each group paired disciplinary with practitioner expertise. The workshop gave one day to each of the three themes set out above. Each group took up the theme of the day from its own angle, so every theme was examined four times over. Poster sessions and rotating discussion rounds carried arguments between the groups, so that a point raised in one could shape the reasoning of another. The four sections that follow come directly from those exchanges, and where the groups disagreed the text says so.
Alongside its call for situated, community-grounded practice, the paper makes a second claim about why African Studies matters to these fields. Many of the problems we examine are not unique to Africa. Ownership and consent, standardisation that erases variation, metadata that does not fit, algorithmic bias, the hidden labour behind “clean” data, and the mismatch between innovation funding and long-term maintenance are familiar across DH and AI in many parts of the world, the Global North included. African contexts, however, make them more visible, more acute, and more theoretically revealing. Colonial and postcolonial asymmetries of control, the suppression of languages in settings of deep linguistic plurality, the extraction of data, access conceived as individual rather than communal, and the fallacy of a “universal” way of knowing inherited from a narrow Euro-American genealogy are all sharper here, and Africa-based practice has produced concrete responses to them. Read this way, African Studies stops being a special case at the margins of DH and AI and becomes a resource for rethinking both. We write for several audiences at once: DH and AI researchers and system designers, African Studies scholars, and the funders, policymakers, and institutions that shape what gets built. To each, the paper offers a shared vocabulary for these problems, a set of situated practices for working with them, and the argument that DH and AI everywhere stand to gain from taking African realities seriously.
The four sections work in different registers but share a substantive vocabulary. Ownership and sovereignty are layered and contested in every domain: manuscript collections, mobile money infrastructure, language data. Access is more than connectivity: language, literacy, devices, electricity, and cost, as well as the publics that shared use makes possible. Sustainability is a structural problem, shaped by funding cycles, succession planning, and the misalignment between those who build collections and those responsible for keeping them alive. Standardisation secures findability and interoperability at costs felt in every section: it flattens dialect variation, privileges Western taxonomies, and defaults to dominant scripts. Running through all of this is a critique of data colonialism (Couldry and Mejias 2019) and data extractivism (Birhane 2023; Kwet 2019): the appropriation of Global South data without adequate return, which links technical questions to political ones. Minimal computing (Risam and Gil 2022) recurs across the sections, recast from a fallback into a deliberate choice. One question, more than any other, returned in every session: for whom and for what purpose? It pushed participants beyond academic audiences towards the communities whose heritage, languages, and histories are at stake.
This paper does not resolve these tensions, and the workshop could not fully overcome the asymmetries it set out to address. European visa procedures prevented one colleague from attending, a direct consequence of the inequalities under discussion. The anglophone bias built into the field’s tools, conferences, and networks meant that we failed to reach many francophone West African colleagues, particularly outside Senegal, where DH infrastructure is thinner than in anglophone contexts. These constraints belong to the argument that follows: the field must describe its own conditions honestly before it can change them.
2 Language Technologies, NLP & Corpora
Natural language processing (NLP) – the computational handling of text and speech behind translation, search, transcription, voice assistants, and most language-facing AI – sits between DH and AI, and depends on the corpora that train it. Language, in the form of corpora, is the material from which much of African Studies’ digital infrastructure is built. The suppression of African languages in favour of European ones runs through the continent’s linguistic history, and both the study of language there and the technologies now built on it carry that history. That inheritance is easy to overlook, because language reaches AI systems as a technical input, and technical inputs look neutral.
For this reason, decisions about what language content to record, in which orthography, under which language code, and with what licence are more than academic questions: they require an ethical and historically informed standpoint. This aligns with the two core ethical questions that we believe all DH and AI projects need to be able to answer before making such decisions: who is the project for, and what purpose does it serve for relevant communities? Before choices can be made between licences that require credit (e.g. CC-BY) and those that waive it (e.g. CC0), one must first determine who profits from the project and how its impact might negatively affect the very communities it claims to give voice to.
For African languages, the corpora and the pipelines that make use of them in NLP have been shaped by infrastructures designed elsewhere, leaving African DH and AI particularly prone to the biases of the Global North. The cases that follow show what this means in practice.
2.1 The “low-resource” label
In computational linguistics, “low-resource” describes a real condition: most African languages have limited annotated corpora, few computational tools, and little representation in the training data of LLMs. While the term describes the data, grammatically it modifies the language. This frames scarcity as a property of the language rather than the resources assembled for it. These languages are not impoverished; what is scarce is the machine-readable material that NLP pipelines can ingest, and this scarcity is a matter of collection and infrastructure.
Adebara (2025) illustrates the politics of such labelling through the idea of “data flaring”. The metaphor is taken from the petroleum industry, where gas flaring describes the wasteful burning of natural gas during oil extraction. African languages, by this account, are not under-resourced so much as under-collected. What gets collected is poorly preserved and rarely put to use. Vast quantities of African linguistic data circulate every day through radio broadcasts, religious gatherings, market exchanges, oral histories, and analogue archives. Yet none of these are visible to the bots behind Common Crawl, an open archive of scraped web text that trains most LLMs, or to the other web-scraping pipelines that feed them. What gets called scarcity, then, is largely about those pipelines: where data sits, and the technologies that have been built to reach it. Hussen et al. (2025) find that only four African languages – Amharic, Swahili, Afrikaans, and Malagasy – are consistently treated across the LLMs they survey, which means that over 98% of African languages remain unsupported.
If scarcity is structural, it can be addressed structurally, and the work need not begin from written text. The Centre de Linguistique Appliquée de Dakar (CLAD) has built corpora for Wolof, Pulaar, Séeréer, and other Senegalese national languages by starting from spoken material. The work spans an English–French–Wolof banking dataset of 9,791 sentences with four hours of audio, a community-driven Wolof monolingual dictionary, a Wolof–French translation pipeline, and automatic speech recognition (ASR) and text-to-speech (TTS) development across Wolof, Pulaar, and Séeréer. All of it is grounded in domains where speech is the medium of everyday transaction. Partnerships with Senegalese banks, hospitals, and public institutions created the demand. Source data is held offline at Cheikh Anta Diop University, while the banking dataset is published openly on Kaggle for reuse.
OlongoAfrica’s Multilingual Anthology[6] resists the same logic from a different angle: the assumption that African-language work should flow outwards, towards non-African readers and pipelines. Most African-language translation work moves out of African languages and into English or French; the anthology commissions translation in the opposite direction, paying writers, translators, and voiceover talent to render published short stories into underserved African languages, with audio versions distributable on WhatsApp for readers who are not literate even in their first language.
The Black Orpheus Revisited[7] project (2025) extends this diagnosis to print and addresses it through DH – in this case digitisation – rather than AI pipelines. Black Orpheus was a major literary journal that published some of Africa’s and the African diaspora’s most important writers between 1957 and 1975. Because literary magazines are ephemeral, physical copies are rare, even in Ìbàdàn, where the journal was published. By digitising extant copies of the magazine and making them available online, Black Orpheus Revisited has digitally repatriated these materials for any African reader who wishes to read them. The project also partners with a Lagos library to preserve the physical objects for on-site access. The journal’s ownership sits in copyright limbo – funded first by the Western Nigerian Ministry of Education, then by a series of international funders and publishers, none of whom now claim it. The project responds by keeping web access free while restricting download and reproduction.
Metadata is itself a form of access: a lightweight layer that can be browsed on the low bandwidth where heavy image scans will not load. Black Orpheus Revisited built such a layer for the magazine’s contents, so that users can search and navigate them at low data cost before opening the heavier scans. The African Literary Metadata (ALMEDA)[8] project collects and structures metadata at a larger scale to address the issue of access to African-language literatures and other forms of spoken, performed, and ephemeral print literatures that are not usually catalogued in libraries and, consequently, do not appear in global databases. This project has the infrastructural capacity to build future corpora for LLM development since it creates Linked Open Data on African language materials in the languages that the materials are created in, as well as in English, French, Portuguese, and Arabic. This provides maximum findability of materials for individual users, but could also be the basis of corpus creation when materials have been digitised.
The term “low-resource language” is therefore an ambivalent one: on the one hand, it accurately describes the conditions computational linguists confront when working with African languages. On the other hand, the term also has negative consequences: if the deficit is understood to lie in the languages themselves, the response to “produce more data” justifies the continued extraction from communities long cast as a “data rich continent” (Birhane 2020). This reproduces the use of Africa as a site of “raw fact: of the minutiae from which Euromodernity might fashion its testable theories and transcendent truths. Just as it has long capitalised on non-Western ‘raw materials’ by ostensibly adding value and refinement to them” (Comaroff and Comaroff 2012, 114). To avoid this extractive logic when applied to African languages, NLP and LLM pipelines, we need instead a new structural design of the field by repositioning ourselves in relation to collection methods, ontologies, and platforms. A good place to start is with speech-first pipelines, partnerships with broadcasters and oral-tradition custodians, investment in capturing naturally occurring language use in the “real world” rather than scraping only online, digitally available language, and a clear accounting of what existing models cannot see.
2.2 The limitations of the textual method
Most NLP pipelines, including the ones now being adapted for African languages, treat language as a finite, divisible, written object. The units assumed by the field – such as tokens, lemmas, sentences, and documents – are all derived from a particular textual ontology, and that ontology was developed for alphabetic Indo-European languages with long traditions of grammar standardisation. While this ontology serves text-classification benchmarks well, it reduces a great deal of what African languages actually do.
African researchers and ventures are not waiting for these systems to improve on their own. On the evaluation side, benchmarks such as AfroBench[9] and IrokoBench (Adelani et al. 2025) document how far LLMs still fall short across African languages. On the development side, grassroots and commercial initiatives are strengthening African-language NLP: Masakhane, a grassroots organisation whose mission is “to strengthen and spur NLP research in African languages, for Africans, by Africans”[10]; Lelapa AI,[11] a commercial venture that in 2024 launched what it described as “Africa’s first” LLM; and a GSMA-led collaboration, announced in October 2025, to develop “inclusive African AI language models” with the continent’s major mobile operators.[12] Much of this work, however, still operates within the textual ontology itself, whose limits show most clearly in the features that writing does not encode.
Verbal gestures, extra-grammatical units that often use sounds outside a language’s phonemic system, are one register that this textual system cannot accommodate (Pillion et al. 2019). Ideophones are another: marked words that depict sensory imagery, first systematically described in West African linguistics and a major lexical class in many languages (Dingemanse 2018). Tonal contrasts that carry lexical weight are captured unevenly by orthographies that were often standardised for missionary or colonial purposes, and they are routinely flattened in downstream processing.
The N|uu dictionary[13] (Sands and Jones 2022) illustrates what is lost when this reduction is applied without correction. N|uu, spoken in northern South Africa and Namibia and now critically endangered, makes systematic use of click consonants and tonal contrasts that no Latin-script orthography fully represents. Twenty years of fieldwork provided language data that allowed for the production of a dictionary covering N|uu, Nama, Afrikaans, and English, but its linguistic value depends on the audio recordings that travel with each entry. The mobile app and online portal foreground these recordings deliberately. In this project, it becomes apparent that written form is just one access point to linguistic study and should not be seen as the canonical default. Had the dictionary treated audio as supplementary metadata rather than as primary linguistic data, the resource would have been unusable for the very community for whom it was built.
The Yorùbá Names[14] project handles a related problem at the level of the orthography of single letters. Yorùbá uses diacritics to mark tone and vowel quality; many users typing on mobile devices cannot or do not enter them. This has been a major problem for the language, from the first colonial collections of the script in the nineteenth century in roughshod orthographies, to the present time of standardised keyboards and software that inadequately integrate Yorùbá diacritics, or fail to do so altogether. As pointed out by Túbọ̀sún (2022), there are numerous errors in the formal catalogue of Yorùbá materials, which negatively impacts findability. The Yorùbá Names project uses diacritic search technology, which matches base characters with accented characters. This means that a search for “Adebola” should return entries for Adébọ́lá even when the searcher does not know the correct diacritics. The system therefore treats tone-marked and tone-stripped forms as equivalent during retrieval while preserving the marked form as canonical for display. The project also generates pronunciation through a concatenated text-to-speech system, so that every name has an audible form alongside its written one. The result is a dictionary built around the recognition that a name is, in the first instance, spoken.
The N|uu dictionary and Yorùbá Names project are partial correctives, not solutions. Both still rely on transcription, search, and retrieval infrastructures built around alphabetic input somewhere in the pipeline. Multimodal pipelines built on top of LLMs trained largely in English and other Western languages do not become structurally multimodal by adding an audio encoder. How to balance the limits of standardisation against the need for reliable, interoperable standards remains an open question.
2.3 Enabling the renegotiation of standardisation
The reduction of complexity imposed by written text being used as the default source of language data is redoubled by the standards used to organise this data. Standards are necessary for data work: cataloguing, searching, interoperability, and any kind of cross-lingual analysis depend on shared identifiers, and abandoning that shared layer would make most of the work in DH and AI impossible. Yet the question remains: what do existing standards encode, to whose benefit and at whose cost?
A major example is the International Organization for Standardization (ISO) coding scheme that indexes named language varieties. This most widely used standard for language description affords English a fine-grained typology – en-US, en-GB, en-NG, en-AU, en-ZA, and many others – while most major African languages are conflated under a single tag. Hausa is hau. Wolof is wol. Yorùbá is yor. Linguistic varieties at least as differentiated as the various Englishes are flattened to a single identifier in the metadata layer used by every search system, app store, content management system, and AI training pipeline that uses them.
What Fricker calls “a wrong done to someone specifically in their capacity as a knower” (2007, 1) describes the effect of such coding from the outside. No individual standards body intended this. The cause lies in a global governance of linguistic localisation that has not drawn on the experience of speakers whose languages do not fit the model, and in standards made to fit already dominant languages such as English. A related reduction operates at the level of script and orthography, which Section 5 takes up in its discussion of Ajami.
Constructive solutions are, however, possible. The Digitizing-Endemann Dictionary[15] project works on Karl Endemann’s 1911 Wörterbuch der Sothosprache, a dictionary covering some 37 Sotho-related languages and dialects. Rather than collapsing the varieties into a single normalised form, the project structures the resource as a bi-directional electronic dictionary that treats each Sotho language and dialect as a distinct entry. The team’s interventions are marked explicitly, so that users can distinguish historical content from contemporary commentary. Endemann’s 1911 phonetic orthography is preserved alongside modern spellings, and shifts in meaning since 1911 are flagged with a graded evaluation rather than collapsed into modern equivalents. The result is more cumbersome than a unified resource and harder to query at scale, but it is a more faithful record of the source.
Standardisation always secures one kind of legibility at the cost of another. This negotiation between the universal and the local is particularly complex when it comes to African languages that have a long history, to their detriment, of being analysed through European languages, standards, and methods. Some practical ways forward are to design standards to accommodate multiple forms, to document what each standardisation choice excludes, and to treat orthographies and other standardisation frameworks as organisational choices, which can be altered, if need be, rather than as immutable technical defaults.
2.4 Licences and the revival of communities
Once a corpus exists, the next area of complexity that needs to be addressed is how it circulates. “Open access” is sometimes spoken of as a single ethical position, as though publishing a dataset under a permissive licence settles the question of who benefits from its onward use. This approach is far too simplistic, especially when considering the African context. Licences that require attribution (e.g. CC-BY) and licences that waive even that (e.g. CC0) assume the social and political equality of all downstream users. This works tolerably well when the asymmetries between users are small. However, when applied to African-language datasets, the lack of such equality becomes a problem: when communities release their linguistic resources for open use, value transfers upwards, from speakers and producers to the actors with the infrastructure to package, fine-tune, and resell it. Ultimately, this can lead to attribution or remuneration at the commercial, rather than the communal, level. Such commercial interests are almost always far removed from the community who provided the “data” for this extractive process.
Mozilla Common Voice[16] is a useful starting case because it was designed in good faith. Crowdsourced speech contributions across many languages, including several African ones, were released under CC0 to maximise reuse for ASR research. The licence accomplished that goal. It also let value generated from those contributions flow in directions the original speakers do not control, with no mechanism of return to the communities that produced the recordings. Because CC0 waives attribution, provenance and authorship are not carried with the data as it circulates, and a contributor has no standing from which to raise a later claim.
The Mozilla Data Collective[17] attempts to correct this problem, from within the same organisation. Rather than collapsing all contributors into a single licensing pool, it lets data creators retain control and licensing of their materials on their own terms. The framework only matters insofar as the licences available are themselves designed for the asymmetries it addresses. Several of the datasets currently published on the platform are licensed under the Nwulite Obodo Open Data Licence 1.0[18] (NOODL-1.0), drafted in 2024 by the Data Science Law Lab at the University of Pretoria, Data Science for Social Impact, and the Centre for Intellectual Property and Information Technology Law at Strathmore University.
The name of this framework comes from the Igbo phrase for raising or reviving the community. To respect and revive what communities contribute, the licence draws a tripartite distinction between dataset providers, recipients in developing countries, and recipients elsewhere. Users in Africa and other developing nations may use, modify, and redistribute the data freely, subject to a share-alike obligation to keep derivatives within the same regime. Users from wealthier regions face the same share-alike obligation, plus a requirement to provide royalties or other tangible benefits to the original dataset owners. The licence is open in the sense that it imposes no permission threshold on the actors most likely to be excluded by paywalls, and it is equitable in the sense that it conditions the use of the data on a return to the people who produced it. Whether the architecture is enforceable depends on jurisdictions, on funders willing to write it into grant requirements, and on the willingness of large training operators to comply with terms they have not historically respected. NOODL is not yet a settled instrument, but it is setting the ethical and political tone for future developments.
NOODL’s logic of collective benefit and community authority converges with principles articulated in the global Indigenous data sovereignty movement: the CARE Principles for Indigenous Data Governance[19] – Collective Benefit, Authority to Control, Responsibility, and Ethics – were formulated to sit alongside the data-centric FAIR (Findable, Accessible, Interoperable, and Reusable) principles (Wilkinson et al. 2016) and make governance answerable to the peoples from whom data originates (Carroll et al. 2020, 2021). The two address the same problem from different continents, yet have developed largely in parallel; bringing the African data commons into explicit dialogue with Indigenous data sovereignty work elsewhere would strengthen both.
Adjacent practices point in the same direction. The South African Centre for Digital Language Resources (SADiLaR)[20] hosts a national repository in which contributing institutions retain copyright and set their own access restrictions, with persistent identifiers and FAIR alignment supplied by the infrastructure rather than as a condition of deposit. This arrangement keeps governance close to the data owners while still producing a discoverable resource at (inter)national scale. Licensing is one of the places where a project’s account of who and what it is for, and to whom it is accountable, is written down. Treating it as boilerplate hands the question to a licensing regime that was never designed for the asymmetries described here.
2.5 Community participation
Licensing dynamics bear directly on community participation, which is widely treated as the way to avoid extractive data practices; “community-built” datasets are often presented as automatically more ethical than scraped ones. Yet, communities can be conscripted as much as consulted, and the difference between participation and crowdsourced extraction lies as much in who controls the pipeline downstream of contribution as in how the contribution was solicited.
Three projects help mark this distinction. The Yorùbá Names project is held by the Yorùbá Names Lexicography Initiative, a non-profit registered in Nigeria. The dictionary admits user input for new names and for vetting existing ones, but the vetting itself is performed by native-language experts, and the project’s open-source code base means that other language communities can replicate the structure on their own terms. Custodianship is deliberately diffuse: the project does not “belong” to any single individual or institution, and decisions about what counts as a name are negotiated rather than hierarchically decided.
The N|uu dictionary occupies a different place on the same spectrum. The orthography itself was co-developed with N|uu speakers and community members, who agreed on the conventions used to render click consonants and tonal contrasts that no prior writing system fully captured. Audio recordings are held by speakers as well as by linguists, and the dictionary, app, and online portal all foreground community use as much as academic research. N|uu is critically endangered, which might at first seem to justify a scholarly intervention for its protection. But an intervention that left the community’s own access to the resulting record mediated entirely by external researchers would not, in our view, meet the basic ethical requirements of a language-protection project. Any language heritage resource must be for the speakers of the language it aims to protect.
The CLAD leads projects on Wolof, Pulaar, and Séeréer that illustrate an even more complicated middle ground between community usefulness and academic methods. The corpora were built through partnerships with banks, hospitals, and other public and private institutions whose multilingual communication needs created a demand for a resource. Speakers gave informed consent, and personal identifiers were removed before publication. This has produced datasets that are domain-rich, including terminology on the banking, medical, and IT sectors alongside everyday conversation. The lines between “community” and “domain” are hard to draw, and these concepts are negotiated, rather than given. In this case it is commercial institutions that decide which Wolof, Pulaar, or Séeréer matters and for which purpose, and partnerships of this kind, however carefully consent is handled, are not the same as community ownership of a resource.
Three patterns stand out. Custodianship without single ownership, as in the Yorùbá Names initiative, lowers the chance that a dataset will be repurposed or sold without consultation, but it also shifts the burden of governance onto a non-profit whose long-term funding is uncertain. Co-development of orthography and infrastructure, as in the N|uu case, ties the resource tightly to the community but is hard to scale beyond a single language. Commercial partnership, as in the CLAD work, can produce data of unusual depth but gives control to entities whose own interests in the resource are not the same as those of speakers themselves. None of the three is the correct model for every project. What matters is understanding the purpose of a given project and then building community participation into its structure, rather than treating it as an opening step that, once completed, can be set aside.
2.6 Sustainability of language infrastructure
A linguistic corpus on its own is rarely useful. Its value emerges in circulation and use, and this value is largely determined by its full bundle of tooling, ontology, documentation, licence, and community of users. SADiLaR is a national research infrastructure that addresses the sustainability of linguistic corpora and projects. It runs three programmes (on digitisation, DH training, and higher-education sector support) through which it builds and hosts the resources it gathers and trains users in them. Each item in the repository receives a persistent handle and is aligned with the FAIR principles, with as much CARE alignment as the underlying data permits. Depositors retain copyright and set their own access conditions, and the infrastructure provides discoverability, preservation, and training in how to use what is preserved. SADiLaR’s model depends on sustained South African public investment that has no equivalent elsewhere on the continent.
The Digitizing-Endemann Dictionary project shows how much can remain unresolved even when a repository is within reach. The team has produced a careful, plural, multi-orthography Sotho dictionary, but the data sit in pre-database tables awaiting design decisions about platform and host. SADiLaR can host the collection and is in conversation with the team about doing so; preserving the data, however, is a step short of building the tool that makes it usable. A searchable dictionary portal requires development capacity that is not yet in place, and no permanent funding stream covers the move from prepared data to a maintained, accessible resource. For now, that step depends on the project leaders and on whoever can fund the development.
The ALMEDA project points to a third possibility. Because it stores only metadata rather than full texts, the computational footprint required to keep the resource alive is small enough that archival sustainability is not the binding constraint. What the project depends on instead is a long-term editorial board, a multilingual ontology that can be extended modularly as new language experts join, and a Wikibase backbone. Resilience requires designing each component so that loss of any one does not strand the others, and being realistic at the proposal stage about what can be sustained on the resources at hand.
2.7 What language work asks of the field
A few points recur across these cases. The term “low-resource” should be applied to the description of pipelines, not to languages themselves. This entails critically questioning what current methods, ontologies, and platforms obscure and what would have to change in them to make existing data legible.
The textual default of NLP is a methodological choice, and projects that begin from speech, audio, or multimodal data address a real ontological gap. Existing standards are neither neutral nor fixed, and need to be actively engaged and questioned. Licences, too, are instruments that must be understood and analysed in context: licences designed to ensure cultural and pecuniary return to communities (such as NOODL-1.0 and the broader African data commons of which it is part) should sit alongside CC0 and CC-BY and be employed whenever appropriate to the context. This expands the field’s vocabulary for what openness can mean for those for whom it is valuable. Community participation, finally, is a structural property of a project rather than a preliminary phase of one, and it does its work only when speakers retain influence over what is done with the resource downstream of contribution.
The tension we did not resolve is the one with which we began. Findability and interoperability require shared standards; fidelity to linguistic variation requires resisting them. The field would benefit from standards that admit their own gaps and accommodate plural forms, schemas that document their own exclusions, and infrastructures that hold variation without flattening it. Whether the field can build at this register, against the pull of benchmarks and pipelines that reward the opposite, is one of the questions the workshop leaves open.
3 African (Digital) Archives in the Digital Realm
The archives working group examined African digital archives through projects run by several of us, which apply computational methods to document and engage with Islamic manuscripts, women’s oral histories, trade union records, visual arts documentation, and West African newspaper collections. They are a selection of what the group works on. We report what has worked in them and what has not, and draw recommendations from both. Our aim is a guide to intentional practice rather than a set of best practices. We returned throughout to how digital archives can serve African communities instead of extracting heritage for external use, and to how that aim plays out at each decision point: what gets digitised, how it is catalogued, how it is presented, and how it is engaged with, whether by human users or by AI systems.
No single answer emerged. How digital archives serve African communities depends on a cascade of decisions about ownership, access, sustainability, and labour, each shaped by the specific communities, materials, and institutional contexts involved. The subsections that follow take these dimensions in turn, with examples from our own projects.
3.1 Intended audience and aims
“For whom?” and “for what purpose?” emerged as foundational questions in our discussions. Not all DH work begins with digitisation. Some projects engage born-digital materials, and others apply computational methods to sources digitised elsewhere. Where digitisation does occur, however, it often conditions what follows. It shapes whether and how researchers can draw on computational methods such as text mining, spatial analysis, and network visualisation — techniques whose value depends on critical reflection about the assumptions they embed. It also shapes how projects come to encounter AI applications, whether by deliberate choice or through unintended exposure, as when open-access materials are ingested into commercial training sets without the knowledge or consent of producers and custodians.
Before one can choose what computational methods or AI applications to use, or an LLM can train on digitised content, someone must decide for whom, for what purpose, what to digitise, and how. That someone may be the researcher in question or, as is often the case, an earlier institution, funder, or project whose selection choices condition what later work can do. These upstream choices continue and exacerbate pre-digital processes through which power shapes what becomes visible and what remains invisible (Trouillot 1995). Digitisation is not a neutral technical process: as Zaagsma has argued, “choices and decisions about selection for digitization, how to catalogue, classify, and what metadata to add are all political in nature and have political consequences” (2023, 831). Chamelot et al. (2020) have shown how this process plays out specifically in the African context, where digitisation has reconfigured power dynamics and governance, situated at the crossroads of political and economic interests. Decisions about what to digitise are often driven by funding priorities and the research interests of individual scholars rather than by the communities whose heritage is involved. For African Studies, these decisions are further shaped by the fact that funders and researchers are usually based outside the continent.
This external/internal divide, however, oversimplifies the dynamic: outside actors are frequently informed by local interlocutors, and selection reflects both translocal and local hegemonies. Groups already marginalised in non-digital archives, such as women authors in Saharan manuscript collections (Frede 2025), women activists in anticolonial struggles, or enslaved people in post-slavery societies, risk having their exclusion reproduced and entrenched in the digital archive. Such decisions carry consequences for whose histories become findable and whose remain in the shadows.
Every subsequent decision about ownership, access, sustainability, and labour depends on the answer to this question, and it must be asked at the outset. The answer determines language choices, metadata systems, interface design, choices about what gets digitised, and, in some cases, whether a project should exist at all. A collection oriented towards broad public engagement in Europe and North America prioritises discoverability and accessibility; a community heritage project in rural Mali prioritises family consent and local custodianship.
The Mapping Senufo project[21] illustrates the complexity of these choices. It is a born-digital publication that set out to map geographic locations linked to objects classified as Senufo across present-day Burkina Faso, Côte d’Ivoire, and Mali. Aimed at art enthusiasts and curious publics, it invites broader audiences to think about the nature of evidence and knowledge production through a single African art style and its related histories, and it acknowledges the partiality inherent in any framing of the material (Gagliardi and Petridis 2021; Gagliardi 2022). The project confronts gaps in knowledge without insisting on filling them, and some gaps may not be appropriate to fill at all.
No digital collection can serve all audiences at once, and the attempt to do so often means serving none of them well by simply defaulting to Anglo-American internet norms. Risam (2018) has called for resisting the universalising tendencies of digital technologies and instead centring local, situated interventions. This held across our cases: clarity about intended audiences and aims should guide all decisions that follow, and failing to specify an audience defers the question until it re-emerges as a problem.
Intended audiences and aims may themselves be multiple and in tension with one another, but unintended audiences and consequences are a distinct matter: they typically emerge after the fact and often outside the team’s control. A collection that makes West African newspapers globally accessible in open access also makes them available for commercial AI training without the consent of the original publishers. A project that preserves manuscripts in a foreign repository protects them from conflict but may alienate the communities that produced them. Good practice here demands more than clarity about such conflicts; it requires attending to the specifics of each project and the relationships within it, recognising that the ethical weight of a decision depends on who is affected and how, and treating negotiation as ongoing rather than settled by a single agreement. Necessary trade-offs remain, particularly between ownership and access.
3.2 Ownership: who has the authority to grant rights?
Ownership in the context of digital archives is rarely singular. The physical object, its digital copy, the metadata, any analytical outputs derived from it, and, in the case of oral archives, the chain of transmission itself, may each have a different “owner”. Families, communities, religious figures, state institutions, and funding bodies may all hold legitimate claims over the same material. Identifying the relevant stakeholders is a necessary first step, but it is often a harder task than it appears, particularly for historical documents whose provenance is uncertain or disputed, or which were produced by organisations that no longer exist.
Our discussions brought out instructive contrasts. The project Archives des Femmes du Mali,[22] initiated in 2016, is building an open-access digital archive of thousands of documents belonging to Malian women activists engaged in political and social struggles from the 1950s to the present (Thiam et al. 2023). The project was born from the recognition that women’s contributions to Mali’s history were largely absent from existing archives. Ownership is vested firmly in the families who produced the documents: the women are the “productrices d’archives”, and the project is built around their right to control their own documents. Two years of trust-building preceded the collection of any material. The project team requires family members to be present during digitisation. The team also hires a family member to handle the physical placement of documents on the scanning equipment, so that nothing is taken without direct oversight. Each family also designates a specific heir as future custodian, so that custody does not rest on one person. Since May 2013, the Malian government has required formal approval from multiple ministries to ensure that personal data remains protected, and the project now operates under a legal framework in which failing to obtain consent carries criminal penalties. In this case FAIR and CARE reinforce each other: the families’ authority over their own documents underwrites open access rather than limiting it.
By contrast, the Mineworkers’ Union of Zambia Archives[23] was digitised from 2018 to 2021 under agreements that gave the union full ownership of the material (Money 2021). Researchers access it with permission at a host institution in Amsterdam, which funded the project, and the union retains both the physical originals and a digital copy. Until recently, the union had few practical means of engaging with its digitised archive beyond reading and retrieving documents. Some union officials have suggested uploading the entire archive into an LLM, largely out of curiosity about what the model would do with it. Such use may incorporate the digitised material into the model’s training data. This makes ownership of that digitised material unclear and was not a possibility considered when the original agreement was reached, yet is technically within the union’s rights.
The Timbuktu manuscripts — around 300,000 handwritten books and documents dating from the thirteenth to the twentieth century, distributed among dozens of private libraries as well as the state-owned Ahmad Baba Institute — illustrate how competing claims compound over time. In a cooperation project (1999–2008) to preserve, digitise, and catalogue the Ahmad Baba collection, the National Library of Norway offered secure deep-mountain storage for digital backups at its cultural heritage facility in Mo i Rana.[24] The Malian side refused, fearing loss of control. Then in 2012 a major conflict erupted: jihadists occupied Timbuktu, manuscripts were smuggled to Bamako, and some of the digitised copies and catalogues from the early 2000s were lost, prompting several partly overlapping projects to redigitise and recatalogue the collections. Meanwhile, the Malian state invokes “national heritage” to claim authority over collections that Timbuktu communities regard as their own.
The Ethiopian case echoes this same dynamic through a different route. The Information Network Security Administration, the national body for digital information infrastructure, hosts digitised manuscripts on a centralised cloud under paid contracts and controls access to them. Foreign researchers face restrictions based on religion and perceived intentions. Ethiopian researchers inside the country may have less access than European researchers consulting copies abroad. The reversal shows how centralised digital infrastructure can cut against the communities an archive is meant to serve.
Community custody is not always the answer, either. Communities can be internally divided, and the notion of “community” is bound up with questions of power: who counts as a member, who represents the group, and who claims authority to speak for it are rarely settled questions. Communities can operate at several scales at once: a single family, a city, a nation, the global scholarly community. Our group discussed thought experiments showing that even starting and ending with full community involvement does not guarantee ethical outcomes: local hierarchies, generational conflicts, political pressures, and religious tensions (such as reluctance to preserve material associated with rival practices) all shape the approach that any “community” might pursue, and different “community” members might prioritise different commitments. Communities, moreover, change over time. For older material, the community that exists today may be only loosely related to the one that produced it, raising further questions about who speaks for heritage whose originators are no longer present.
3.3 Rights and agreements: what formal instruments cannot settle
Copyright is narrower than the overlapping claims set out above suggest. The World Intellectual Property Organization defines it as “the rights that creators have over their literary and artistic works”.[25] It attaches in the first instance to the creator, runs for a term that varies by jurisdiction, and does not pass to an owner, custodian, or institution by virtue of their holding the object. The claims described here fall outside it: those over the physical object, over communal or intergenerational custody, over circulation on cultural or spiritual grounds. They can still prove more restrictive in practice than copyright would be, and for manuscripts several centuries old, in which no copyright subsists, they are the only claims in play.
Yet formal agreements, whatever rights they allocate, remain useful instruments for protecting owners, co-producers, and hosting institutions when external actors attempt to claim material after the fact. Several participants reported experiences in which families or government authorities approached them after the work was completed, accusing them of “stealing” documents. What matters is that such agreements are concluded at the outset and between all the relevant parties: producers and collectors, collectors and their employees, hosting institutions and collectors, funders and collectors. They also ensure that the voices of the archives’ creators are acknowledged and preserved, and that researchers working with the material can contact the primary rights-holders for additional context. This documentation keeps the archives grounded in human connection rather than reducing them to impersonal data.
Agreements of this kind reach their limit at AI. Even when digitisation agreements include forward-looking clauses covering future technologies, the rapidly changing nature and applications of AI mean that the organisations, communities, and researchers who signed them are unlikely to have anticipated all practical implications. Consenting to digitisation is not the same as consenting to having one’s materials used for training LLMs. There are end users who have received digital copies and are already feeding material into LLMs without re-consulting the original custodians, and there is little that the digitising researchers can do about this. Open-access collections, meanwhile, are routinely crawled by commercial AI systems.
This situation raises the question of whether formal agreements alone can bear the weight of ethical obligation. Agreements are better treated as living documents, revisited as technologies and circumstances change. Communities may also change their minds generationally: a family that consented to digitisation may, a decade later, want materials removed from public access. Sustainability planning must account for this possibility, and decisions taken about the material must reflect the wishes and intentions of the owners.
A further limit surfaces after digitisation is complete. Digitised material may contain information that its owners were not aware of: exploitative relationships, power dynamics, sensitive personal details that become visible only when documents are aggregated, transcribed, or cross-referenced. This is especially the case for large archives developed over several decades. Does the researcher have the prerogative — or the obligation — to “unmask” what the archive reveals and to inform its owners? Our discussions produced no single answer. Some participants argued for the researcher’s analytical independence; others stressed that owners must retain control over what is disclosed. The important thing, we concluded, is to confront the question rather than ignore it, and to recognise that the answer may differ from one context to the next.
What can be generalised is the obligation to document the choice. Behind every digital archive lies a series of choices: why this material and not that, which standards, whose classification, what was excluded and why. Making this labour visible is itself an ethical practice. Money (2021) published a candid account of the selection and organisation choices behind the Mineworkers’ Union of Zambia digitisation, including an assessment of how it might reinforce existing historiographical biases. This kind of reflexive transparency was endorsed by others in the group as a model worth emulating. It does not eliminate the power asymmetries embedded in archival work, but it allows others to identify, assess, critique, and build on the choices that were made. Transparency has limits, however: a record’s apparent neutrality can itself “generate opacity” and carry colonial categories unexamined into digital systems, so digitisation is not decolonisation unless the classifications themselves are opened (Gibson 2024).
3.4 Access: making material findable, usable, and equitable
Digitisation makes material accessible to researchers who cannot travel. It saves time, overcomes visa-related travel barriers, and, when archives are endangered by conflict, provides a form of protection that may prove decisive. It also reduces the need to handle fragile originals, protecting both researchers and the objects themselves, since even careful use accelerates wear. Digital copies of Ethiopian archives, held abroad, survived the war in northern Ethiopia (2020–22), while the physical originals in the country were damaged or destroyed. Sudanese archival collections, feared to have been damaged or destroyed in the war that began in 2023, illustrate the cost of not having copies elsewhere. Yet digitisation does not preserve the original: termites, water damage, dust, and neglect erode physical objects whether or not they have been scanned, and a digital surrogate cannot fully reproduce the material evidence that an original carries.
Nor is online access universally beneficial. Paywalls, language barriers, lack of digital literacy, and poor internet infrastructure all limit who benefits. In Ethiopia, students in rural areas lack the connectivity and institutional support to access digitised collections that are freely available in principle. In Timbuktu, the Ahmad Baba Institute kept its digitised manuscripts accessible only on site, concerned that online availability would reduce physical visits to heritage centres and the income stream tied to them. Cairo’s al-Azhar Library adopted a similar policy for its large digitised manuscript collection. Any digitisation strategy must reckon with the possibility that making material globally available may undermine the local institutions and livelihoods that had previously sustained the material. One middle ground the group discussed is to restrict public access to metadata and keep the digitised material available only to authorised researchers, possibly for a fee scaled to the researcher’s means, with proceeds going to the local archive. Discoverability need not require full exposure.
Discoverability, though, is itself a product of cataloguing choices, and those are never neutral: they determine who can find what. Archives catalogued using an organisation’s own terminology — “Copperbelt Industrial Service Bureau”, for instance — are meaningful to insiders but opaque to outsiders; externally imposed categories, in turn, make material findable but strip local meaning. As Loukissas (2019) has argued, all data are shaped by the local conditions of their production. One response is to refuse standardisation altogether: in 2012, Burkina Faso’s national geographical names commission deliberately declined to standardise place names, because non-standardisation allowed multiple communities to maintain competing claims (Gagliardi 2023).
Yet standardisation and plurality are not strictly opposed, and digital systems can accommodate multiple variants and co-existing terms for the same entity more easily than paper catalogues ever could. The Ethiopian case points to challenges such systems would need to absorb: localised terminologies (“Sheikhota”, a respectful plural for a single scholar) and orthographic diversity within single languages (Ajami variants, Oromo dialects) that make common keywords nearly impossible to establish. AI-powered search and classification tools compound the difficulty by absorbing standardisation biases, defaulting to dominant forms, and reproducing inherited classification frameworks with less transparency than human-made systems.
The practical conclusion is not that standards should be abandoned, since findability requires some shared vocabulary, but that their limits should be acknowledged and their assumptions documented. Good practice means designing for the audiences a collection is built to serve while remaining open to revision as those audiences change, accommodating variant forms where the technology allows, and being transparent about what is lost.
AI bears on each of these choices: what is catalogued, in whose terms, and who can reach it. It offers both genuine possibilities and risks. Several applications proved productive: LLM-assisted triage of large, unprocessed collections (“getting things out of the drawer”, as one participant put it); optical character recognition (OCR) and handwritten text recognition (HTR) for materials with complex layouts, multiple fonts, or handwriting; automatic entity extraction to generate or enrich metadata; and multilingual content summaries to make vast collections navigable. Work by one member of our group on the Hofheinz Collection — a corpus of over 600 Arabic manuscript pages from eighteenth- to twentieth-century Sudan, part of the University of Bergen’s Sudan Collection (Khalīfa et al. 2025) — illustrates the pragmatic case for this approach. Even imperfect HTR output can be used for AI-powered content analysis and summarisation, which cuts the time needed to identify materials that warrant in-depth textual scholarship. The goal is effective triage: scarce human expertise then goes to the documents that most require it.
The experience of another project in our group illustrates both the promise and the limits of these applications. Work on the Islam West Africa Collection (IWAC)[26] has shown that multimodal LLMs can markedly improve OCR for newspapers whose complex multi-column layouts frustrate conventional OCR systems, at a fraction of the cost and time of manual processing, while also automating named entity recognition and metadata enrichment across a collection of over 17,500 documents. Yet clean output can conceal fabrication, over-normalisation, or silent modernisation of spellings, unlike traditional OCR, which signals failure through obvious garbling. The shift from traditional OCR to LLM-based processing trades visible failure for invisible failure, a distinction with significant consequences for verification and trust.
For metadata generation, LLMs can produce semantically rich suggestions, associating place names with countries, inferring historical periods from dates, and integrating external knowledge beyond simple keyword extraction. But they also hallucinate: fabricating publication details, inventing authors, and producing confident output that is factually wrong. Detecting such errors is increasingly difficult, since LLM output now mimics the phrasing, sourcing, and expressions of uncertainty of a human cataloguer closely enough that verification typically requires a domain expert. Error rates that may be acceptable for discoverability keywords become problematic for authoritative catalogue records. Even so, AI-assisted cataloguing with human verification can require less time than fully manual processing. The challenge is to build verification workflows into projects from the outset.
A deeper tension emerged in our discussions between the aspiration to carefully design an ontology or metadata system and the speed of LLM-based metadata attribution, which can process at scale but with less control over the categories imposed. Conventional database structures tend to rely on rigid, hierarchical taxonomies rooted in Western academic abstractions, which do not always reflect the fluid and dynamic character of the knowledge systems they attempt to describe (Eisenhuth et al. 2023). The Islamic Cultural Archive (ICA),[27] a research platform linking scholars in Germany, Kenya, Mauritania, Senegal, and Tanzania, has shown that a promising approach involves building ontologies collaboratively and iteratively, keeping classification structures open and extensible rather than fixing categories at the outset. This kind of careful, participatory ontology design takes time, but it produces systems that grow with the research rather than constrain it.
We did not reach a consensus on how to balance human-designed ontologies with AI-generated metadata. What we agreed on was that AI-generated summaries and metadata are a first step that makes collections navigable enough for human engagement to begin, while leaving substantive interpretation to human study.
Most digital collections default to English. Multilingual interfaces and AI-generated abstracts in several languages widen who can use a collection, and cross-lingual databases, though technically demanding and resource-intensive, let a researcher working in one language find material catalogued in another. The ICA’s ontology operates across Arabic, English, and French, so that data catalogued in one language can be found and linked from the other two (Eisenhuth et al. 2023). Datasets that a monolingual system would keep siloed become searchable together.
3.5 Sustainability: beyond project-based funding
Digital projects are structurally fragile. German university libraries, for instance, generally offer ten years of storage, which is an obvious limitation for projects that aim for long-term preservation. Server contracts expire and hosting infrastructure strains over time. The Copenhagen-hosted database of the IslHornAfr Project,[28] which documents Islamic manuscript traditions of the Horn of Africa, has outgrown its original server capacity; the host has repeatedly requested upgrades or migration, placing the digitised materials at risk and leaving the principal investigator to negotiate relocation to secure their long-term preservation. National libraries may store metadata but not the data itself. As one participant observed, the perceived duplication of digitisation efforts, meaning the same collections digitised multiple times by different projects, is itself a symptom of this structural failure: each wave of funding creates a new collection, but none sustains what came before. Not all stakeholders share the same investment in long-term sustainability. Funding agencies tend to prioritise new digitisation over the maintenance of existing collections, and institutional incentives reward creation over stewardship. The result is a structural misalignment: the people and entities who build digital collections are rarely the same actors responsible for keeping them alive. Digitisation does not equal preservation. Without sustained investment in storage, backup, institutional commitment, and ongoing format migration as hardware and software grow obsolete, digital copies are no more durable than physical ones, and sometimes less so.
Fragility is personal as well as infrastructural. Many DH projects are built and maintained by individual researchers. The researcher’s contract, fixed-term or permanent, makes little difference: when they move, retire, or die, the project’s precarity becomes evident. Tenure and institutional standing do not fully resolve this: even when a project’s significance is widely recognised, host institutions must weigh continued investment in its upkeep against support for current and future faculty. A researcher who moves may migrate the project, often to an institution with no obligation to support it. The IWAC illustrates this kind of trajectory. What began as a research database at the University of Florida was transferred to the Leibniz-Zentrum Moderner Orient in Berlin, and the researcher has since moved again to a third institution. Each transition depends on institutional goodwill rather than structural support, and on someone willing to advocate for the project and absorb the work of maintenance. A model that rests on individual persistence and institutional negotiation is neither scalable nor fair.
One common answer to both kinds of fragility is redundancy. Storing copies in multiple locations, including outside the country of origin, may appear as a rational preservation strategy: distributed storage reduces the risk of total loss from conflict, infrastructure failure, or neglect.[29] As the Ethiopian and Sudanese cases discussed above illustrate, geographic redundancy can be decisive. It is not always, however: redundancy protects only if all parties retain access to the copies they rely on, and a sudden political shift can leave an originating country cut off from its own material held abroad. Geographic dispersal may conflict with ownership and community sovereignty. Owners may want material to remain under their control, and the political symbolism of storing heritage “abroad” carries political symbolism: as the Timbuktu case showed, even well-intentioned offers of foreign storage can be experienced as a form of dispossession or digital imperialism (Breckenridge 2014). Non-commercial platforms (Zenodo, Dataverse) and minimal computing approaches offer partial solutions that reduce dependency on expensive infrastructure and commercial providers. But no technical arrangement resolves the underlying political questions of who controls access to heritage.
Those questions are settled, in practice, by trust rather than by contract. Most digitisation projects depend on personal trust between individual researchers and specific people within the organisations, groups, and communities whose materials are involved. As the ownership discussion made clear, none of these are homogeneous entities; formal agreements with institutions often rest on a single person trusting a single researcher, whether a family elder, a union official, or a library director. That foundation creates acute sustainability risks: when the researcher is no longer available, the trust the work relied on goes with them. Nor are institutional frameworks sufficient on their own: institutions can also change priorities, lose funding, or even disappear. Hybrid models are needed: institutional scaffolding that supports but does not replace personal trust, and team-building across career stages and institutions so that no project depends on a single individual.
Building such models requires investment in relationships, succession planning, and institutional cultures that value continuity as much as innovation. “Sustainability” here does not mean perpetuity. Many projects will not and need not run forever. What matters is that resilience to life events is built in from the start, and that publications, datasets, and reflective accounts of process are recognised as outputs that can carry a project’s knowledge forward even when the digital platform itself cannot be maintained.
What gets sustained is also uneven by format. Textual material is over-emphasised at the expense of audio recordings, visual heritage, and born-digital documents, all of which are also at risk, and some of which are disappearing faster than manuscripts. Born-digital documents pose particular challenges because they are rarely treated as archival material until they have disappeared.
Not everything is meant to be preserved. Some knowledge systems assume ephemerality: materials exist in specific contexts and are meant to come into being, circulate, transform, or disappear (Brett-Smith 2001). The Western impulse to archive everything in perpetuity should not be imposed uncritically. A researcher’s desire for completeness or preservation should not override a clear commitment to ephemerality. But communities rarely speak with one voice on such questions: members may disagree among themselves, and a researcher who is themselves part of the community may find their own judgement at odds with those who claim authority within it. Such differences offer opportunities for discussion and reconsideration of what it means to know and preserve.
Preservation also raises a question projects rarely answer: what happens to the physical collection once it has been digitised? Some projects preserve the physical original alongside the digital copy; others damage or alter it in the act of scanning. Does digitisation reduce the perceived value of the original, or does it generate new interest and new claims? Even when both versions are preserved, they are not interchangeable: engaging with physical materials is an embodied practice involving touch, sound, and presence that no digital surrogate can reproduce. And what is the physical archive’s fate when funding for the digital project ends? These questions arise in every project represented in our group, and they do not have generic answers.
3.6 Labour: the invisible work of digitisation
Creating digitised material and metadata is often time-consuming and tedious, and scholars readily overlook the intellectual dimensions of this work. The working conditions of those engaged in scanning, cataloguing, transcribing, and verifying deserve explicit attention, and their contribution should be recognised as intellectual as well as manual, both in project planning and in acknowledgement once the material is digitised.
Project teams must make choices about employment: direct employees or contractors, formal or informal arrangements, each with trade-offs. Formal employment provides labour rights such as sick pay, parental leave, and protection against dismissal but may require transferring a significant portion of the budget to a partner organisation or navigating unfamiliar regulatory frameworks. In some contexts, workers do not have bank accounts, creating basic logistical problems for payment. More broadly, transferring funds from European or North American institutions to African partners runs into currency regulations, banking infrastructure gaps, and administrative procedures that are rarely accounted for in project budgets, timelines, or staffing.
Project teams need knowledge of national labour law, and assumptions about employment practices that hold in one country may not apply in another. A related question is whether the archive producers themselves — the families, communities, or organisations whose materials are being digitised — receive any compensation, or whether the benefits of digitisation flow primarily to researchers and institutions.
Training is one benefit that projects can share directly. Digitisation projects can create opportunities for people to build skills and expand capacity for future work, but only if training is intentional rather than incidental. Training has to run both ways: technical instruction for those new to digital tools, and grounding in the physical, linguistic, and cultural dimensions of the material for those who arrive with the technical skills but without it. Understanding the material and cultural dimensions of a collection is as essential to a responsible digitisation project as understanding the scanning equipment.
Our group and other workshop participants raised the risk of “recolonising through technology”. When digital literacy gaps mean that European researchers end up doing the technical work, building the taxonomy, configuring the database, and running the AI tools, while African researchers carry out the manual work of scanning and data entry, asymmetries of knowledge and power are reinforced rather than dismantled. The ICA project experienced a version of this: despite the intention to collaboratively develop a cross-lingual ontology and tagging system, uneven digital literacy and insufficient funding meant that much of the technical design work fell to the better-resourced European partners. Addressing this means budgeting skills development as a central objective of the project from the start.
Who receives that training depends on who is hired, and the selection of project staff involves power dynamics that are rarely made explicit. Who decides whom to hire, based on what criteria, and with whose interests in mind? Interdisciplinary, intergenerational, and transnational and transcontinental teams are widely endorsed as good practice, but they are also difficult to manage across different institutional cultures, expectations, working norms, and funding constraints. Making these decisions transparent — like making archival choices transparent — is part of the ethical labour of running a digitisation project.
4 Infrastructure, Governance & Access
The spread of digital technologies and AI across Africa has deepened long-standing debates about relevance, usability, and sustainability in digital design (see, for example, Gikunda 2023). Many DH and AI systems are still built on assumptions of neutral, uniform users, stable infrastructures, and linear governance pathways (Caruso and Spadaro 2024; Chun and Elkins 2023). These assumptions are often embedded in database structures, interface designs, algorithmic benchmarks, and policy frameworks, where they carry a universal epistemology that has shaped both computing and the humanities for decades. That epistemology generalises from a narrow range of Euro-American infrastructures, institutions, and users, then treats the result as the default, so that everything else looks like deviation. Everyday practice across much of the continent departs from that default: people navigate multiple languages and dialects, work around intermittent connectivity and power, operate within regulatory regimes that are still being built, and rely on informal networks and human intermediaries such as agents, librarians, and community stewards to reach and make sense of digital technologies. Intermediation of this kind is universal. Help desks, IT departments, and knowledgeable relatives do the same work in the settings that produced the default, where institutions absorb the labour and it goes unrecorded. African cases make that mediation visible, along with the linguistic plurality and infrastructural improvisation that surround it, and on a scale no design can treat as an edge case. The model that passes for universal is then one particular case among others, with no better claim to neutrality. We need frameworks answerable to how technology is actually used, adapted, and sustained, on the continent and beyond it.
The case for situated, locally responsive design has moved beyond the academy: it is increasingly written into continental and international policy. At the continental scale, the African Union’s Continental Artificial Intelligence Strategy calls for locally produced and controlled datasets, NLP for African languages, and the preservation of African cultural heritage for the benefit of African communities (African Union 2024, 29, 48, 55). Internationally, most African states have endorsed the UNESCO Recommendation on the Ethics of Artificial Intelligence (UNESCO 2023), and at least twenty-eight have begun applying UNESCO’s Readiness Assessment Methodology, with these states pressing for “an AI adoption that is not just ethical, but also contextualised, localised and culturally relevant to spur a transformation integrating sociocultural values as well as linguistic diversity” (UNESCO 2025). These instruments align closely with the priorities this paper sets out. Whether they shape design or remain declaratory will depend on implementation, resources, and political will, and we urge policymakers, funders, and project leaders to treat them as binding commitments rather than aspirational language.
At the national level, governments and policy institutions are increasingly turning to stakeholder interactions, including surveys, focus groups, and scenario-building exercises, to shape national AI strategies and digital governance frameworks (Kwarkye 2025). While these approaches can surface public perspectives, they often replicate extractive, limited consultation models in which communities contribute insights but have little real influence over design, implementation, or long-term stewardship. These patterns point to a deeper tension. What is the appropriate role of standardisation in societies where heterogeneity is the norm? How can standards protect users and enable interoperability without flattening linguistic, infrastructural, and cultural differences? And what forms of governance (bottom-up, networked, or hybrid) can ensure that DH and AI systems remain accountable to the communities they are meant to serve?
In responding to these questions, this section of the position paper contends that design should be understood as emerging downstream of three interconnected conditions: Access, Governance, and Sovereignty. Each condition is expressed through material infrastructures, linguistic environments, trust relationships, custodial arrangements, and policy capabilities, all of which shape what technologies can achieve and whose interests they serve. We therefore propose the following formula:
(a) Access + (b) Governance + (c) Sovereignty = (d) Design Intention
With this formula, we think about our design intention in a simple and even simplistic way. Design emerges from the conditions people actually live within. To clearly illustrate this point, we draw on the example of a mobile money system that works across both online and offline contexts (M-PESA). This case will show how effective systems arise when design principles align with realities such as linguistic diversity, infrastructural inconsistency, community stewardship, and national policy priorities. African digital environments are distinct design contexts that call for new ways of thinking about knowledge, technological systems, and governance.
In the discussion below, we argue that Access (physical tools, infrastructure, and language) and Governance (trust and accountability) shape what can realistically work, while Sovereignty (over infrastructure, data policy, and funding) shapes who decides the direction of development. Design should come from a careful understanding of these three conditions rather than from a template applied everywhere.
4.1 Access for DH and AI
Access is more than connectivity. We define access as the practical capacity of diverse users to discover, understand, and use DH and AI services reliably and safely under variable infrastructure and linguistic conditions, without coercion or extraction. The definition is deliberately operational: it converts into specifications that design teams can build to and test against, and into criteria by which policymakers, funders, and other stakeholders can assess a project. It also puts the obligation the right way round. Systems should be built to the needs, circumstances, and dignity of the people who use them, rather than requiring those people to adapt to whatever the technology happens to support. Framed this way, the definition shows where access breaks down in practice: at the infrastructural layer (what is technically possible) and at the audience layer (how people engage with these systems).
Infrastructure layer: Solutions designed for an individual with reliable mobile data do not translate to shared or resource-constrained settings. In many communities on the continent, where connectivity is intermittent or expensive, devices are communal, shared among families, neighbours, and co-workers. In such a context, people rely on local networks, offline sharing (for example via Bluetooth or memory cards), and cached content that can be accessed without a live connection. As a result, systems built around constant, real-time cloud access can quickly become unusable. What matters more in these environments is the approach to peer-to-peer distribution that allows information to move gradually and reliably rather than instantly. These constraints do not stop at the handset. The compute and network layers beneath it, from data centres and fibre to landing stations and the power that runs them, decide what can be trained, hosted or served on the continent at all, and they are far more concentrated than mobile access is. Where that capacity is now being added, it rests largely on imported stacks, which repeats at the compute layer the pattern we identify in M-PESA below: infrastructure sited on the continent, control over its components held elsewhere.
The devices themselves constrain what is possible. Many users rely on basic feature phones or low-cost smartphones with limited memory and processing power. Such devices may run older operating systems, may support fewer applications, and struggle with complex interfaces. Beyond specification, repairability also plays a crucial role here. If a device breaks and spare parts are unavailable or even unaffordable locally, it may simply fall out of use. These factors shape design choices in very practical ways by limiting how feature-rich an interface can be, how large a model can run effectively, and how frequently software can be updated without disrupting usability.
Access also depends on power supply. In settings where electricity is unstable, unavailable for long periods, or even costly to access, even a well-designed system can fail if it assumes frequent charging or continuous use. Charging a phone may require travel, waiting time, or additional expenses, which changes how and when people engage with technology. This makes power-aware design essential. Offline-first approaches, where data is stored locally and synchronised only when connectivity and power are available, become a requirement. Similarly, smaller data formats reduce both bandwidth and energy consumption, while energy-efficient processing helps extend usability. Access, then, depends on whether systems can function within the real constraints of everyday life.
Audience layer: Access clearly shows how people actually live, communicate, and build trust. Many people move fluidly between being online and offline depending on cost, connectivity, or daily routines, so systems need to support that reality rather than assume constant access. This might mean simple fallbacks like SMS or USSD, the ability to store information for later use, or syncing data only when a connection becomes available. Trust, too, grows from clarity and accountability. People are more likely to engage when it is clear who created or contributed information, how it can be used and what recourse exists if something goes wrong. Language and engagement also play an equally important role in the audience layer. Supporting users means recognising dialects, mixed-language use, and varying levels of literacy; translation into a dominant language is only the first step. In many cases, speech-based interfaces can open up access where reading or writing may be a barrier. At the same time, meaningful access depends on fair and respectful engagement. People should be able to understand consent in terms that make sense to them, in their own language and have the ability to withdraw that consent at any time. Participation should also lead to tangible benefits for the communities involved. Access should mean systems where people can see value, retain agency, and participate on fair terms.
In practice, an access checklist should include type of connectivity (offline, USSD, SMS), device level, stability of power, language fit (including dialects and speech), literacy mode (text, voice, or icons), visibility of consent and data origin, cost and time demands (fees, data use, travel), and whether communities benefit in return. Treating these as essential requirements is what separates systems that only reach people from those that truly work for them.
4.2 Governance for trust
Trust comes from how systems are governed. People-centred governance goes beyond extractive methods like surveys and focus groups to focus on shared responsibility and clear ways to address problems. In practice, this means co-designing systems with community representatives, conducting participatory evaluations with open issue tracking, and creating advisory boards with real decision-making authority over data policies, feature updates, and system changes. Additionally, complaint and redress systems should be built directly into services with clear steps, timelines, and escalation pathways, rather than hidden behind unclear support channels.
Language also becomes a key part of governance. When public consultations, policy documents and terms of service are available in local languages, more people can take part and feel confident engaging with technology. Simple and clear “model cards” and “data cards”, readable consent receipts, and summaries in local languages are governance tools as much as design features. For instance, translating AI policy drafts or service terms into Twi, Ewe, or Dagbani can expand participation in Ghana beyond English-speaking groups. In the same way, workshops held in national and regional languages can reveal local risks and benefits that technical terms may often overlook.
Since no single group can manage everything alone, networked governance spreads responsibility across different actors, including community organisations, local governments, national regulators, and regional bodies. Community groups can guide content practices and manage consent, while government institutions can set safety rules and enforce them. At the same time, regional bodies can support open flexibility standards that allow systems to connect while still adapting to local needs. In digital archives, this may involve community editors, clear licensing rules, and stewardship boards. In the context of mobile money, this may include coordination among telecommunications companies, regulators, and consumer protection policies, as well as oversight of local agents, to build trust.
Finally, governance needs to move from asking people for inputs to giving them a say in how data is handled. Tools like data trusts allow communities to decide who can use their data, change their minds later, and benefit from how the data is used. We should also work to protect those who make DH and AI systems possible, including data annotators and content moderators. Their contracts should guarantee fair pay, mental health support, limits on exposure to harmful materials, and transparent conditions. These steps are essential to DH and AI systems that are trustworthy and sustainable.
4.3 Sovereignty beyond adoption
Sovereignty means going beyond simply using DH and AI tools designed by others. It is about having the power to shape these systems by setting the rules, deciding what matters, and ensuring the technology serves the community’s needs rather than the other way around.
Infrastructure sovereignty begins with who controls the computing power and storage that digital systems rely on and whether those systems run directly on local devices or in the cloud. It also includes decisions about how the data is kept, moved, or shared; who owns or manages the communication networks and the land they sit on; and how power (electricity) systems are organised and maintained. These constraints are often described as obstacles, but they can actually inspire smarter, more context-aware design. For example, running speech-recognition tools directly on a device can reduce delays and protect privacy in places where internet access is slow, not available, or costly. Likewise, storing data locally and syncing it only when connections are stable can make technology more reliable in areas with frequent power cuts, without necessarily disrupting user experience.
Data sovereignty and stewardship turn on clear community rules about who can access data, how it is tracked, and how it can be reused. This includes decisions about where data is stored, limiting its use to specific purposes, setting different access levels for sensitive information, and allowing users to withdraw consent with proper records in place. In community archives, stewardship groups can decide how data is shared and protect cultural practices. In public services, clear consent records and visible data histories, such as showing how decisions are made and what data is used, may be helpful in understanding and challenging system outcomes.
Policy capacity determines whether governments can act on either. When policymakers and public institutions better understand how DH and AI systems work and how they affect individuals, they can make stronger decisions.
With these insights, governments can set procurement rules that favour open standards, ensure different systems work well together, and place sensible limits on energy use. Policymakers can also build labour protections directly into contracts, ensuring that the people who support these systems are treated fairly. Strong policy capacity also helps prevent governments from becoming overly dependent on a single company or platform by requiring clear plans for data portability and long-term maintenance. Finally, bringing technical experts into public bodies can make a significant difference. Their presence will strengthen oversight, improve accountability, and create quicker and more meaningful feedback loops between real-world experiences and policy decisions.
Economic sovereignty determines whether any of it survives. Long-term systems cannot depend on a single source of funding. Using a mix of public funding, philanthropy, social business models, and membership support reduces reliance on any single actor. Safeguards should limit corporate control, require open designs for shared projects, and make budgets and environmental impact public. In this way, sovereignty shapes which systems can be built, how data is handled, and how funding influences (or does not control) governance.
4.4 An infrastructural analogy: M-PESA
M-PESA (M for mobile, plus pesa, Swahili for money) is a money transfer service created in 2007 in Kenya. It has since expanded well beyond Kenya as the leading mobile money service, and is often hailed as a Kenyan success story. The story is less straightforward: M-PESA was first developed by Vodafone UK. The infrastructure of M-PESA is Kenyan, but the Intellectual Property Rights over its components do not belong to Kenyan companies (Foster 2024). That gap is what makes M-PESA a useful infrastructural analogy for future DH projects.
Access: How does M-PESA achieve inclusive access across different devices, connectivity conditions, and user needs?
M-PESA can be used by everyone with a phone to borrow and repay loans, buy airtime, pay bills, or simply transfer money. Services may differ depending on the country, but it works with SIM-based technology through either the SIM Application Toolkit (STK) or USSD, two user interfaces which are simple to use, cost-effective, and efficient in an environment where the signal can be weak.
Governance: How does M-PESA build trust and ensure user protection through regulation and human support systems?
The Central Bank of Kenya regulates M-PESA, which is underwritten by local commercial banks running the security layers required of payment systems. Customers and banks must satisfy Know Your Customer (KYC) requirements: the identity of M-PESA subscribers is verified through authorised retail agents cleared by the telco. Telcos such as Safaricom train and monitor their agents, and ensure that agents hold enough cash to keep the system trusted.
Sovereignty: How does M-PESA maintain national control and policy flexibility while balancing inclusion, privacy, and interoperability with domestic priorities?
Regulations might vary from one country to another, but the phone companies operating M-PESA follow financial regulations in each country. In countries such as Kenya, M-PESA is not nationalised but is regulated by the Central Bank of Kenya (CBK) which can dictate certain changes. For example, during Covid-19, the CBK asked M-PESA to temporarily waive transaction fees.
Design: Combining these three (access, governance, and sovereignty), what is the design outcome?
M-PESA shows how access constraints, governance agreements, and sovereignty choices translate into design. Weak signal and basic handsets produced STK and USSD interfaces rather than an app; the KYC regime and the retail agent network placed a trained human intermediary at the point of use; and regulation by the CBK, rather than nationalisation, left room for intervention of the kind seen during Covid-19. The result is a service that is easy to use and perceived as reliable. What the case does not settle is ownership: the design answers to Kenyan conditions, while the intellectual property behind its components does not.
4.5 Rethinking standardisation in the African context
Africa’s socio-technical fluidity, including the many languages on the continent, the changing human support systems, uneven access to the internet, and uncertain long-term stability (both political and economic), should be treated as a core starting point when designing DH tools and AI systems. Multilingual habits such as code-switching, dialect variation, and mixed-language use make text-only designs difficult. At the same time, unreliable electricity and fluctuating network strength may disrupt cloud-based systems, while human support roles such as agents, community annotators and moderators often play key roles in how people actually use technology. Systems that begin with these realities in mind tend to fit better into daily life: M-PESA’s USSD systems keep cost, reach, and confirmations clear even on basic phones and weak networks, while community-driven archives that focus on local knowledge and care show how issues like ownership and consent can be built into the system from the start rather than added later on.
In contrast, many DH and AI systems are effectively designed on the working assumption that users are broadly interchangeable, that operating conditions are stable, and that systems can proceed through neat, standardised processes from start to finish (Gikunda 2023). This focus on sameness leads to tools built for an imagined “average” user (someone with steady internet access, who uses a dominant language such as English, and who operates within simple system cycles and top-down control). In more complex and changing environments, such as those in Africa, this creates a clear problem. People are left out at the interface level due to language barriers or literacy gaps; systems fail in practice due to power and connectivity issues; and trust breaks down because communities are treated as sources of data rather than as active decision-makers. Low usage, improvised workarounds, and one-sided engagement all point to a mismatch between design and reality.
This position paper does not argue for rejecting standards; it argues for rethinking the process of standardisation and its purpose. Instead of being fixed and rigid, standards should work as flexible scaffolding that promotes the flow of ideas for innovation, safety, compatibility, and connection, while still allowing change, local adjustments, and the option to move away from them when needed. Standards only become harmful when they prevent change, lock users into particular systems or formats, reinforce stereotypes, or are unavailable locally. The question is then whether to adopt, adapt, or build, judged against local context, the flexibility required, and practical need.
5 Epistemologies, Decoloniality & Ethical Frameworks
We were tasked with interrogating the assumptions and ideologies embedded in Global DH and AI tools, through an African epistemological and decolonial lens. This section, thus, aims to surface key concepts, approaches, and ethical considerations that are rooted in critical African Studies traditions and Africa-based perspectives and experiences. The following questions framed the majority of our exchanges:
a) What key concepts and strategies do African and Africa-based scholarship and practice introduce to DH and AI? What embedded assumptions do they trouble and/or address?
b) How can critical traditions – as well as recent insights and debates emerging from African Studies, which have significantly shaped decolonial and ethical frameworks and advanced plural epistemologies – inform and diversify DH and AI development in and on Africa?
c) How can the emergent praxis of DH and AI development in Africa influence and/or redirect global discourse, practice, and ethics in these fields?
d) Conversely, how might innovations in global DH and AI contribute to reshaping and expanding these critical traditions and praxis in a genuinely reciprocal exchange?
e) How do these questions all come to bear on the ethical imperatives of DH and AI integration into African Studies?
Participants brought expertise in postcolonial DH, digital research infrastructure, digital heritage and restitution, AI ethics, and African historical knowledge systems. Our exchanges opened up further questions of their own, which the sections below take up.
5.1 Mobilising African epistemologies and decolonial critique
African Studies, as an interdisciplinary field focused on the continent and its global context, together with digital interventions in Africa by Africans, has the potential to shape knowledge production in the field of DH and, more broadly, ways of thinking around AI.
Interdisciplinary, decolonial perspectives and epistemic forms from African humanities and social science research push back against the re-inscription of colonial epistemologies and the erasure of African cultural and historical knowledge systems in new technologies, through a range of approaches and with different aims. They do so, for instance, by reasserting African knowledge and ethical frameworks in digitised artefacts, heritage and born-digital objects, and by calling for genuinely inclusive understandings of the world. The stakes of this approach centre on a critical, postcolonial DH that “addresses underexplored questions of power, globalisation, and colonial and neocolonial ideologies that are shaping the digital cultural record in its mediated, material form” (Risam 2019, 9). African Studies is therefore becoming more relevant to conversations about DH and about the expanding role of AI within and beyond scholarship.
Similarly, digital interventions on the continent, at times attached to scholarly institutions or pursuits, are rich, undertheorised terrain for discourse and technical development. African language projects like NEH Ajami Project,[30] Túbọ̀sún’s Yoruba Language Projects,[31] and Ngue Um’s work on endangered / less-endowed languages in Cameroon hold critiques and learnings around interfacing holistic, fluid, multi-modal datasets with rigidly categorising and Latin alphabet / text-based tools. Africa-based, public-facing projects such as Archivi.ng and Open Restitution Africa offer untapped insights into frugal digital interventions, meaningfully aggregating fragments, and centring people and their agency in content, design, and exploration. Africa-wide projects such as ALMEDA, meanwhile, exemplify the hurdles and possibilities of crowdsourcing and linking inherently incomplete databases / sets.
Yet, mainstream digital scholarship and engagements with digital records, as well as broader conversations of what an ethical use of AI might look like, often uncover the absence of both African Studies and African ways of knowing and doing. The absence matters most where technologies and their effects take concrete form, and where decision-makers sit, in tech and policy environments alike. These absences have material consequences for the Global Majority, particularly in African contexts, where technologies, often designed and governed through an abstracted “Global North” lens, become increasingly deployed at scale – often under the banner of “development”. In such cases, assumptions become embedded in systems that shape access to resources and the terms in which futures can be imagined. They also mediate visibility and inform high-stakes decisions that often reinforce existing inequalities while rendering local epistemologies invisible or legible only on extractive terms.
At the same time, the lack of approaches that appropriately value African sources of knowledge points to a methodological gap. Empirical and conceptual frameworks are needed to contest, measure, and dismantle lingering colonial legacies in DH and AI. This is evident in data representations and visualisations that offer incoherent narratives of Africa (and, indeed, the world), as well as in AI infrastructures and datasets that fail to account for the Global Majority’s situation.
Thus, we contend that knowledge and practical digital initiatives from Africa belong at the centre of integrating AI and DH with African Studies, rather than at the margins of current AI and DH discourses, where the granular details of African cultural records can sometimes be erased. Placing them there is essential both for more rigorous interdisciplinary learning and for rethinking who participates in and influences the design and governance of these technologies, and who should. That rethinking also means ensuring the meaningful inclusion of African scholars, practitioners, and institutions in the decision-making spaces where priorities are set and systems are designed.
5.1.1 Gaps, silences, inequalities, and possibilities in bidirectionality
Conversely, we recognise that Global DH and AI contribute to reshaping and expanding African Studies in a genuinely reciprocal exchange, particularly by offering access to new tools, methods, and knowledge infrastructures that, when critically adapted, can support new forms of enquiry, preservation, and analysis of African histories, languages, and cultural records. Uncritical adoption, however, is likely to echo the same epistemic hierarchies, extractive practices, and representational distortions that African Studies developed in order to contest.
One central question that needs persistent foregrounding is what precisely the cost of inclusion is; or, as Couldry and Mejias (2019) invite us to consider in The Costs of Connection: how do data and AI infrastructures colonise life for capitalist systems and repressive governments? For instance, does integrating images, languages, and performances from African societies empower those societies while also exposing them to surveillance by corporations and states? With regard to archives, how seriously have we grappled with early warnings such as O’Connell’s: “If the white walls of the archive are extended into the unlimited space of the Internet, what will be the price that … [Africans] … once again will be expected to pay?” (2008, 60). Which violences are being repeated and exacerbated in the act of digitisation – particularly, as Odumosu (2024) points out, at the level of consent? What strategies have been used to mitigate these harms, given that colonial cultural institutions and their records are often what Sprute (2023) terms “places of non-knowledge”?
In 2016, Gallon called for a “technology of recovery” (2016), a critical approach to the digital infrastructure used in Black and African DH. Ten years later, in 2026, the practical implementation of “technology of recovery” remains challenging, as persistent problems continue to affect efforts to democratise digital spaces, infrastructure, and African knowledge production, much of which is still dominated by “Global North” systems. These conditions continue to create a domino effect, keeping computational processes within such infrastructures grounded in racialised colonial logics (Noble 2019; Onuh 2025). In 2025, Błoch, Martos Oms, and Santana conducted experiments on how modern AI models not trained on African historical data affect results. They found that both machine learning techniques and LLMs consistently misclassified entities containing African vocabulary, with custom models – despite being tailored to the task – showing the same limitations as pretrained ones. When misclassified terms were replaced with Western language equivalents, model performance improved significantly. They observed that data imbalance, rather than model infrastructure or level of customisation alone, was responsible for the poor results in African historical resources (Błoch et al. 2025).
5.1.2 Possibilities and emerging practices in African DH and AI
Rather than treating the gaps, silences, and inequalities of data curation, digital infrastructure, and institutional capacity in and about Africa as grounds for fatalism, we note that many ongoing initiatives are creating new archives and datasets, or facilitating new insights altogether. AI and DH are helping to expand these endeavours.
For example, despite the reality that the majority of national African archives result from colonial projects, in “Out of the Ashes: Rethinking Loss in the African Archive” Ashie-Nikoi writes: “In its diverse languages and forms, the African archive can challenge the language and form in which African Studies is produced” (2024, 35). In Mali, Sidibé’s chance discovery of Malian women’s written works and photographs from critical transition periods in the country led to the development of the Archives des Femmes du Mali project. Once at risk, these personal collections are now digitally preserved, and they allow Malian women to be written back into the country’s political and social transformations as active agents rather than passive observers. In Lagos, Archivi.ng, an independent initiative, has relied on readily available and emergent AI-powered software and applications to digitise deteriorating and aged Nigerian newspapers and magazines (Ameenat 2024). Their inclusive material selection process has involved partnering with local publishers and receiving donations from everyday family archives. The result decentres and disrupts the “illusion of totality and continuity” presented by both the colonial and the post-colonial state archive (Mbembe 2002, 21). The digitisation has been coupled with an ancillary online publication, The Archivist; an annual open-to-all gathering centred on collective memory, the August Event; and, more recently, an interdisciplinary fellowship that drew over 2,500 applicants. These additional efforts to process, refine, and enliven the digital archive have put forth Nigerian-centric uses and sense-making of the content.
Thus, whereas DH and AI have thrown a spotlight on and encouraged the production of LLMs, we cannot overlook how the digital turn is increasing the integration of African voices and multimedia from Africa into African Studies. Texts are being digitised, recirculated, and interpreted next to a growing mass of local audiovisual content and in-person engagement, in line with the intermedial ontology of African oral aesthetics and indigenous performances. Furthermore, while knowledge, along with digital tools and archives, is never neutral, decolonising digital infrastructure and archives, for example, can take the form of developing innovative methods of sharing digital projects and initiatives that transcend the standard practices often dictated by funding agencies. Designing effective means of knowledge sharing is key to ensuring that the increasing number of digital initiatives being undertaken across Africa and other parts of the world reach the communities from which their material was collected, as well as students, scholars, and the general public, as well as scholarly audiences.
Digitised content can be disseminated in many ways: sharing with local partners and institutions the links to digital archives hosted at Western institutions or on Western servers; linking the websites of the partners involved, wherever possible; supporting local partners to develop their own websites; depositing copies of digital archives with local partners; returning offline archival copies to the communities of origin; and “experiential sharing of the archives”, in which the content is viewed together with local communities and the experiences of community members are documented, thus creating “living archives”. This latter approach has been central to Mark-Thiesen’s project digitising 1980s television content from the Liberia Broadcasting System (Mark-Thiesen 2026). At Open Restitution Africa, the development and publication of an open data platform[32] presenting pan-African perspective-driven research data on restitution processes has occurred alongside a data socialisation strategy that recognises the varied levels at which people enter into the subject matter, as well as inequalities in digital access and digital literacy. This has included short-form multimedia publications on social media, African practitioner-centred gatherings, briefings and guides, and pairing the screenings of an introductory explainer video series[33] with local case study researchers’ presentations of their findings in their respective cities or communities across the continent. This kind of work is in line with how social media platforms, despite their not being archival infrastructures, are also being used by individuals and local communities as digital archives that document the African past and, specifically, urban experiences (Yékú 2024).
Local actors and communities are more than consumers of such content. They have a part to play in knowledge production itself: enriching data and metadata, handling the post-processing such work requires, and producing “metaknowledge” (knowledge about knowledge). This enrichment, drawing on African sources and datasets, is an urgent priority if the “technology of recovery” is to work in practice. Meeting that priority calls for measures on two timescales. In the short term, digital archives and platforms such as Wikipedia need revising so that African languages and knowledge systems are properly represented. SADiLaR’s SWiP project,[34] a collaboration with Wikimedia South Africa and the Pan South African Language Board, trains community members to create and expand Wikipedia content in their own languages, beginning with isiNdebele and since extending to other South African languages (Matfunjwa et al. 2025).
Metadata needs the same attention. Many such platforms do not adequately recognise non-mainstream languages in their metadata descriptions, and so impose Western standards on African languages and knowledge systems. Metadata development and the post-processing of archival materials must therefore be strengthened. That means hyperlinks for non-indexed or underrepresented terminologies (Martino 2014), corpora built from scratch (Ngué Um et al. 2022), and active epistemic intervention, whether through open-access software (Jethro 2023) or through digital annotation that overlays new interpretive layers onto historical expressions (Odumosu 2025). Linked Open Data principles would take the work further, allowing archives and datasets to be linked and cross-referenced. In the longer term, more inclusive and equitable infrastructures depend on digitising African sources and democratising access to digital knowledge. Such work would support open data practices, more inclusive datasets, and communities formed around data sharing.
We recognise that all archives and data carry power dynamics. Rather than making a general case for the benefits and harms of digitisation and AI, we want to centre two concrete questions: How can we celebrate achievements in the preservation of diverse materials and generation of new information such as the above, while also remaining aware of what elements are being overshadowed? In addition, what new concerns surrounding ownership and surveillance emerge amid digitisation processes? These are only some of the questions we encourage fellow scholars to continue surfacing – and that funders and other decision-makers should recognise as deserving further exploration.
5.1.3 Incompleteness, conviviality, and collective sense-making as decolonising methods
Recognising the value of fluidity is urgent if researchers are to resist the reinscription of simplistic, colonial epistemologies in their use of AI and digital tools. Nyamnjoh’s invitation (2017) to celebrate and embrace incompleteness as a normal order of things should guide this effort, as it points to the grandiosity inherent in ambitions to preserve and claim wholeness. Conviviality, in this framing, is a way of extending ourselves beyond what we know and through potencies that other humans or super- and suprahuman forces can grant us. The question becomes what this may mean in the context of an increasingly habitualised use of digital means – for scholarly purposes and otherwise.
Incompleteness, understood as an everyday practice of negotiating difference and interdependence, shows that “the digital” lets people remake sociality through improvisation, friction, and mutual accommodation, though often within asymmetrical infrastructures shaped by global tech power. In reflecting on the making of ALMEDA’s repository, which seeks to link African literary metadata, Harris notes that “no data repository ever represents the entire picture of knowledge production on any given topic” (2025, 5). The resultant exercise of tracing African authors across Global and African repositories revealed the incompleteness of each of the databases, but also bias in represented regions, literary forms, and languages, as well as the invisibilised and quickly disappearing physical materials that hold the key to subverting this. These discoveries affirm the necessity of “technologies of resistance” – at both a critical and technical level – but also the importance of consultation and collaborations across disciplines and skill sets, in-process publishing that alerts others to limitations in technologies, and refiguring research value rationales that ensure the “future of the field” rather than singular researchers (Proferes et al. 2020; Harris 2025, 8, 15).
With AI scaling old questions of resource inequities and related discrepancies in determining who becomes visible (and at what cost), it has become more important than ever to reiterate both pluriversal and incomplete epistemic traditions – and to renew attention to local knowledge, transparency of process, collective and collaborative design, knowledge and sense-making practices (Junck 2024), as well as scholarly and applied studies in and of Africa. Understanding how knowledge circulates in partnering communities (including but not limited to political, institutional, bureaucratic, diplomatic, and cultural realities) is also important for successful preservation, development and dissemination of digital projects that do justice to local perspectives and enhance AI and DH infrastructures and outputs. But, of course, there is no singular “African” or “local” perspective, so why and how can we acknowledge incompleteness / trade-offs in our local partnerships / engagements, as well?
5.2 Ethical frameworks for a pluriversal DH and AI
Africa-centred perspectives bear directly on the ethical frameworks that govern AI and DH implementation. But how can we establish the value of such frameworks beyond the field of African Studies? A pluriversal approach offers essential possibilities to thinkers disenchanted with the existing model of “universal” thought. While thinking about ethical norms at a rational human level may provide an easy solution in the sphere of technical work, we need to embed critical (including scientific) outlooks in DH and AI endeavours. This starts with treating technology as a human product rather than an agent, and centring the fallible human aspects of such work. Recognising the limits of technologies, and the situated judgements built into them, makes research and AI work an ongoing, never-complete effort that leaves room for plural ways of understanding grounded in diverse lived realities.
Ethical norms are inseparable from the narratives, stories, and subjectivities of any given community (whether national, institutional, scholarly, or disciplinary); therefore, being transparent about these influences is necessary. Conceptual work should also leave room for ethics and governance as processes that are renegotiated continually, always in relation to particular engagements and concrete temporalities. Such processes yield the best available answer at a given moment, but each answer has to be revisited and remade, not least because a specific ethical norm can be invoked to justify a wide range of practices. As an example, in terms of decolonial praxis, an ethical approach could emphasise “social justice” or “care” for sensitive materials. Beyond technical achievements, how we write about projects like Uduah’s Biafra War Memories[35] must be carefully managed to respect the delicate nature of the archive, particularly regarding a history the Nigerian state often seeks to either erase or marginalise.
The ethical questions digital humanists face today include scholarly and transdisciplinary concerns: open access and the moral obligation to share research outputs with the communities that have a stake in them, while accounting for risks of extraction and misappropriation. Ethical considerations are both material (the ethics of data storage, the curation of cultural heritage and records) and immaterial (the careful handling of sensitive materials). Additional worries surround informed user consent in the context of big data research, data use and reuse, researcher reflexivity in scholarship, and the ethical complexities of research funding while maintaining humanist ideals (Proferes et al. 2020). These questions are further distilled into ethical approaches that balance access with the agency of knowledge holders and consider how value is distributed across the whole chain of contributors, custodians, and institutions that a project involves. Those approaches in turn shape the relational and technological infrastructures that get built.
Ngom, who has led several digital projects on classical Arabic and Ajami manuscripts, has arrived at a set of key ethical practices (Ngom and Castro 2019). These include identifying the rightful owners of archives according to the community’s own perspectives, and developing credit and copyright agreements the owners can understand, translated into the local language and script and read aloud where necessary. Such agreements go beyond Western institutions’ standard copyright agreements: they incorporate local concerns, respect local privacy and safety, and protect both the owners and the digitisation team. Equally important is having experts well versed in local culture, knowledge systems, and perspectives vet archival content for accuracy and against misrepresentation, so that the digital archives produced meet the highest ethical standards. Integrating equitable resource flows across each project’s ecosystem has also led to innovations in grant writing.
While scholars in African Studies have always grappled with ethical dilemmas – such as the “objectivity” of knowledge – the characteristic methods and approaches of DH scholarship in its African / postcolonial articulations raise additional questions, some of which, in non-digital contexts, have previously been raised by African Studies scholars. For example, as it relates to the subject of “Global North”–“Global South” collaborations, in Yékú’s digital archive project, Digital Nollywood,[36] the primary movie posters constituting the collection were made possible by colleagues in Nigeria who first collected and later digitised the work. For a project based in a Western academic institution in the United States, using these materials in a non-extractive way that is cognisant of the labour of these initial partners is another ethical move that reflects best practice in both DH and traditional African Studies research.
These are moves that could help ensure that computational methods do not simply digitise colonial categories but instead create space for African ways of organising and transmitting knowledge – as part of individual scholarly enquiries but also, hopefully, beyond academic silos.
6 Conclusion: From Principles to Commitments
The cases examined here yield no universal model for DH and AI in African Studies. They do support a set of commitments against which projects, infrastructures, funding programmes, and policies can be judged, all anchored in the question of our title: for whom is this work undertaken, and what purpose does it serve? “For whom?” asks about access and governance: who can reach a system, and who trusts it enough to use it. “For what purpose?” asks about sovereignty: who sets a system’s aims, and who can change them. Design, as Section 4 argues, emerges from these conditions. Because they shift over time, ethics stays a process rather than a deliverable.
- Begin with purpose, not technology. Identify intended communities, benefits, harms, and excluded users before selecting tools; sometimes the responsible decision is not to digitise at all.
- Treat communities as rights-holders, with authority over selection, description, access, reuse, withdrawal, and the distribution of benefits.
- Make ethics an ongoing process: consent, licences, agreements, and governance arrangements are living documents, revisited as technologies and communities change.
- Design from actual conditions of connectivity, power, devices, language, literacy, and trust; offline-first, low-bandwidth, multilingual, and multimodal approaches count as deliberate design choices.
- Reject deficit framings: “low-resource” describes pipelines, not languages; invest in speech, oral knowledge, plural orthographies, and naturally occurring language use.
- Use standards as revisable scaffolding that accommodates variation and documents what it excludes.
- Define openness through equity: CARE and community-return licences such as NOODL belong in decisions about how data and heritage circulate.
- Keep AI assistive and proportionate: demonstrable benefit to scholarship and to source communities, human verification, visible provenance, and honest accounting of the annotation labour and environmental burden that fall disproportionately on the continent.
- Fund stewardship as seriously as innovation: fair labour, local capacity, maintenance, succession, and the physical originals alongside the digital copies.
- Place African knowledge and practice at the centre of global DH and AI. The problems set out here run through both fields everywhere; African contexts make them legible, and Africa-based work on incompleteness, conviviality, and linguistic plurality offers methods for addressing them.
These commitments translate into steps on two timescales. Any project can begin now: write down its audience and purpose, reopen agreements to cover AI uses their signatories could not foresee, label AI-generated records, and publish accounts of its selection choices. The intermediate work falls to funders, who should write stewardship, equitable licensing, and capacity-building into grant conditions; to institutions, which should sustain public repositories that supply persistent identifiers, preservation, and training while depositors retain copyright and set their own access conditions, with centralisation decided collection by collection; and to policymakers, who should carry the Continental AI Strategy and the UNESCO Recommendation into procurement, labour protections, and policy published in national languages.
The commitments apply to the workshop that produced them. Anglophone bias in the field’s tools, conferences, and networks shaped who was in the room: francophone African colleagues were under-represented. The account given here is partial in the ways the paper itself describes, and any continuation of this work has to be convened differently.
None of this dissolves the tensions between access and sovereignty, standardisation and variation, preservation and ephemerality, innovation and sustainability. We did not resolve them among ourselves either: we disagreed about whether a researcher who uncovers something an archive’s owners did not know has an obligation to tell them, and about how far AI-generated metadata should be allowed to run ahead of human-designed ontologies. The commitments require such tensions to be made explicit, negotiated in context, and revisited over time. DH and AI in African Studies should be judged less by novelty, scale, or technical sophistication than by who shapes them, who benefits, who bears the costs, and whether communities retain meaningful authority over the futures built from their languages, histories, and heritage.
References
Adebara, Ife. 2025. AI and Language Data Flaring in Africa: Addressing the Low-Resource Challenge. No. 216. Policy Brief. Centre for International Governance Innovation. https://www.cigionline.org/publications/ai-and-language-data-flaring-in-africa-addressing-the-low-resource-challenge/.
Adelani, David Ifeoluwa, Jessica Ojo, Israel Abebe Azime, et al. 2025. ‘IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models’. arXiv:2406.03368. Preprint, arXiv, January 23. https://doi.org/10.48550/arXiv.2406.03368.
African Union. 2024. ‘Continental Artificial Intelligence Strategy: Harnessing AI for Africa’s Development and Prosperity’. https://au.int/sites/default/files/documents/44004-doc-EN-_Continental_AI_Strategy_July_2024.pdf.
Ameenat, Ayodimeji. 2024. ‘The Economics of Archiving Old Newspapers’. Archivi.Ng, August 28. https://archivi.ng/the-archivist/stories/issue-1/the-economics-of-archiving-old-newspapers.
Ashie-Nikoi, Edwina D. 2024. ‘Out of the Ashes: Rethinking Loss in the African Archive’. Social Dynamics 50 (1): 26–42. https://doi.org/10.1080/02533952.2024.2327248.
Barringer, Terry, Jos Damen, Peter Limb, and Marion Wallace. 2014. ‘Introduction’. In African Studies in the Digital Age: Dis/Connects?, edited by Terry Barringer and Marion Wallace. Brill. https://doi.org/10.1163/9789004279148_002.
Birhane, Abeba. 2020. ‘Algorithmic Colonization of Africa’. SCRIPTed 17 (2): 389–409. https://doi.org/10.2966/scrip.170220.389.
Birhane, Abeba. 2023. ‘Algorithmic Colonization of Africa’. In Imagining AI: How the World Sees Intelligent Machines, edited by Stephen Cave and Kanta Dihal. Oxford University Press. https://doi.org/10.1093/oso/9780192865366.003.0016.
Błoch, Agata, Guillem Martos Oms, and Clodomir Santana. 2025. ‘Decolonizing Archival Narratives: Exploring Digital Bias in the Catalogs of Portuguese-Colonized African Territories’. The Journal of African History 66: e19. https://doi.org/10.1017/S0021853725100601.
Breckenridge, Keith. 2014. ‘The Politics of the Parallel Archive: Digital Imperialism and the Future of Record-Keeping in the Age of Digital Reproduction’. Journal of Southern African Studies 40 (3): 499–519. https://doi.org/10.1080/03057070.2014.913427.
Brett-Smith, Sarah C. 2001. ‘When Is an Object Finished? The Creation of the Invisible among the Bamana of Mali’. Res: Anthropology and Aesthetics 39: 102–36. https://doi.org/10.7282/T3668B97.
Carroll, Stephanie Russo, Ibrahim Garba, Oscar L. Figueroa-Rodríguez, et al. 2020. ‘The CARE Principles for Indigenous Data Governance’. Data Science Journal 19 (1): 43. https://doi.org/10.5334/dsj-2020-043.
Carroll, Stephanie Russo, Edit Herczog, Maui Hudson, Keith Russell, and Shelley Stall. 2021. ‘Operationalizing the CARE and FAIR Principles for Indigenous Data Futures’. Scientific Data 8 (1). https://doi.org/10.1038/s41597-021-00892-0.
Caruso, Mariflora, and Alessandro Spadaro. 2024. ‘Digital Humanities and Artificial Intelligence: An Accelerationist Perspective of the Future’. Proceedings 96 (1): 10. https://doi.org/10.3390/proceedings2024096010.
Chamelot, Fabienne, Vincent Hiribarren, and Marie Rodet. 2020. ‘Archives, the Digital Turn, and Governance in Africa’. History in Africa 47: 101–18. https://doi.org/10.1017/hia.2019.26.
Chun, Jon, and Katherine Elkins. 2023. ‘The Crisis of Artificial Intelligence: A New Digital Humanities Curriculum for Human-Centred AI’. International Journal of Humanities and Arts Computing 17 (2): 147–67. https://doi.org/10.3366/ijhac.2023.0310.
Comaroff, Jean, and John L. Comaroff. 2012. ‘Theory from the South: Or, how Euro-America is Evolving Toward Africa’. Anthropological Forum 22 (2): 113–31. https://doi.org/10.1080/00664677.2012.694169.
Couldry, Nick, and Ulises A. Mejias. 2019. The Costs of Connection: How Data Is Colonizing Human Life and Appropriating It for Capitalism. Stanford University Press.
Dingemanse, Mark. 2018. ‘Redrawing the Margins of Language: Lessons from Research on Ideophones’. Glossa: A Journal of General Linguistics 3 (1): 4. https://doi.org/10.5334/gjgl.444.
Eisenhuth, Philipp, Myriel Fichtner, Britta Frede, and Rüdiger Seesemann. 2023. ‘Developing Crosslingual Ontologies in WissKI: Transcontinental Research Collaboration in the Africa Multiple Cluster of Excellence’. Modern Languages Open, no. 1: 39. https://doi.org/10.3828/mlo.v0i0.445.
Foster, Christopher. 2024. ‘Intellectual Property Rights and Control in the Digital Economy: Examining the Expansion of M-Pesa’. The Information Society 40 (1): 1–17. https://doi.org/10.1080/01972243.2023.2259895.
Frede, Britta. 2025. ‘Historical Amnesia at Work: Reflecting About the “Absence in the Presence” of Female Authors in Saharan Arabic Manuscript Collections’. Hawwa 23 (3–4): 433–69. https://doi.org/10.1163/15692086-12341446.
Fricker, Miranda. 2007. Epistemic Injustice: Power and the Ethics of Knowing. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780198237907.001.0001.
Gagliardi, Susan Elizabeth. 2022. ‘Mapping Senufo: Process, Collaboration, and Generous Thinking’. Mande Studies 24: 251–63. https://doi.org/10.2979/mnd.2022.a908479.
Gagliardi, Susan Elizabeth. 2023. Seeing the Unseen: Arts of Power Associations on the Senufo-Mande Cultural ‘Frontier’. Indiana University Press. https://doi.org/10.2979/seeingtheunseen.
Gagliardi, Susan Elizabeth, and Constantine Petridis. 2021. ‘Mapping Senufo: Reframing Questions, Reevaluating Sources, and Reimagining a Digital Monograph’. History in Africa 48: 165–209. https://doi.org/10.1017/hia.2021.5.
Gallon, Kim. 2016. ‘Making a Case for the Black Digital Humanities’. In Debates in the Digital Humanities 2016, edited by Matthew K. Gold and Lauren F. Klein. Debates in the Digital Humanities 2. University of Minnesota Press. https://dhdebates.gc.cuny.edu/read/untitled/section/fa10e2e1-0c3d-4519-a958-d823aac989eb.
Gibson, Laura. 2024. ‘Digitization Is Not Decolonization: South Africa’s Amagugu Ethu Museum Project and Colonial Documentation in Digital Times’. Museum Worlds 12 (1): 31–46. https://doi.org/10.3167/armw.2024.120104.
Gikunda, Kinyua. 2023. ‘Empowering Africa: An In-Depth Exploration of the Adoption of Artificial Intelligence Across the Continent’. arXiv:2401.09457. Preprint, arXiv, December 28. https://doi.org/10.48550/arXiv.2401.09457.
Guiliano, Jennifer, Roopika Risam, Leah Junck, and James Yékú. 2024. ‘Editors’ Note: April 2024’. Reviews in Digital Humanities 5 (4). https://doi.org/10.21428/3e88f64f.e16773b2.
Harris, Ashleigh. 2025. AI and African Literary Studies. December 6. https://doi.org/10.5281/zenodo.17839280.
Hussen, Kedir Yassin, Walelign Tewabe Sewunetie, Abinew Ali Ayele, Sukairaj Hafiz Imam, Shamsuddeen Hassan Muhammad, and Seid Muhie Yimam. 2025. ‘The State of Large Language Models for African Languages: Progress and Challenges’. arXiv:2506.02280. Preprint, arXiv, June 26. https://doi.org/10.48550/arXiv.2506.02280.
Isaacman, Allen, Premesh Lalu, and Thomas Nygren. 2005. ‘Digitization, History, and the Making of a Postcolonial Archive of Southern African Liberation Struggles: The Aluka Project’. Africa Today 52 (2): 55–77. https://doi.org/10.1353/at.2006.0009.
Jethro, Duane. 2023. ‘History Uploaded: Digital Archives After Thirty Years of Democracy’. South African Historical Journal 75 (3): 376–99. https://doi.org/10.1080/02582473.2024.2351828.
Junck, Leah. 2024. ‘Unlocking Human Centered AI: Building Inclusive Futures Through Community Engagement’. Global Center on AI Governance, May 13. https://www.globalcenter.ai/research/unlocking-human-centered-ai-building-inclusive-futures-through-community-engagement.
Khalīfa Muḥammad ʿUmar, Alexandros Tsakos, and Marianne Paasche. 2025. ‘Raqmanat al-makhṭūṭāt wa’l-wathāʾiq al-Sūdāniyya bi-Maktabat Jāmiʿat Bērgen: Majmūʿat al-Khurṭūm namūdhajan [Digitization of Sudanese Manuscripts and Documents at the University of Bergen Library: The Khartoum Collection as a Case Study]’. Ādāb: Majallat Kulliyyat al-Ādāb Jāmiʿat al-Khurṭūm [Journal of the Faculty of Arts, University of Khartoum] 53: 211–228. https://onlinejournals.uofk.edu/JFA/article/view/2093.
Kyomuhendo, Adam. 2025. ‘An African Perspective to Ethical Questions Posed by Artificial Intelligence and the Intersections with Climate Justice, Resilience, and Equity’. In Oxford Intersections: AI in Society, edited by Philipp Hacker. Oxford University Press. https://doi.org/10.1093/9780198945215.003.0029.
Kwarkye, Thompson Gyedu. 2025. ‘“We Know What We Are Doing”: The Politics and Trends in Artificial Intelligence Policies in Africa’. Canadian Journal of African Studies / Revue Canadienne Des Études Africaines 59 (3): 437–55. https://doi.org/10.1080/00083968.2025.2456619.
Kwet, Michael. 2019. ‘Digital Colonialism: US Empire and the New Imperialism in the Global South’. Race & Class 60 (4): 3–26. https://doi.org/10.1177/0306396818823172.
Loukissas, Yanni Alexander. 2019. All Data Are Local: Thinking Critically in a Data-Driven Society. The MIT Press. https://doi.org/10.7551/mitpress/11543.001.0001.
Mark-Thiesen, Cassandra. 2026. ‘Digitizing ELTV Television of the Liberia Broadcasting System (LBS) – Liberia’. Africa PID Alliance, April 21. https://docid.africapidalliance.org/docid/20.500.14351%2Fcb69e80b6183a422e4e9.
Martino, Enrique. 2014. ‘Open Sourcing the Colonial Archive – A Digital Montage of the History of Fernando Pó and the Bight of Biafra’. History in Africa 41: 387–415. https://doi.org/10.1017/hia.2014.15.
Muzi Matfunjwa, Nomsa Skosana, and Lebogang Boemo. 2025. ‘A Report for the SADiLaR-Wikipedia-PanSALB Project for South African Languages’. Forum for Linguistic Studies 7 (5): 598–603. https://doi.org/10.30564/fls.v7i5.9068.
Mbembe, Achille. 2002. ‘The Power of the Archive and Its Limits’. In Refiguring the Archive, edited by Carolyn Hamilton, Verne Harris, Jane Taylor, Michele Pickover, Graeme Reid, and Razia Saleh. Springer Netherlands. https://doi.org/10.1007/978-94-010-0570-8_2.
Money, Duncan. 2021. ‘Rebalancing the Historical Narrative or Perpetuating Bias? Digitizing the Archives of the Mineworkers’ Union of Zambia’. History in Africa 48: 61–82. https://doi.org/10.1017/hia.2021.6.
Musa, Imad. 2025. ‘Who Will Own and Control Africa’s AI Energy Future?’ The Republic, August 24. https://rpublc.com/story/2025/08/24/climate-change/africa-ai-energy.
Ngom, Fallou, and Eleni Castro. 2019. ‘Beyond African Orality: Digital Preservation of Mandinka ʿAjamī Archives of Casamance’. History Compass 17 (8): e12584. https://doi.org/10.1111/hic3.12584.
Ngué Um, Emmanuel, Émilie Eliette, Caroline Ngo Tjomb Assembe, and Francis Morton Tyers. 2022. ‘Developing a Rule-Based Machine-Translation System, Ewondo–French–Ewondo’. International Journal of Humanities and Arts Computing 16 (2): 166–81. https://doi.org/10.3366/ijhac.2022.0289.
Noble, Safiya Umoja. 2019. ‘Toward a Critical Black Digital Humanities’. In Debates in the Digital Humanities 2019, edited by Matthew K. Gold and Lauren F. Klein. Debates in the Digital Humanities 5. University of Minnesota Press. https://dhdebates.gc.cuny.edu/read/untitled-f2acf72c-a469-49d8-be35-67f9ac1e3a60/section/5aafe7fe-db7e-4ec1-935f-09d8028a2687.
Nyamnjoh, Francis B. 2017. ‘Incompleteness: Frontier Africa and the Currency of Conviviality’. Journal of Asian and African Studies 52 (3): 253–70. https://doi.org/10.1177/0021909615580867.
O’Connell, Siona. 2008. ‘No Hunting : Finding a New f. Stop for the Bushmen’. Master, University of Cape Town. http://hdl.handle.net/11427/11762.
Odumosu, Temi. 2024. ‘Approaching Colonial Photographs with Care’. Digital Benin, January 1. https://digitalbenin.org/documentation/approaching-colonial-photographs-with-care.
Odumosu, Temi. 2025. ‘Annotating The New Union Club: A Case Study on Critical Praxis for Digital Art Histories’. Nineteenth-Century Art Worldwide 24 (2). https://doi.org/10.29411/ncaw.2025.24.2.2.
Onuh, Frank. 2025. ‘Decolonial Prompting: Rewriting AI Toward Black Futures’. Crossing Boundaries and Recovering Intellectual Traditions (ASR 2025), November 20. https://doi.org/10.5281/zenodo.17655158.
Perrigo, Billy. 2026. ‘African Content Moderators Have Worse Mental Health than Global Peers, Study Finds’. Time, March 27. https://time.com/article/2026/03/26/africa-content-moderators-mental-health-study/.
Pickover, Michele. 2005. ‘Negotiations, Contestations and Fabrications: The Politics of Archives in South Africa Ten Years after Democracy’. Innovation 30: 1–11. https://doi.org/10.4314/innovation.v30i1.26493.
Pillion, Betsy, Lenore A. Grenoble, Emmanuel Ngué Um, and Sarah Kopper. 2019. ‘Verbal Gestures in Cameroon’. Theory and Description in African Linguistics (Berlin), August 13, 303–22. https://doi.org/10.5281/zenodo.3367152.
Proferes, Nicholas, Kristen Schuster, and Stuart Dunn. 2020. ‘What Ethics Can Offer the Digital Humanities and What the Digital Humanities Can Offer Ethics’. In Routledge International Handbook of Research Methods in Digital Humanities. Routledge. https://doi.org/10.4324/9780429777028-29.
Risam, Roopika. 2018. ‘Decolonizing the Digital Humanities in Theory and Practice’. In The Routledge Companion to Media Studies and Digital Humanities, edited by Jentery Sayers. Routledge. https://doi.org/10.4324/9781315730479-8.
Risam, Roopika. 2019. New Digital Worlds: Postcolonial Digital Humanities in Theory, Praxis, and Pedagogy. Northwestern University Press. https://doi.org/10.2307/j.ctv7tq4hg.
Risam, Roopika, and Alex Gil. 2022. ‘Introduction: The Questions of Minimal Computing’. Digital Humanities Quarterly 16 (2). https://doi.org/10.63744/b49fzhuz9hhz.
Sands, Bonny, and Kerry Jones, eds. 2022. Nǀuuki Namagowab Afrikaans English ǂXoakiǂxanisi/Mîdi Di ǂKhanis/Woordeboek/Dictionary. African Sun Media.
Shringarpure, Bhakti. 2020. ‘Africa and the Digital Savior Complex’. Journal of African Cultural Studies 32 (2): 178–94. https://doi.org/10.1080/13696815.2018.1555749.
Sprute, Sebastian-Manès. 2023. ‘Chaos im Museum: Bestandsaufnahme und Wissensordnung’. In Atlas der Abwesenheit: Kameruns Kulturerbe in Deutschland, edited by Mikaél Assilkinga, Lindiwe Breuer, Fogha Mc. Cornilius Refem, et al. Arthistoricum.net. https://doi.org/10.11588/arthistoricum.1219.
Thiam, Madina, Devon Golaszewski, Moussa Beïdy Tamboura, Oumou Sidibé, and Gregory Mann. 2023. ‘Le projet Archives des femmes : archiver, numériser et diffuser les luttes des femmes maliennes’. Revue d’histoire contemporaine de l’Afrique, no. 8. https://doi.org/10.51185/journals/rhca.2023.stc03.
Trouillot, Michel-Rolph. 1995. Silencing the Past: Power and the Production of History. Beacon Press.
Túbọ̀sún, Kọ́lá. 2022. ‘An Overview of the British Library Yorùbá Language Collection’. Africa Bibliography, Research and Documentation 1: 47–62. https://doi.org/10.1017/abd.2022.3.
UNESCO. 2023. ‘The UNESCO Recommendation on the Ethics of AI: Shaping the Future of Our Societies’. https://unesco.org.uk/resources/the-unesco-recommendation-on-the-ethics-of-ai-shaping-the-future-of-our-societies-english-version-german-national-commission.
UNESCO. 2025. ‘From Readiness to Ethical AI Adoption and Localization: UNESCO’. 23 July. https://www.unesco.org/en/articles/readiness-ethical-ai-adoption-and-localization-unesco-elevates-africas-voice-bangkok.
Wilkinson, Mark D., Michel Dumontier, IJsbrand Jan Aalbersberg, et al. 2016. ‘The FAIR Guiding Principles for Scientific Data Management and Stewardship’. Scientific Data 3 (1): 160018. https://doi.org/10.1038/sdata.2016.18.
Yékú, James. 2024.’Social Media Images as Digital Sources for West African Urban History’. In Oxford Research Encyclopedia of African History. Oxford University Press. https://doi.org/10.1093/acrefore/9780190277734.013.977.
Zaagsma, Gerben. 2023. ‘Digital History and the Politics of Digitization’. Digital Scholarship in the Humanities 38 (2): 830–51. https://doi.org/10.1093/llc/fqac050.
For questions and comments, please contact Frédérick Madore (frederick.madore@uni-bayreuth.de) and Vincent Hiribarren (vincent.hiribarren@kcl.ac.uk).