Papers

Que sait une machine ? Épistémologie, gestion des données et enjeux de l’IA pour les langues africaines

KeynoteThu 24 · 09:30FR

Presented by

Abstract

“What does a machine know?” is usually asked as a question about capability: how well a system transcribes, translates or speaks. This presentation asks it as a question about epistemology. A learning machine knows its data, and only its data. It does not observe language; it minimises error over statistical regularities in whatever it has been given, stores those regularities as prior information, and thereafter perceives the world by inference from those priors (Hyvärinen 2024). What its data lacked, it cannot know; what its data contained by accident, it treats as law.

A second point is less often stated. A machine does not only know its data; it infers “knowledge” from the epistemology that presides over the design of its datasets. Every dataset encodes decisions about boundaries, categories and relevance before a single example is recorded: what counts as one language rather than two, which variety is the reference and which the “dialect”, which orthography is standard, which accent is native, which hesitation is noise. These categories present themselves as neutral descriptions of an objective linguistic reality. They are nothing of the kind. The enumeration of Africa’s languages into countable, mappable, codable units is the product of historically and politically situated epistemologies: colonial cartography, missionary orthography, the nation-state’s need for administrable populations, and the classificatory registries that now serve as default references for technology (Ngue Um 2025, forthcoming). Built into locale codes, benchmarks and validation rules, such categories are not merely inherited by the machine but re-enacted at scale, and the cost of the mismatch between category and lived linguistic experience is borne by speakers in their capacity as knowers (Fricker 2007).

If the epistemology of a dataset determines what a machine can know, the decisive question is who designs, publishes and governs the datasets that generate knowledge about a community’s linguistic experience. I call this capacity data stewardship: the ability of grassroots stakeholders of language work — speakers, teachers, community linguists, students — to take part not only in contributing data but in deciding how it is structured, under what terms it circulates, and to whom. Stewardship is the practical form of epistemic justice in the data economy.

The final part draws on the experience of the Mozilla Data Collective, a platform on which dataset owners publish, license and curate their own resources. Through the Institute of African Digital Humanities, more than 150 language and cultural datasets have been created and published by community agents: speech corpora in Afaan Oromoo, isiXhosa, Yoruba, Igbo, Hausa, Tiv, Lingala, Ewondo, Adamawa Fulfulde and many smaller languages; multimodal lexical resources; parallel corpora; and sociocultural datasets that document a community’s history, kinship, religion, economy and naming practices rather than a “language” abstracted from them. These datasets are governed by the agents who produced them, under a licence that requires users to declare their purpose and share value with the originating community. The result is a pathway of collaboration between those who use the data and the communities from which it comes, in an unmediated interaction that neither extractive data collection nor conventional academic publishing has allowed.

The stakes of AI for African languages, I conclude, are not first about access to technology but about which epistemology becomes the prior. Machines will hold beliefs about African languages whether or not Africans design the data; stewardship is the name for taking responsibility for those beliefs.

In the programme Keynote · Thu 24 · 09:30 View this session

DH & AI in African Studies

21–24 September 2026
STIAS, Stellenbosch, South Africa

Key Dates

  • Submission Deadline
  • Notification of Acceptance
  • Deadline for Full Papers
  • Workshop Dates

Supported by

Point SudSTIAS — Stellenbosch Institute for Advanced StudyDeutsche Forschungsgemeinschaft (DFG)Goethe University FrankfurtUniversity of Bayreuth / Africa MultipleKing's College LondonSADiLaR

© 2026 Frédérick Madore, Vincent Hiribarren, Emmanuel Ngue Um, Menno van Zaanen. Content under CC BY 4.0; paper titles, abstracts and biographies remain their authors'.