Research Project 2026- Deutsche Forschungsgemeinschaft (EXC 2052)

Digital Research Environment

A research programme on how African research data is described, modelled, and made findable.

The Digital Research Environment (DRE) is the digital infrastructure unit of the Africa Multiple Cluster of Excellence at the University of Bayreuth. We design, build, and maintain the data systems that connect researchers across the Africa Multiple Research Centres (AMRCs) and partner institutions worldwide: research data management, knowledge-graph development, AI-assisted processing, and digital literacy, grounded in the FAIR and CARE principles. I joined as Data Curator in 2026.

Two principles govern the work. Curation is shared: each AMRC (Ouagadougou, Makhanda, Lagos, Eldoret) describes the data it knows best, and the DRE runs the infrastructure that connects them. Storage is distributed: data remains in its local repository while Bayreuth holds the metadata layer pointing to it, so research data becomes findable without being relocated.

That metadata layer is public as the Africa Multiple Interactive Research Atlas (AMIRA), an Omeka S site that makes the collections searchable from one place : nearly 4,000 research items, more than 90 projects, some 1,600 people and 600 organisations, in 28 languages across 39 countries.

Federated description

The DRE works with each centre to assess what it holds and what can be linked to the platform: structuring and quality-checking collections, drafting data management plans, and preparing datasets for citation and long-term preservation.

Underneath the service is a research question: whose description counts. Good metadata makes a researcher’s work more discoverable, but the vocabularies that carry it are rarely neutral. Subject terms, where they exist at all, tend to derive from Library of Congress Subject Headings, often poorly suited to African sources and knowledge systems. We are working towards taxonomies in African languages, and towards modelling the relations between entities rather than only tagging them, so the records become a graph one can reason over. Reconciliation with Wikidata would anchor those entities outward, but its African coverage is thin, which makes the work a contribution as much as a lookup.

AI-assisted processing

The metadata researchers supply is often thin, and subject description is almost always the weakest part of it. We are experimenting with AI-assisted processing as an additional service: OCR for print and handwriting, transcription of audio and video, summarisation, entity extraction, tagging and classification. The models run locally, on open-source weights and the Cluster’s own hardware, which keeps the pipeline GDPR-compliant for material that is often special-category personal data, and keeps costs down.

Access and dissemination

A collection that cannot be interrogated is not findable in any sense that matters. Alongside searchable databases and project websites, the DRE builds dashboards, network graphs, and maps that foreground connections: the subjects that co-occur across collections, the collaboration networks no single catalogue record makes visible.

Users increasingly put their questions to an AI assistant instead of searching a catalogue. The AMIRA MCP server lets an assistant query the collection directly, so answers are drawn from the records and cite them, which reduces the hallucination that comes from asking a model what it already knows. Building and evaluating that layer is the current strand of the work.