Structured capability evaluation
Benchmark data for African languages, written natively and traceable to its source.
AfriChallenge collects benchmark material across fifty languages. Every item is authored by a qualified speaker, independently reviewed, adjudicated where needed, validated and exported with its full history.
The challenge
AfriChallenge is a structured capability evaluation. Its data is not translated from another language: contributors write questions and answers natively in the target language, so the evidence reflects how the language is actually used.
What AfriChallenge collects
Four components, one platform. Knowledge is first; the others attach their own data to the same records rather than replacing them.
Knowledge
Questions and reference answers written natively in the target language, each with structured source evidence, in five domains.
Speech and Prosody
Audio with recording metadata, transcripts, second-transcriber verification and consent and provenance records.
LINDA
Paradigms containing several minimal pairs, with phenomenon metadata and native-speaker verification records.
Code-Switching
Items involving two or more language codes, with pair and register metadata, the basis for attested switching, and specialised validation.
Knowledge covers five domains, each with five sub-domains: Health, Agriculture, Finance, Education and Government services.
How the data journey works
The usual path is short. Adjudication is an exception, used when it is needed.
Step 1
Author
A qualified contributor writes an item natively, against an assigned blueprint, and cites where the answer is supported.
Nativenot translated
Step 2
Review
A different qualified contributor reviews the exact submitted version independently, without seeing the author or other reviewers.
Blindauthor hidden
Step 3If needed
Adjudicate
Where a reviewer flags a problem or reviewers disagree, the Language Lead decides and must give a reason.
On flagreason recorded
Step 4
Validate
An item is validated only when every gate it needs is complete, including expert sign-off in high-consequence areas.
All gatesbefore release
Step 5
Export
Validated items are exported with their metadata, sources, reviews, decisions and history.
JSONLwith manifest
Quality and integrity
The platform is the evidence system behind the benchmark, not just a form that stores answers.
- Submitted versions never change
- A correction creates a new version. The version that was reviewed stays exactly as it was.
- Independent review is enforced
- Reviewers are matched by language qualification, never receive their own items, and do not see other reviewers’ outcomes before they submit.
- Every step leaves a record
- Who did what, to which version, when, and under which rules, is written to an append-only log.
- Nothing is silently overwritten
- Removals are retirements, not deletions. Original records are kept.
- Public and held-out data are separated
- Assignment is recorded from the start and hidden from contributors and reviewers.
- Rules are versioned
- Length guidance, similarity checks and the review rubric are configuration with recorded versions, so a submission can be explained later.
Languages
The platform is built for fifty languages. Contributors work only in languages where their qualification has been confirmed.
Writing systems
ኢትዮጵያ
Ethiopic script, including Amharic and Tigrinya
العربية
Arabic script, written right to left
ⵜⴰⵎⴰⵣⵉⵖⵜ
Tifinagh script
ɓ ɗ ƙ ŋ ẹ
Latin script with language-specific diacritics and special characters
For contributors
- Accounts are created by the project team after your language qualification is confirmed.
- You keep one identity. Your queue shows authoring tasks and review tasks, limited to your languages.
- Drafts save as you write and are kept if your connection drops.
- You never review your own work, and your reviews are recorded against the exact version you saw.
Already qualified and invited?
Sign in to see your authoring and review tasks.
Sign in