AfriChallenge

Structured capability evaluation

Benchmark data for African languages, written natively and traceable to its source.

AfriChallenge collects benchmark material across fifty languages. Every item is authored by a qualified speaker, independently reviewed, adjudicated where needed, validated and exported with its full history.

The challenge

AfriChallenge is a structured capability evaluation. Its data is not translated from another language: contributors write questions and answers natively in the target language, so the evidence reflects how the language is actually used.

What AfriChallenge collects

Four components, one platform. Knowledge is first; the others attach their own data to the same records rather than replacing them.

  • Knowledge

    Questions and reference answers written natively in the target language, each with structured source evidence, in five domains.

  • Speech and Prosody

    Audio with recording metadata, transcripts, second-transcriber verification and consent and provenance records.

  • LINDA

    Paradigms containing several minimal pairs, with phenomenon metadata and native-speaker verification records.

  • Code-Switching

    Items involving two or more language codes, with pair and register metadata, the basis for attested switching, and specialised validation.

Knowledge covers five domains, each with five sub-domains: Health, Agriculture, Finance, Education and Government services.

How the data journey works

The usual path is short. Adjudication is an exception, used when it is needed.

  1. Step 1

    Author

    A qualified contributor writes an item natively, against an assigned blueprint, and cites where the answer is supported.

    Nativenot translated

  2. Step 2

    Review

    A different qualified contributor reviews the exact submitted version independently, without seeing the author or other reviewers.

    Blindauthor hidden

  3. Step 3If needed

    Adjudicate

    Where a reviewer flags a problem or reviewers disagree, the Language Lead decides and must give a reason.

    On flagreason recorded

  4. Step 4

    Validate

    An item is validated only when every gate it needs is complete, including expert sign-off in high-consequence areas.

    All gatesbefore release

  5. Step 5

    Export

    Validated items are exported with their metadata, sources, reviews, decisions and history.

    JSONLwith manifest

Quality and integrity

The platform is the evidence system behind the benchmark, not just a form that stores answers.

Submitted versions never change
A correction creates a new version. The version that was reviewed stays exactly as it was.
Independent review is enforced
Reviewers are matched by language qualification, never receive their own items, and do not see other reviewers’ outcomes before they submit.
Every step leaves a record
Who did what, to which version, when, and under which rules, is written to an append-only log.
Nothing is silently overwritten
Removals are retirements, not deletions. Original records are kept.
Public and held-out data are separated
Assignment is recorded from the start and hidden from contributors and reviewers.
Rules are versioned
Length guidance, similarity checks and the review rubric are configuration with recorded versions, so a submission can be explained later.

Languages

The platform is built for fifty languages. Contributors work only in languages where their qualification has been confirmed.

Writing systems

For contributors

  • Accounts are created by the project team after your language qualification is confirmed.
  • You keep one identity. Your queue shows authoring tasks and review tasks, limited to your languages.
  • Drafts save as you write and are kept if your connection drops.
  • You never review your own work, and your reviews are recorded against the exact version you saw.

Already qualified and invited?

Sign in to see your authoring and review tasks.

Sign in