Africa Temporal Intelligence Corpus

Africa Temporal Intelligence Corpus (ATIC) is an open temporal intelligence corpus for Africa.

We transform African datasets into machine-readable timelines of observations, states, events, transitions, relationships, and evidence so that AI systems can reason more accurately about change over time.

ATIC is not just a dataset collection. It is a temporal memory layer for Africa.

Mission

Our mission is to build the world's highest-quality open temporal intelligence corpus for Africa: a trusted, continuously improving public resource for temporal reasoning, retrieval, forecasting, evaluation, and AI research.

We focus on one question:

What happened, when did it happen, what changed, and what evidence supports that interpretation?

What We Publish

ATIC publishes structured dataset packages across African countries and domains.

Each package may include:

Our initial domains include:

Temporal-First Design

Time is the primary axis of ATIC.

Every record is designed to support chronology, comparison, sequence modeling, and temporal reasoning. We care not only about isolated facts, but about how states change, how events unfold, and how evidence accumulates.

Examples of temporal annotations include:

Relationship to Electric Sheep Africa

ATIC uses Electric Sheep Africa as a key upstream source layer.

Electric Sheep Africa provides African datasets and data infrastructure. ATIC builds on top of that foundation by adding temporal structure, annotation layers, provenance, review status, and reasoning-oriented corpus formats.

In simple terms:

Corpus Principles

Time is the primary axis

Everything in ATIC is ordered in time. No annotation is complete without temporal context.

Every fact has provenance

Annotations should link back to source datasets, source records, evidence, generation methods, annotation versions, and reviewers.

Observations are separate from interpretation

ATIC distinguishes between:

This separation helps prevent raw data, model outputs, and human interpretation from being mixed together.

Models propose, humans approve

ATIC uses statistical methods and language models to assist annotation, but important annotations require human review. Review status is part of the data.

Corpus first, benchmark second, product third

ATIC is optimized for trust, reproducibility, and long-term research value. Benchmarks and applications are built from the corpus, not the other way around.

Dataset Package Format

ATIC dataset packages are versioned like software.

A typical package contains:

raw.parquet
labels.parquet
events.parquet
relationships.parquet
reasoning.jsonl
metadata.yaml
README.md
LICENSE
CHANGELOG.md

Each package is designed to answer:

  1. Where did this come from?
  2. Why was it labeled this way?
  3. Can it be reproduced?

Annotation Layers

ATIC organizes annotations into levels:

Level Layer Description
0 Raw Original or lightly normalized observations
1 Cleaned Typed, deduplicated, unit-normalized data
2 Statistical Trends, anomalies, missingness, seasonality, structural breaks
3 Temporal Increasing, declining, stable, recovering, accelerating
4 Events Elections, floods, outbreaks, budgets, policy changes
5 Relationships Before, after, during, lead, lag, overlap, follows
6 Reasoning Model-generated temporal explanations with provenance

Intended Uses

ATIC is built for: