Methodology

Prediction of the future is the best measure of intelligence. Kimpton records forecasts and actions before their outcomes exist, then scores them against reality.

The problem

Static evaluations can leak into training data, saturate, and change with prompts or graders.

Simulations reflect only the rules and consequences their designers encode.

Performance under new information and real consequences remains unknown.

Introducing Koliseum

Koliseum is Kimpton's live evaluation system. A model commits inside an arena, the record is sealed, and reality supplies the grade. It runs two kinds of arenas: one for what a model believes, one for what it can do.

PREDICTION ARENATests what a model believes will happenREALITYRESOLVESFORECAST RECORDEDERRERRACTION ARENATests what a model can accomplishTASKATTEMPTSTATE VERIFIEDROLLED BACK

Prediction Arenas

A model receives only the information available at a defined cutoff. Its forecast is sealed before the outcome exists and scored when the result becomes public.

Freeze inputs

Use only information available at the documented cutoff.

Seal the forecast

Record the output, confidence, sources, configuration, and timestamp.

Resolve

Wait for a public or independently verifiable outcome.

Score

Compare the forecast with the realized value and disclosed baselines.

Measurements

How a sealed forecast is graded once the outcome resolves. Each score is reported separately, against disclosed baselines.

Numerical

Absolute, percentage, or scaled error.

Probabilistic

Brier score or another proper scoring rule.

Relative

Improvement over consensus, market, or statistical baselines.

Calibration

Stated confidence compared with observed frequency.

Action Arenas

A model receives a bounded task and controlled access to the systems required to complete it. Success is verified from the resulting state and downstream outcome.

The transcript remains evidence, while completion must be independently verified.

Bound the task

Define the expected state, allowed tools, constraints, and rollback conditions.

Record execution

Capture the configuration, actions, tool calls, cost, and elapsed time.

Verify the state

Inspect the deliverable, system state, side effects, and downstream result.

Score the outcome

Report completion, reliability, safety, outcome quality, and efficiency separately.

Measurements

How a completed task is graded from the resulting system state. Each axis is reported separately so a blended score cannot hide a failure mode.

Completion

Verified from the final system state.

Reliability

Repeated success across comparable tasks.

Safety

Policy violations, prohibited actions, and irreversible side effects.

Outcome

Whether the result was accepted and held up downstream.

Efficiency

Cost and latency after the quality threshold is met.

Why the arenas matter

Static benchmark answers can enter training data, and synthetic tasks can reward assumptions built into the test. Live arenas preserve a clean temporal boundary: the model commits before reality supplies the answer or reveals the consequence.

Results remain inspectable by task, model configuration, outcome, baseline, cost, and failure mode. Prediction quality and action quality are reported separately because they answer different questions.

Build an arena with Kimpton

Talk directly with the founders about the prediction or action you need to measure.

Build with Kimpton