Methodology
Prediction of the future is the best measure of intelligence. Kimpton records forecasts and actions before their outcomes exist, then scores them against reality.
The problem
Static evaluations can leak into training data, saturate, and change with prompts or graders.
Simulations reflect only the rules and consequences their designers encode.
Performance under new information and real consequences remains unknown.
Introducing Koliseum
Koliseum is Kimpton's live evaluation system. A model commits inside an arena, the record is sealed, and reality supplies the grade. It runs two kinds of arenas: one for what a model believes, one for what it can do.
Prediction Arenas
A model receives only the information available at a defined cutoff. Its forecast is sealed before the outcome exists and scored when the result becomes public.
Freeze inputs
Use only information available at the documented cutoff.
Seal the forecast
Record the output, confidence, sources, configuration, and timestamp.
Resolve
Wait for a public or independently verifiable outcome.
Score
Compare the forecast with the realized value and disclosed baselines.
Measurements
How a sealed forecast is graded once the outcome resolves. Each score is reported separately, against disclosed baselines.
Numerical
Absolute, percentage, or scaled error.
Probabilistic
Brier score or another proper scoring rule.
Relative
Improvement over consensus, market, or statistical baselines.
Calibration
Stated confidence compared with observed frequency.
Action Arenas
A model receives a bounded task and controlled access to the systems required to complete it. Success is verified from the resulting state and downstream outcome.
The transcript remains evidence, while completion must be independently verified.
Bound the task
Define the expected state, allowed tools, constraints, and rollback conditions.
Record execution
Capture the configuration, actions, tool calls, cost, and elapsed time.
Verify the state
Inspect the deliverable, system state, side effects, and downstream result.
Score the outcome
Report completion, reliability, safety, outcome quality, and efficiency separately.
Measurements
How a completed task is graded from the resulting system state. Each axis is reported separately so a blended score cannot hide a failure mode.
Completion
Verified from the final system state.
Reliability
Repeated success across comparable tasks.
Safety
Policy violations, prohibited actions, and irreversible side effects.
Outcome
Whether the result was accepted and held up downstream.
Efficiency
Cost and latency after the quality threshold is met.
Why the arenas matter
Static benchmark answers can enter training data, and synthetic tasks can reward assumptions built into the test. Live arenas preserve a clean temporal boundary: the model commits before reality supplies the answer or reveals the consequence.
Results remain inspectable by task, model configuration, outcome, baseline, cost, and failure mode. Prediction quality and action quality are reported separately because they answer different questions.
Build an arena with Kimpton
Talk directly with the founders about the prediction or action you need to measure.
Build with Kimpton