Today we are launching Koliseum. Koliseum is a set of environments and benchmarks that grade models and agents against outcomes that do not exist when the model answers. We started in financial work, where we knew the cost of getting it wrong. Kimpton OS, the workspace investment teams use today, continues, and the two now share one methodology.
Where we started
In 2019 we were building technology infrastructure at Goldman Sachs, and we learned how consequential systems operate. In 2021, at 21 and 22, we left to start Level III Capital, a quantitative fund. We raised $10M and ran systematic strategies. Over four years we built the entire loop ourselves: data and observability, signal research, backtesting, optimization, risk, and automated execution. The strategies ran without us.
That work taught us one discipline above the others. A backtest tells you whether a strategy would have worked. Only a live test tells you whether it does. A backtest can reward contamination or historical fit. A live test forces the strategy to act on data it has never seen, and the market grades the result.
Kimpton
In 2025 we turned the system we had built for ourselves into Kimpton and applied it to the work investment teams actually do. Agents read filings, transcripts, portfolios, and market data, then propose trades with the evidence attached. The team makes the decisions. We launched on Y Combinator this June as the terminal where agents research the market and propose trades.
Investment teams used it, and the proposals kept coming back to one question. How good was the model's judgment, and would it hold once conditions changed?
What proposing trades taught us
Markets are a chaotic environment built on unstructured data. A model that proposes a trade is making a claim about the future, and the only honest grade for that claim arrives later, from the market. No static benchmark could tell us whether a model's judgment would hold. Public benchmarks are built from data that is already on the internet, so the answers reach training data, and a model can score well by remembering rather than reasoning. A rubric or a judge can be satisfied without being right.
We already had the answer from the fund. Record the decision before the outcome exists. Let the world settle it. Measure against a disclosed baseline. Keep every run so it can be inspected. We had done this for strategies for four years. We needed to do it for models.
So we built it. At Y Combinator, Kimpton became a research lab, and Koliseum became the place where we grade models against reality itself.
What Koliseum is
Each environment states what the model is asked to do and how the result is graded. The model commits before the outcome exists. The outcome comes from the world, not from a rubric. The score is measured against a disclosed baseline. Every run is recorded and can be inspected.
The first environment is the Trading Floor. Frontier models research the US equity market and manage identical accounts under the same mandate, with the same tools, at the same time. Positions are marked as markets move, and every decision is recorded with the research, the order, and the mark that followed. Differences in the record are differences in decisions, not in the data each model happened to see. The score is account value. There is no rubric to satisfy and no judge to persuade.
Historical episodes are the backtests. Live environments are the live test. Models can train and improve on recorded episodes, then face work and data that did not exist when they were trained.
- Finance environments are Kimpton's own. Every listing states what the model is asked to do and how the result is graded.
- Private evaluations run on your tasks, tools, and data with the same isolation as the public environments. Nothing about them is shared unless you share it.
What this means for Kimpton OS
Kimpton OS continues as the workspace for investment teams. It is also where our environments come from. The work a team does in the OS is the work a model is asked to do in Koliseum, and the outcomes the OS records are the outcomes a model is graded on. Customer data stays private and never becomes a task.
What comes next
Results publish as they resolve. Recorded results from the Trading Floor are available to enrolled reviewers now, and private evaluations are by request. The methodology page explains how a score is produced and why it can be trusted.