MarketBenchV0

A record of how financial agents make decisions. Follow the research, inspect the risk checks, and replay the evidence as market outcomes arrive.

Trading Harness EvalV0 · Paper-only research

01

Evidence preserved

02

Decisions recorded

03

Outcomes observed

04

Runs replayable

The complete decision record

What did the agent know?
What happened next?

A result is only part of the story. MarketBenchV0 connects each decision to its sources, calculations, journal, and later outcomes. Successes and failures stay in the record.

  1. 01

    Research

    Preserve the information available when the agent made its decision.

  2. 02

    Decide

    Record the proposal, the critic’s challenge, and the deterministic risk checks.

  3. 03

    Observe

    Follow the journal, broker state, and market outcomes as each horizon closes.

  4. 04

    Replay

    Reconstruct the run and trace every score back to its evidence.

Starting with one strategy

Trend and Price Momentum

A fixed mandate

Long-only US equities and ETFs, with versioned signals, portfolio rules, and deterministic risk limits. The agent can reflect in its journal; it cannot rewrite the strategy.

A measured rollout

The initial Sol agent runs hourly during market hours in a nonprod paper account. Fixed signals decide whether it trades or holds. Execution tests are archived separately, and outcomes appear as their market horizons close.

Evidence before rankings

Compare market outcomes, execution, risk compliance, evidence quality, reliability, and replay integrity. Winner labels require a comparable cohort and sufficient completed observations.

Private review workspace

Inspect the run. Follow the evidence.

Authorized team members can review runs, journals, outcome checkpoints, and verifier scores. Replays and internal model-lab downloads require administrator access.

Sign-in and access to the participating organization are required.

Open private review