Cohort and selection
The frozen cohort contains 450 attempts across 30 synthetic software repair, database, and simulated API tasks grouped into 11 shared templates. Three model-and-harness configurations each completed five repeats per task. The pre-dispatch selection was fixed before execution, with no replacements or post-dispatch exclusions.
Every selected attempt remains in the outcome and success-rate denominator. Infrastructure and execution failures count as unsuccessful. The study includes all attempts rather than conditioning results on successful execution.
Scoring and uncertainty
Each attempt receives a task utility from 0 to 1. The success threshold is 1.0, with higher utility treated as better. Comparisons use the same task IDs and repeats across configurations. Confidence intervals use paired bootstrap resampling by shared task template, keeping task variants and repeats together across systems.
There are 11 template clusters, so uncertainty estimates describe a limited synthetic workload. The comparisons are descriptive and are not adjusted for multiple comparisons. Intervals that include zero leave a difference unresolved and do not establish equivalence.
Resource and cost scope
Measured resource cost applies frozen, versioned prices to recorded model-token quantities and sandbox CPU and memory measurements. A resource quantity or price that is unavailable remains unknown. The reported scope is not total economic cost, collected revenue, or an invoice.
Reasoning output tokens are unavailable for all 150 Claude Haiku 4.5 attempts. The aggregate output quantity is retained where recorded, but the reasoning-token partition remains unknown and is not estimated. Provider direct costs, customer ledger debits, and measured resource cost remain separate evidence.
External tool and API fees, storage, network, human effort, platform overhead, and other unmeasured costs are outside this analysis. Sandbox counters are resource telemetry, not independent invoice evidence. Cost per success describes only the measured priced-resource scope for this cohort.
Interpretation and limits
The observed quality and cost comparisons apply only to these tasks, configurations, execution controls, and frozen price basis. They do not establish universal model rankings, causal savings, or an invoice-equivalent cost. Unsuccessful-attempt spend summaries are descriptive heuristics, not estimates of avoidable cost.
The public release is bound to canonical source evidence SHA-256 98e6afe2787a8e498383d7d2fac00a38198c6b000bc4987a2ce3725785c13772. The 22-page paper PDF has SHA-256 bc5c4e2870b346286fe5272da4baed58e3465fe6e22bc4ee45a7e183423f31d5. The reproduction capsule has SHA-256 31a56d0dcc5e0e3e6d319cd642206d62d5135f3a3fe481c0750915d4cbbda2d9. Offline reproduction matched the 23 published derived outputs.
CostBench v3 is a separate Dev self-service publishing exercise. It does not reproduce or extend the frozen v2 cohort.