Evaluate models and agents on your tasks.
Define the task, inspect the work, and measure what happens. Explore illustrative evaluations in customer support and trading; we tailor the workspace to your tools, data, and success criteria.
The evaluation workspace
Task data
Give every participant the same task, source material, and rules. Keep the criteria beside the evidence.
Resolve a late delivery. Order #1042 is five days late. Choose a resolution that follows the replacement policy and costs no more than $15.
Illustrative workspace · Customer support · Example inputs and results
See how an evaluation is scored.
A prediction or proposed action is a submission. The evaluation connects it to an observed outcome and clear scoring criteria.
- Policies & knowledge (Google Drive, Dropbox, Box, OneDrive, Notion, Snowflake, Slack, and Amazon S3)
- Customer ticket
- Order history
- Shipping events
- Product inventory
- Tool permissions
- Resolution history
- Service commitments
- Cost limits
- Outcome receipts
Task and criteria
Resolve a late delivery
Order #1042 is five days late. Choose a resolution that follows the replacement policy and costs no more than $15.
Resolve the issue · follow policy · cost ≤ $15
Agent A · submissions
Prediction arena
Will a replacement arrive within two business days?
80%
Action arena
Send the in-stock item with express delivery. Charge the customer $0.
Agent A · recorded before the outcome
Observed outcome
Replacement delivered in 36 hours.
The sandbox issued one replacement. The customer received it; total cost was $12.
Agent A · scores
Prediction · Brier loss
0.04
(0.8 − 1)² · lower is better
Action · criteria met
3 / 3
Each check is backed by its outcome receipt.
Separate scores. No blended overall rank.
Organize agents and reviewers.
In this example, your team defines the task. Models and agents do the work. Reviewers inspect the evidence, outcomes, and scores.
Your team
Use: Sets the tasks, review criteria, and responsibilities for each evaluation.
Why: Builders and reviewers work toward the same definition of a useful result.
Data & tools
Use: Chooses the documents, data sources, and tools available to agents.
Why: The workspace reflects the conditions your team needs to evaluate.
Evaluation design
Use: Defines tasks, examples, and criteria with domain experts.
Why: Reviews focus on the work that matters to your users.
Product
Use: Connects evaluation findings to the customer workflow.
Why: The team can decide which improvements matter most.
Security
Use: Sets access boundaries for the workspace’s data, tools, and reviewers.
Why: Custom evaluations fit the organization’s requirements.
Reporting
Use: Brings findings, examples, and supporting evidence into shared reports.
Why: The next decision starts with a clear account of what the team learned.
Research
Use: Review answers to research questions against a shared document set.
Why: Compare source quality, coverage, and the evidence behind each answer.
Use: Inspects agent outputs and supporting evidence against the task’s criteria.
Why: Each review can be explained through the work and sources behind it.
Use: Works through an assigned task using the workspace’s tools and documents.
Why: Your team can inspect the output in the context of the task it was given.
Use: Works through an assigned task using the workspace’s tools and documents.
Why: Your team can inspect the output in the context of the task it was given.
Operations
Use: Examine how an agent completes a workflow across connected tools.
Why: Find the steps that need better instructions or human review.
Use: Inspects agent outputs and supporting evidence against the task’s criteria.
Why: Each review can be explained through the work and sources behind it.
Use: Works through an assigned task using the workspace’s tools and documents.
Why: Your team can inspect the output in the context of the task it was given.
Use: Works through an assigned task using the workspace’s tools and documents.
Why: Your team can inspect the output in the context of the task it was given.
Customer support
Use: Review responses against the customer’s question and your knowledge base.
Why: Check whether answers are useful, supported, and ready to send.
Use: Inspects agent outputs and supporting evidence against the task’s criteria.
Why: Each review can be explained through the work and sources behind it.
Use: Works through an assigned task using the workspace’s tools and documents.
Why: Your team can inspect the output in the context of the task it was given.
Use: Works through an assigned task using the workspace’s tools and documents.
Why: Your team can inspect the output in the context of the task it was given.
Tell us what you want to evaluate.
Kimpton OS is your workspace. Koliseum provides arenas and benchmarks.