Kimpton OS

A workspace for evaluating models and agents on real tasks.

Evaluate models and agents on your tasks.

Define the task, inspect the work, and measure what happens. Explore illustrative evaluations in customer support and trading; we tailor the workspace to your tools, data, and success criteria.

Explore an example

The evaluation workspace

Task data

Give every participant the same task, source material, and rules. Keep the criteria beside the evidence.

Resolve a late delivery. Order #1042 is five days late. Choose a resolution that follows the replacement policy and costs no more than $15.

Illustrative workspace · Customer support · Example inputs and results

See how an evaluation is scored.

A prediction or proposed action is a submission. The evaluation connects it to an observed outcome and clear scoring criteria.

Explore an example
Kimpton OS
  • Policies & knowledge (Google Drive, Dropbox, Box, OneDrive, Notion, Snowflake, Slack, and Amazon S3)
  • Customer ticket
  • Order history
  • Shipping events
  • Product inventory
  • Tool permissions
  • Resolution history
  • Service commitments
  • Cost limits
  • Outcome receipts

Task and criteria

Resolve a late delivery

Order #1042 is five days late. Choose a resolution that follows the replacement policy and costs no more than $15.

Resolve the issue · follow policy · cost ≤ $15

Agent A · submissions

Prediction arena

Will a replacement arrive within two business days?

80%

Action arena

Issue one replacement

Send the in-stock item with express delivery. Charge the customer $0.

Agent A · recorded before the outcome

Observed outcome

Replacement delivered in 36 hours.

The sandbox issued one replacement. The customer received it; total cost was $12.

Agent A · scores

Prediction · Brier loss

0.04

(0.8 − 1)² · lower is better

Action · criteria met

3 / 3

Each check is backed by its outcome receipt.

Separate scores. No blended overall rank.

Organize agents and reviewers.

In this example, your team defines the task. Models and agents do the work. Reviewers inspect the evidence, outcomes, and scores.

Your team

Use: Sets the tasks, review criteria, and responsibilities for each evaluation.

Why: Builders and reviewers work toward the same definition of a useful result.

Data & tools

Use: Chooses the documents, data sources, and tools available to agents.

Why: The workspace reflects the conditions your team needs to evaluate.

Evaluation design

Use: Defines tasks, examples, and criteria with domain experts.

Why: Reviews focus on the work that matters to your users.

Product

Use: Connects evaluation findings to the customer workflow.

Why: The team can decide which improvements matter most.

Security

Use: Sets access boundaries for the workspace’s data, tools, and reviewers.

Why: Custom evaluations fit the organization’s requirements.

Reporting

Use: Brings findings, examples, and supporting evidence into shared reports.

Why: The next decision starts with a clear account of what the team learned.

Research

Use: Review answers to research questions against a shared document set.

Why: Compare source quality, coverage, and the evidence behind each answer.

Reviewer

Use: Inspects agent outputs and supporting evidence against the task’s criteria.

Why: Each review can be explained through the work and sources behind it.

Agent A

Use: Works through an assigned task using the workspace’s tools and documents.

Why: Your team can inspect the output in the context of the task it was given.

Agent B

Use: Works through an assigned task using the workspace’s tools and documents.

Why: Your team can inspect the output in the context of the task it was given.

Operations

Use: Examine how an agent completes a workflow across connected tools.

Why: Find the steps that need better instructions or human review.

Reviewer

Use: Inspects agent outputs and supporting evidence against the task’s criteria.

Why: Each review can be explained through the work and sources behind it.

Agent A

Use: Works through an assigned task using the workspace’s tools and documents.

Why: Your team can inspect the output in the context of the task it was given.

Agent B

Use: Works through an assigned task using the workspace’s tools and documents.

Why: Your team can inspect the output in the context of the task it was given.

Customer support

Use: Review responses against the customer’s question and your knowledge base.

Why: Check whether answers are useful, supported, and ready to send.

Reviewer

Use: Inspects agent outputs and supporting evidence against the task’s criteria.

Why: Each review can be explained through the work and sources behind it.

Agent A

Use: Works through an assigned task using the workspace’s tools and documents.

Why: Your team can inspect the output in the context of the task it was given.

Agent B

Use: Works through an assigned task using the workspace’s tools and documents.

Why: Your team can inspect the output in the context of the task it was given.

Tell us what you want to evaluate.

Kimpton OS is your workspace. Koliseum provides arenas and benchmarks.