Back to Docs/Evaluation overview

Evaluation overview

Measure models and agents with datasets, graders, and experiments.

Evaluation answers whether a model or agent performs a defined task well. It also lets you check whether the grader measures the intended outcome. Neither operation changes model weights.

Follow the Console workflow

  1. Open Datasets, create or import examples, and inspect their input, references, provenance, and execution requirements.
  2. Open Graders, define the scoring criteria, and validate them against known passing, failing, and ambiguous examples.
  3. Open Experiments, select the dataset and target, review execution mode and limits, and start the comparison.
  4. Inspect case-level results, scores, errors, and retained evidence before deciding what to improve.
ResourcePurpose
DatasetsVersioned tasks and their evaluation contract
GradersRepeatable scoring, including required gates
ExperimentsExecutions and results for a chosen target and dataset
EnvironmentsThe runtime contract required to reproduce attempts
Human reviewLabels and explanations for recorded responses

Publish reviewed releases before comparing targets. Keep dataset and grader versions fixed across a baseline and candidate comparison. Separate held-out tasks from training and from examples used to improve the judge.

If an experiment identifies a repeatable failure, choose the smallest effective change: fix task data, improve a grader, adjust the Harness, or prepare a separately approved Training run.