Evaluation overview
Measure models and agents with datasets, graders, and experiments.
Evaluation answers whether a model or agent performs a defined task well. It also lets you check whether the grader measures the intended outcome. Neither operation changes model weights.
Follow the Console workflow
- Open Datasets, create or import examples, and inspect their input, references, provenance, and execution requirements.
- Open Graders, define the scoring criteria, and validate them against known passing, failing, and ambiguous examples.
- Open Experiments, select the dataset and target, review execution mode and limits, and start the comparison.
- Inspect case-level results, scores, errors, and retained evidence before deciding what to improve.
| Resource | Purpose |
|---|---|
| Datasets | Versioned tasks and their evaluation contract |
| Graders | Repeatable scoring, including required gates |
| Experiments | Executions and results for a chosen target and dataset |
| Environments | The runtime contract required to reproduce attempts |
| Human review | Labels and explanations for recorded responses |
Publish reviewed releases before comparing targets. Keep dataset and grader versions fixed across a baseline and candidate comparison. Separate held-out tasks from training and from examples used to improve the judge.
If an experiment identifies a repeatable failure, choose the smallest effective change: fix task data, improve a grader, adjust the Harness, or prepare a separately approved Training run.