Experiments
Preview
Run fresh attempts or score recorded evidence against pinned dataset and grader releases.
Open Experiments to create executions and inspect their results. You can also start an experiment from a Dataset or selected recorded activity. An experiment retains its target, dataset release, graders, and execution context so the result can be inspected later.
Choose what you are measuring
| Mode | What happens | Use when |
|---|---|---|
| Fresh model execution | The selected model generates new responses | Measure the model on held-out tasks |
| Model with Harness or agent behavior | Attempts use the selected published behavior and permitted runtime | Measure the complete agent workflow |
| Recorded responses | Existing attempts are graded without regenerating them | Review captured work or validate a grader |
Available modes depend on the dataset, target, runtime, and captured evidence. A recorded result does not establish that a fresh model can reproduce it.
Start an experiment
- Select the workspace and optional Project, then create an experiment or open the Dataset's experiment action.
- Choose the dataset and tasks. If using a draft, review and publish the dataset release required by the setup.
- Select the target and execution mode. For a model request, review prompt, token limit, and sampling settings; for agent behavior, select the exact published Harness or Agent source.
- Select the grading rules and review the execution limits shown by the setup.
- Start the experiment and open its execution to follow progress and results.
Read the result
Inspect per-task responses, tool evidence, grades, and errors as well as the aggregate score. Infrastructure failure or unavailable grading evidence is not a passing result. Use Human review when scores need interpretation or the judge disagrees with the evidence.
Compare a baseline and candidate using the same dataset release, grader versions, and meaningful limits. If the grader changes, recheck the baseline under the new grader before claiming improvement. Keep judge prompt examples separate from independent judge validation and model evaluation tasks.
Continue into training
A completed experiment can open Train from completed Experiment. The handoff retains the exact target, dataset, result, and grader releases. Training examples are selected and approved separately: a result does not itself authorize copying responses, spending money, or activating a model. See Training and Models.