Back to Docs/Experiments

Experiments

Preview

Run fresh attempts or score recorded evidence against pinned dataset and grader releases.

Open Experiments to create executions and inspect their results. You can also start an experiment from a Dataset or selected recorded activity. An experiment retains its target, dataset release, graders, and execution context so the result can be inspected later.

Choose what you are measuring

ModeWhat happensUse when
Fresh model executionThe selected model generates new responsesMeasure the model on held-out tasks
Model with Harness or agent behaviorAttempts use the selected published behavior and permitted runtimeMeasure the complete agent workflow
Recorded responsesExisting attempts are graded without regenerating themReview captured work or validate a grader

Available modes depend on the dataset, target, runtime, and captured evidence. A recorded result does not establish that a fresh model can reproduce it.

Start an experiment

  1. Select the workspace and optional Project, then create an experiment or open the Dataset's experiment action.
  2. Choose the dataset and tasks. If using a draft, review and publish the dataset release required by the setup.
  3. Select the target and execution mode. For a model request, review prompt, token limit, and sampling settings; for agent behavior, select the exact published Harness or Agent source.
  4. Select the grading rules and review the execution limits shown by the setup.
  5. Start the experiment and open its execution to follow progress and results.

Read the result

Inspect per-task responses, tool evidence, grades, and errors as well as the aggregate score. Infrastructure failure or unavailable grading evidence is not a passing result. Use Human review when scores need interpretation or the judge disagrees with the evidence.

Compare a baseline and candidate using the same dataset release, grader versions, and meaningful limits. If the grader changes, recheck the baseline under the new grader before claiming improvement. Keep judge prompt examples separate from independent judge validation and model evaluation tasks.

Continue into training

A completed experiment can open Train from completed Experiment. The handoff retains the exact target, dataset, result, and grader releases. Training examples are selected and approved separately: a result does not itself authorize copying responses, spending money, or activating a model. See Training and Models.