Evaluation
Compare a baseline and candidate against a frozen Taskset.
Evaluation
Evaluation runs a defined model or behavior against a frozen Taskset and records the result, grader output, version identity, and review context.
Before you compare
Use the same Taskset version, grading rules, and meaningful resource limits for the baseline and candidate. Inspect representative failures as well as the aggregate score; a higher score cannot justify a regression in a critical task.
After an evaluation
Choose one of three outcomes: keep the current release, investigate a bounded Harness change, or approve a candidate for a further training/activation gate. Evaluation does not silently deploy a candidate.
Next: Training methods.