OpenPond
Back to Docs/Evaluation

Evaluation

Compare a baseline and candidate against a frozen Taskset.

Evaluation

Evaluation runs a defined model or behavior against a frozen Taskset and records the result, grader output, version identity, and review context.

Before you compare

Use the same Taskset version, grading rules, and meaningful resource limits for the baseline and candidate. Inspect representative failures as well as the aggregate score; a higher score cannot justify a regression in a critical task.

After an evaluation

Choose one of three outcomes: keep the current release, investigate a bounded Harness change, or approve a candidate for a further training/activation gate. Evaluation does not silently deploy a candidate.

Next: Training methods.