Review and label tasks
Preview
Apply a grader's criteria to responses and decide which tasks are useful for learning.
A task describes what the model should attempt. A response records one attempt. A human label records your judgment of that response against specific criteria. The same task can have several responses with different scores.
The preview's response review loads the exact attached grader criteria and scale. It retains that grader/rubric identity with ratings while keeping corrections and training eligibility as separate decisions.
Review a response
- Open a Model's Tasks page and choose Review on a recorded response.
- Read the attached grader's rubric or checks and its scoring scale.
- Compare the automatic result with the evidence. Decide whether the model failed, the grader was wrong, or the task lacks enough information.
- Record a label and explanation against the relevant criterion. Mark cases you cannot assess rather than assigning an arbitrary score.
- Under Training use, decide separately whether the task is suitable for training, then choose Save review.
For a support judge, a criterion might be “Checks refund eligibility before promising a refund.” A failed label should explain which promise or missing check supports the judgment. That gives a judge improvement a concrete target.
Distinguish the review outcomes
| Finding | Next step |
|---|---|
| Model failed; grader correctly detected it | Consider the task for training with the current grader |
| Grader gave the wrong score | Use the label to check/refine the grader before relying on that score |
| Task input or expected result is wrong | Correct the task, then evaluate the revised example |
| You can supply a better answer | Save a correction/reference and review its intended use |
| Evidence is insufficient or reviewers disagree | Keep uncertainty visible and resolve it before treating the label as authoritative |
Approving a task for RL means allowing fresh attempts at it. It does not approve copying an incorrect response. A suggested answer is also different from a human score: one shows what to say; the other judges a particular attempt.
Use labels to improve scoring
Selected labels can become examples in the judge prompt or test cases for the judge. Record which role each example serves. A test of the judge hides the human score from it and compares the returned score with your judgment.
Start with diverse successes, failures and ambiguous cases. Ten or fifteen examples can reveal useful disagreements, but the count alone does not prove that the judge is reliable.
From a retained rating, choose Improve grader, select the reviewed labels to use as development examples, and continue into the candidate editor. Add separate independent validation fixtures and check both the published grader and candidate on the unchanged fixtures. Saving a label or candidate alone does not change the active grader or train the model.
Next: Improve an LLM judge or use a code verifier.