Back to Docs/Training

Training

Preview

Evaluate tasks, review results, improve a grader, and train a candidate model.

Model training changes model weights. Continual learning repeats that process with new approved tasks while checking that the model retains earlier skills. The grader supplies the scores that guide reinforcement learning.

Start with Continual Learning to import OpenClaw, Hermes, or another local agent, choose graders, and schedule recurring updates. The guide also explains on-policy training and holdout evals.

Open Training under Reinforcement Learning. The page collects training runs and preparations. A new run starts by selecting or creating a training configuration; it does not start compute merely because you opened the form or chose a model.

Prepare the evaluation and reward

Your starting pointGuide
Build or import tasksDatasets
Establish a baselineExperiments
Label good and bad responsesReview and label tasks
Success requires judgmentImprove an LLM judge
Success can be checked with rules or testsUse a code verifier
Scoring criteria changedChange a grader safely

Keep training tasks separate from held-out evaluation. Resolve grader and runtime readiness problems before approving a run; importing a benchmark does not establish that it can execute or train with the selected target.

Create a training run

  1. Choose New training run and select a saved training configuration or Create training configuration. Review the starting model, training dataset, reward binding, recipe, and held-out evaluation release.
  2. Choose Run once or the recurring-learning option. For a one-off GRPO run, select or compose an approved reward-training batch compatible with the configuration's reward binding. Selected captured tasks have their own preparation and approval flow.
  3. Give the run a name and set its maximum spend. Review the prepared inputs and readiness result before starting it.
  4. Choose whether to attach a post-training evaluation. Its manual or automatic mode, selected experiment, checks, and spending limit are separate from the training inputs. Omitting a check does not make the candidate accepted.
  5. Start the reviewed preparation, then open the run to follow execution, cost, failures, and resulting artifacts.

Starting from Train from completed Experiment retains the exact source result, target, dataset, and grader releases. Select and approve the training examples separately; the experiment handoff does not make held-out evaluation tasks into training data.

Review runs and candidates

A retained preparation can be revisited or discarded through its controls. After submission, inspect the job's state before retrying a failed request so you do not create duplicate work. Use the available cancellation controls for submitted work rather than treating closing the editor as cancellation.

Review the resulting version in Models, compare it with the baseline, and explicitly decide whether to use it. Deployment through Serving is a further action. Configure recurring runs through Continual Learning.

A continual-learning cycle

  1. Collect tasks from an approved source, imported Taskset or reviewed work.
  2. Evaluate attempts and inspect the results. Human labeling is optional when the automatic grader already checks the intended outcome reliably.
  3. Approve useful tasks for training. A failed response can identify a useful task; approval does not mean copying that failed response.
  4. If the grader misunderstood your criteria, validate and adopt a revised grader before starting the next run.
  5. Train: the current model generates fresh attempts, the selected grader scores them, and the optimizer uses those rewards to update the model.
  6. Evaluate the candidate on separate held-out tasks and compare it with the current model under the same scoring rules.
  7. Apply the candidate decision: review it yourself, or use the automatic activation authority saved in a conversation learning policy.

With scheduled learning enabled, runs use the approved configuration and eligible tasks. Task arrival does not change the judge prompt. No eligible batch means no training job. Candidate acceptance remains a separate check. Conversation learning can require your approval or activate a passing candidate on an explicitly authorized serving target; see Continual Learning for setup.

What the model learns from

ComponentReceives
Model making an attemptTask instructions, permitted context, tools and environment observations
GraderThe response or actions, scoring rules and permitted reference information
RL optimizerGenerated attempts and their rewards
Evaluation reportResults for comparison and human inspection

Human labels can help develop and test the grader. They are not automatically inserted into the model's prompt. Private reference answers and judge examples stay with the grader. Saving a corrected answer does not by itself train the model to reproduce it.

Start from a benchmark

A benchmark Profile can supply tasks, environment instructions and native scoring. Inspect the source, selected cases, grader and execution requirements before running it. You should not have to recreate scoring that the benchmark already supplies.

For example, τ Retail tasks can test whether a customer request produced the correct database outcome. Its native verifier can check that outcome without asking you to label every conversation. A conversation that sounds successful can still fail the state check.

Check the selected release and runtime readiness rather than treating a visible benchmark Profile as a trainable configuration. A local verifier check does not qualify execution in a different hosted environment.

Training methods

Hosted continual learning currently uses GRPO, a reinforcement-learning method. Supervised training from corrected demonstrations and preference optimization are different methods; recording a correction or preference does not select either automatically. A Refiner proposal also does not launch model training.

See Evaluation for fair comparisons and Continuous learning for the separate Harness refinement and recurring-review loop.