Continual Learning
Preview
Import local agent conversations, choose graders, and schedule training with independent holdout evaluations.
Connect the agents you already use, define what a good answer looks like, and choose when OpenPond can learn from their work. Each cycle evaluates new conversations, prepares verified training tasks, trains a candidate, and checks it against the current model before it can become active.
The setup has three steps: import conversations → pick graders → schedule learning. This guide covers both setup and the evaluations behind the loop.
1. Import your conversations
Run the command for your agent on the computer where its conversations are stored. Install the OpenPond CLI first if you do not already have it:
npm install -g openpond@latestChoose a source below. These commands import the past week and keep syncing new activity after you approve the connection.
Import from
openpond import connect --source openclaw --range weekFollow the sign-in and source approval prompts. Without a destination flag, the importer uses your personal Imported conversations Project. To send activity to a particular Project and workspace, open Connected agents, choose the source, and copy its configured command. That command includes your Project and workspace IDs. Connect each additional source to the same Project if you want one learning policy to use them together.
Use --range day or --range all to change the initial history window. Add
--once for a single import. If you have multiple installations or profiles,
use openpond import discover and pass the intended directory with
--source-path /path/to/agent-data.
The importer installs a background collector, so closing the terminal monitor does not stop ongoing sync. The source computer must be awake and connected for new local activity to arrive. Check it with:
openpond import statusInspect the imported conversations and tool evidence in OpenPond. Importing makes that evidence available; enabling a learning policy separately authorizes its use. OpenPond's own hosted Chat and Work activity can be selected through the Ponder Pal connection without a local import command.
2. Pick your graders
Select the destination Project, open Continual Learning, and create a conversation learning policy. You can also begin from a connected agent's Continual learning control. In Sources & rubrics, select the connections that should contribute and the evaluation rubrics that describe success.
The default conversation graders cover:
| Grader | What it checks |
|---|---|
| Correctness | Whether the response follows the request and produces a correct, complete result. |
| Groundedness | Whether material claims are supported by the conversation and recorded tool evidence. |
| Conciseness | Whether the response includes the necessary detail without irrelevant material or repetition. |
Use task-specific graders when those better describe the work. Complete text conversation evaluation uses calibrated LLM judges. If you choose an existing training configuration, its pinned graders are shown instead; code verifiers and other task-specific checks belong to that configuration.
Review passing and failing examples before relying on a grader. A short answer must still be correct, and a plausible answer is not evidence of a verified fact. Every selected check must pass the policy's admission threshold; a strong conciseness score cannot compensate for failed correctness.
Saved policies pin grader versions. Importing more conversations does not rewrite a rubric. When your criteria change, validate a new grader version and compare models under the same updated scoring rules.
3. Choose when and how to learn
In Model, choose the training configuration and the amount of automation:
| Mode | What happens |
|---|---|
| Evaluate only | Score and verify conversation evidence without launching weight training. |
| Train automatically, approve candidate | Train eligible batches and compare candidates; you decide whether to accept the result. |
| Train and activate automatically | Train and compare, then activate a passing candidate on the explicitly authorized OpenPond serving target. |
Training needs a compatible model, reviewed training setup, executable rewards, and an independent acceptance plan. Importing activity from a proprietary model does not make that model's weights available for training. Choose an available trainable model as the target.
An optional Answer improvement model proposes corrections for failed answers. Separate grading checks those proposals before they can be admitted. Without it, the workflow uses recorded answers and any supported reference answers from the selected task setup.
In Schedule & spending, set From, Until, and Timezone. These fields define a recurring nightly start window, rather than a one-time date. Check Next scheduled run on the saved policy for the actual date and time. A run that starts within the window can finish afterward; Run now starts immediately.
Set maximum spending per run, per day, and per month. Evaluation, answer improvement, verification, and training consume the authorized budget. In Advanced, review the minimum training examples, maximum tasks, correction attempts, GPU duration, selection limits, and acceptance thresholds. Choose whether exceptions block the cycle or allow the independently verified remainder to proceed.
For automatic activation, select a registered serving target and save the configuration paused. Authorize its exact settings in OpenPond's Model learning page, then enable the policy. Activation also runs a runtime canary; a failed canary restores the recorded prior binding. Availability depends on your workspace's training access and target readiness.
What happens in each cycle
- Collect: find new completed conversation snapshots from the selected sources. Already processed unchanged snapshots are skipped.
- Evaluate: score the recorded answers and tool evidence with the pinned graders. Missing context, abstentions, and failed checks remain visible.
- Prepare: retain verified successes and, when configured, propose and verify corrections. Keep the original answer alongside its proposed correction.
- Select: admit eligible task families within the source, diversity, batch, and budget limits. Protected evaluation families cannot enter training.
- Train: generate fresh attempts from the current training policy and update model weights using the configured rewards.
- Compare: run the baseline and candidate on the same fresh holdout window with the same evaluator versions and execution settings.
- Decide: reject an unsuccessful candidate, request your approval, or activate a passing candidate within the authorization you saved.
The schedule provides an opportunity to learn, not a guarantee of a training job. Insufficient verified examples, exhausted budgets, missing authority, or an exhausted holdout pool prevent a new training cycle from proceeding.
How graders, prompts, and reinforcement learning fit together
A grader defines how to score an attempt. An LLM judge implements that scoring with a rubric and a selected judge model; a code verifier implements it with executable checks. A training configuration binds compatible graders to the reward used by the optimizer. The independent acceptance plan also pins the graders used to evaluate the candidate.
On-policy means the model being trained generates fresh attempts under its current policy. It does not mean inserting the private judge rubric into the model's task prompt. Task instructions and allowed context tell the model what to do; the grader scores what it actually did.
| Part of the loop | Receives |
|---|---|
| Model generating an attempt | Public task input, permitted conversation context, tools, and environment observations. |
| LLM judge or verifier | The attempt, scoring rules, and authorized reference or tool evidence. |
| RL optimizer | Fresh attempts and their rewards, under the selected training recipe. |
| Independent evaluation | Baseline and candidate attempts on protected tasks, scored under the pinned acceptance plan. |
The hosted continual-learning training path uses GRPO (Group Relative Policy Optimization). The model generates a group of attempts for a task, the reward function scores them, and their relative rewards guide the weight update. This gives the model feedback about which behaviors worked on that task.
Imported answers help identify tasks and evaluate past behavior. A verified correction can supply reference evidence for a supported task; it does not silently switch the job to supervised imitation. The next training attempts are generated afresh. Judge-only qualitative evidence keeps a judge-validated label, so it can be distinguished from reference-backed verification.
How holdout evals protect the result
Training rewards answer “did this attempt satisfy the grader?” Holdout evals answer “did the resulting model improve on tasks it did not train on?” A higher training reward alone cannot establish that.
OpenPond keeps three purposes separate:
| Split | Purpose |
|---|---|
train | Tasks used to generate attempts and update weights. |
validation | Validation used by the training recipe. |
frozen_eval | Protected tasks reserved for independent baseline/candidate comparison. |
The conversation loop protects evaluation families before discovering and constructing training examples. Family and input identities prevent the same protected work from being admitted through another source or a correction. Frozen evaluation tasks and answers stay out of the training artifact.
When a cycle has enough training evidence, it reserves a fresh window of unused protected families from the configured evaluation pool. It records that exposure, including when a cycle later fails. If too few unused families remain, the loop requires more protected evaluation data before another cycle; it does not weaken the holdout or silently recycle the same test window.
Baseline and candidate must have sufficient scored cases on the same population. The candidate must meet the required quality improvement margin and every required floor and regression check. Missing results or insufficient evidence are inconclusive, not a pass. Candidate acceptance and serving activation are separate decisions, even when your policy automates both.
Follow progress or pause learning
The policy's Overview shows the next run, last cycle, verified examples, and spending. Use Sources, Evaluation, Data, Training, and Activity to inspect inputs, original and corrected answers, verification, training runs, and candidate outcomes. Unresolved cases that need your decision appear in the Inbox.
Pause the policy to prevent new learning cycles. Use Stop current cycle for
work already in progress. Turning off a source's learning participation is
separate from pausing its importer; to pause sync, take the connection ID from
openpond import status and run openpond import pause <connection-id>.
For one-off runs, see Training. For how to build and validate scoring rules, see Graders and LLM judges.