Continual Learning From the Agents You Already Use

October 4, 2026
6 min read
OpenPond
Continual LearningReinforcement LearningEvaluation

Your agent's everyday work contains useful signals: requests it completed well, answers that missed the point, and corrections that show what should have happened. OpenPond brings that activity into a recurring learning process. Connect your agents, choose how to measure quality, and set a time window for learning. The system evaluates new work, prepares verified tasks, trains a candidate model, and checks whether the update deserves to be used.

You can start with evaluation alone, let training run while keeping candidate approval manual, or authorize automatic activation after independent checks. The same process retains the original conversations, scoring decisions, and model lineage so you can inspect what changed.

1. Connect the agents you already use

Install the CLI with npm install -g openpond@latest, then run the command for your source on the computer where its conversations live:

Import from

Run locally in your terminal
openpond import connect --source openclaw --range week

After sign-in and approval, the importer brings in the past week and keeps syncing through a background collector. Keep that computer online for new activity to arrive. Add --once for a single import, or change the initial history window to day or all.

These commands use your personal Imported conversations Project. To connect several agents to a chosen Project, copy each source's configured command from Connected agents. It includes the destination Project and workspace. OpenPond's hosted Chat and Work activity is available through Ponder Pal without a local importer.

Importing an agent's activity supplies evidence. You separately choose the model to train and authorize learning from the selected sources. A conversation produced by a proprietary model does not give OpenPond access to that model's weights.

2. Choose what good looks like

A learning loop needs a useful definition of success. In Sources & rubrics, select the graders that fit the work. The default conversation graders cover correctness, groundedness, and conciseness. A specialist task can use its own rubric or executable verifier through a compatible training configuration.

These criteria work together. A concise answer that gets the calculation wrong should not enter training because its average score looks acceptable. Every selected admission check must pass. When evidence is missing or a judge cannot verify a claim, the workflow keeps that uncertainty visible.

An optional answer improvement model can propose corrections. The proposal goes through separate grading before it is eligible; the original answer stays in the record. Qualitative judge checks carry a judge-validated label, distinct from reference-backed verification. Unsupported or unresolved cases can go to review, and you decide whether they block the cycle or allow the verified remainder to proceed.

3. Set the schedule and automation mode

In Schedule & spending, choose a nightly start window and timezone, then review the actual next date and time on the saved policy. A run can finish after the window closes. Run now starts immediately.

Set per-run, daily, and monthly spending caps, plus limits on batch size, correction attempts, and GPU duration. The schedule does not force an empty training run: there must be enough eligible evidence and budget to proceed.

Choose the mode that matches the authority you want to grant:

ModeResult
Evaluate onlyRecurring evaluation and verification, without weight training.
Train automatically, approve candidateAutomatic training and comparison, followed by your acceptance decision.
Train and activate automaticallyAutomatic training, comparison, and activation on an explicitly authorized OpenPond target.

Automatic activation is saved paused until its exact serving settings are authorized in OpenPond's Model learning page. Passing the quality checks is still required. A runtime canary checks the activated model, and a failed canary restores the prior binding. Training options depend on workspace access and a compatible, ready target.

How reinforcement learning updates the model

The hosted continual-learning path uses GRPO, or Group Relative Policy Optimization. For an eligible task, the model generates a group of fresh attempts. The configured reward function scores them, and the optimizer uses their relative rewards to update model weights. Repeating that process provides feedback about which behaviors satisfy the task's criteria.

This is on-policy learning: the attempts come from the policy currently being trained. Imported conversations help identify useful tasks and assess past behavior; they are not simply replayed as answers for the model to copy. A verified correction can become reference evidence in the task setup, while the model still generates new attempts during RL.

The grader defines the reward, but its private rubric is not automatically the model's prompt. The model receives task instructions and permitted context. The judge receives the attempt, its scoring rubric, and authorized reference evidence. The optimizer receives attempts and rewards. Keeping these roles separate lets us describe what the model learned without confusing it with what the evaluator was allowed to see.

Grader versions are pinned when the configuration is saved. New conversations do not rewrite the judge. If you change the criteria, validate the new grader and compare models under those same revised rules.

Holdouts check whether the update actually helped

A higher training reward is only part of the story. The candidate also needs to improve on work it did not train on, without unacceptable regressions. OpenPond separates training tasks, optimizer validation, and protected frozen_eval tasks used for independent comparison.

The conversation loop protects evaluation families before constructing training examples. Family and input checks keep the same protected work from returning through a different source or a proposed correction. Once a cycle is ready to train, it reserves a fresh window of unused protected families from the configured evaluation pool. Exposure is recorded even if the cycle later fails. When that pool runs out, another cycle needs more protected data.

The baseline and candidate face the same window, grader versions, and execution settings. Acceptance requires sufficient scored cases, the configured quality improvement margin, and all required floors and regression checks. Missing results or too little evidence produce an inconclusive outcome. They do not turn into a pass because training completed successfully.

A recurring process you can inspect

Each cycle moves from collection through evaluation, verification, selection, training, comparison, and a candidate decision. The policy's Data and Activity views retain the evidence and outcomes. Exceptions requiring your input appear in the Inbox; routine automatic work remains visible in the policy history.

Continual learning is a sequence of bounded model updates. It can handle recurring work automatically while preserving the decisions that matter: which sources are allowed, how quality is measured, how much can be spent, and when a candidate may replace the active model.

Follow the continual learning guide for the complete setup, import controls, grader explanation, and holdout behavior.