Back to Docs/Datasets

Datasets

Author, import, review, and publish tasks for experiments and training.

Open Datasets in Evaluation. A Dataset contains task inputs, reference information, provenance, grading dependencies, and any execution environment needed to reproduce an attempt. Published Taskset releases are the immutable evaluation contract behind these datasets.

Create or import examples

Create a working dataset, author examples, or import a selected dataset release. You can also select captured activity from Connected agents, review the import, and retain its source context. Marketplace acquisition is described separately in Marketplace.

For each task, check the input, expected outcome, permitted context, source, and split. Keep reference answers and private grader state separate from what the model sees. Label corrected, synthetic, and source-derived examples clearly.

Inspect the dataset

Use the dataset's Tasks, Graders, Experiments, and Versions views to inspect examples, attached scoring, executions, and published releases. Check the selected version when viewing tasks or graders: an edit to a working draft does not rewrite a previous release or experiment.

Attach the Graders needed to measure success. If the task needs tools or a simulator, inspect its Environment and readiness before execution. Publish the reviewed dataset so experiments can pin an exact release.

Evaluate before training

Start an Experiment from the dataset and review the selected task population, target, scoring, and limits. An imported benchmark is not automatically executable in every environment.

Human review and approval for training remain separate. A failed response can identify a useful training task without becoming a demonstration to copy. Keep related examples out of both training and held-out splits, then use the approved tasks in Training.