Terminal · Coding · Multimodal

Agent training data that proves itself.

We provide terminal, coding, and multimodal training data for agents. Every task ships with its own reference solution and executable verifier, and every trajectory is verified and rewritten into one consistent harness, calibrated to the difficulty your model needs next.

  • Verified by execution
  • Calibrated to your checkpoint
  • Fresh every month
  1. T1Warm-up
  2. T2Core
  3. T3Hard
  4. T4Expert
  5. T5Frontier

Tasks run on the software practitioners actually use

  • CadQuery
  • OpenSCAD
  • FreeCAD
  • KiCad
  • Blender
  • OpenFOAM
  • MuJoCo
  • URDF
  • Tesseract
  • Coq
  • CMake
  • Git
  • Docker
  • Nuxt
  • Vue
  • Flask
  • SQLite
  • STEP

What we provide

Terminal, coding, and multimodal data.

Terminal work is coding work: to finish a terminal task, the agent writes, runs, tests, and debugs real programs. The same verified environments extend to visual inputs, from engineering drawings to rendered pages.

Terminal

Long-horizon command-line work across real tools, services, builds, and data: hundreds of commands toward one verified end state.

  • Seeded from real recorded terminal sessions
  • Ops, data, build & release workflows
  • RL environments with executable verifiers

Coding

Every terminal task is a coding task. Agents write, test, and debug real programs, from rebuilding scientific software to shipping full-stack apps.

  • Software reconstruction from real program behavior
  • Repository repair & feature work
  • Full-stack apps verified in a browser

Multimodal

Visual inputs, code outputs: engineering drawings, renders, screenshots, meshes, and scientific images.

  • CAD from engineering drawings
  • Web pages and apps from designs
  • Charts, SVG & documents to code
Every task ships with
  • Reference solution
  • Executable verifier
  • Difficulty tier
  • Harness-consistent trajectories

Examples

Watch real software run. Rebuild its behavior.

A sample of our software-reconstruction tasks. The agent sees public runs of real scientific and engineering software, writes a program that reproduces it, and is graded on hidden configurations it never saw. Every visual is rendered from the task's own files.

Rendered fan-duct mesh from a public OpenSCAD run

Spatial, Visual & 3D

OpenSCAD / BOSL2 parametric fan duct

Regenerate a swept fan-duct mount STL whose inlet size, outlet size, bend angle, and rib count change the geometry.

  • Terminal
  • Coding
  • Multimodal
  • 3D mesh
  • OpenSCAD
  • BOSL2
OpenFOAM dam-break mesh drawn from the public polyMesh

Physical & Engineering Simulation

OpenFOAM 2-D dam-break CFD

Reconstruct a 2-D dam-break CFD workflow across water height, mesh, gravity, and simulation time.

  • Terminal
  • Coding
  • Multimodal
  • CFD mesh
  • OpenFOAM
KiCad board with design-rule findings plotted at their positions

Spatial, Visual & 3D

KiCad LED-matrix driver PCB

Reconstruct LED-matrix driver board outputs (DRC report and project files) under row, column, driver, and routing interventions.

  • Terminal
  • Coding
  • Multimodal
  • PCB layout
  • KiCad
Container image layers and attestation from a public run

Software & Formal Systems

Container SBOM and supply-chain attestation

Infer coupled container base-image, package-set, label, entrypoint, and SBOM-scope behavior.

  • Terminal
  • Coding
  • Hadolint
  • Syft
  • Skopeo
Blender cloth scene rotation dial across public scenarios

Spatial, Visual & 3D

Blender cloth flag scene

Reconstruct cloth-scene behavior under wind, pin-group, frame, and subdivision interventions.

  • Terminal
  • Coding
  • Multimodal
  • 3D scene
  • Blender
MuJoCo joint-position trajectories across public scenarios

Physical & Engineering Simulation

MuJoCo / URDF mass-inertia repair

Infer robot state and trajectory behavior under actuator, collision, pose, joint-limit, and target interventions.

  • Terminal
  • Coding
  • MuJoCo
  • URDF/MJCF tooling
Microscopy label masks from public runs

Life & Molecular Sciences

Microscopy cell segmentation and counting

Reconstruct cell segmentation and measurement behavior across channel, illumination, threshold, object size, and measurement feature.

  • Terminal
  • Coding
  • Multimodal
  • Microscopy images
  • CellProfiler
  • scikit-image
  • Bio-Formats
Astronomy FITS cutouts from public runs

Earth & Space Sciences

Astronomy FITS/WCS cutout and photometry

Reconstruct FITS/WCS cutout and photometry behavior under aperture, background, clipping, threshold, and coordinate-frame interventions.

  • Terminal
  • Coding
  • Multimodal
  • FITS images
  • Astropy
  • Photutils
OCR page preview from a public run

Data & Document Workflows

Tesseract table-image OCR

Reconstruct how DPI, font, language, metadata, and page range affect OCR behavior.

  • Terminal
  • Coding
  • Multimodal
  • Document images
  • Pandoc
  • Poppler
  • Tesseract
  • QPDF
Color-pipeline previews from public runs

Spatial, Visual & 3D

Color-space and tone-curve pipeline

Infer a color pipeline across color space, white point, tone curve, demosaicing, and compression quality.

  • Terminal
  • Coding
  • Multimodal
  • Images
  • colour-science
  • ImageMagick
  • scikit-image

Terminal & coding examples

Real terminal work, verified end to end.

Seeded from real recorded terminal sessions, then extended round by round. Every task ships with its environment, a reference solution, and a verifier that passed in a fresh sandbox. Each replay shows commands from the task's own reference solution.

Software engineering

Repair and cross-build an AArch64 assembly program

Fix the assembly source at /app/horas_seg.s so it prompts for integer hours and correctly prints the equivalent seconds (e.g., 2 → 7200 segundos.).

  • 239-line reference solution
  • Verified in a fresh sandbox

Networking

Wire a network namespace with a veth pair

Set up a Linux network namespace t1 with a veth pair.

  • 213-line reference solution
  • Verified in a fresh sandbox

Containers & cloud

Align a container build with its Kustomize manifest

Align the Dockerfile at /environment/Dockerfile with the files referenced by the kustomization in /app (kustomization.yaml), notably config.yaml.

  • 311-line reference solution
  • Verified in a fresh sandbox

Databases

Stand up InfluxDB 2 against a compliance policy

Set up InfluxDB 2 using the configuration at /etc/influxdb/init.json.

  • 329-line reference solution
  • Verified in a fresh sandbox

Version control

Version a dataset with DVC and reconcile its history

Set up a DVC-tracked dataset under /app/dataset/ with two items: dataset/foo and dataset/test/0.

  • 313-line reference solution
  • Verified in a fresh sandbox

Data & scientific computing

Install an HPC scheduler's Python bindings, verified

Install the Flux Python package from the tarball at /app/flux-python-0.48.0rc6.tar.gz after verifying its integrity against the provided checksum file.

  • 171-line reference solution
  • Verified in a fresh sandbox

System administration

Rebuild a GRUB boot sequence from a broken system

Assemble a three-line GRUB boot sequence (linux, initrd, boot) in /app/result.txt for the system rooted at /mnt/target.

  • 214-line reference solution
  • Verified in a fresh sandbox

Debugging & performance

Diagnose and fix a Python dependency bug

In /app, diagnose and fix the environs URL parsing bug.

  • 202-line reference solution
  • Verified in a fresh sandbox

Security analytics

Count attack telemetry with the EQL shell

Working in /app, count each event_type in the dataset '/app/normalized-atomic-red-team.json.gz' using the EQL shell and save results to counts.json.

  • 242-line reference solution
  • Verified in a fresh sandbox

Breadth across real engineering work

Share of terminal tasks by area, from system administration and databases to security and scientific computing.

  • Software engineering19%
  • Scripting & automation17%
  • System administration14%
  • Environment setup11%
  • Version control8%
  • Containers & cloud7%
  • Security6%
  • Databases4%
  • Data & scientific computing3%
  • Files & storage3%
  • Networking2%
  • Other7%

Why verified synthesis

Scale without losing the proof.

Hand-written environments are slow to produce. Prompted generation is fast, but a task can look perfect and be silently broken. We generate aggressively and then try to break every task, so only environments that survive the gates ship.

Verified by execution

A task is accepted only if its reference solution passes in a fresh sandbox, a do-nothing agent fails, and every check is stated in the instruction.

Difficulty you can dial

Tasks grow harder round by round. Every task carries a measured tier, so you train on a curriculum instead of a pile.

Fresh every month

New, never-published environments each cycle, synthesized for you and audited against public benchmark test sets before delivery.

One consistent harness

Passing runs from many harnesses are rewritten into the single harness your model trains and runs under, so it learns the skill, not the scaffolding.

Capabilities

Six task families. One standard of proof.

Each family pairs what the model sees with what it must produce and how we check it. Pick one, or combine them into a single training mix.

Terminal & coding agents

Multi-hour workflows where the agent explores a real environment, writes and runs programs, and checks its own work. Terminal tasks are coding tasks.

Model sees
A live workspace, an instruction, and the tools a practitioner would reach for.
Model produces
Working code and a verified end state after hundreds of commands: rebuilt programs, migrated services, repaired builds.
We verify
Artifacts and system state, with checks that reject placeholders, hard-coded outputs, and skipped steps.
agent@sandbox:/app
$ jq -c .axes_to_vary observations/manifest.json
["inlet_size","outlet_size","bend_angle","rib_count"]
$ ls /tmp/obs/public_1/native
metrics.json  product.stl
$ jq -c .parameters /tmp/obs/public_1/scenario.json
{"bend_angle":238,"inlet_size":11.15,"rib_count":5}
$ python3 workspace/reconstruct.py < hidden.json
verifier ▸ hidden configurations vs. real OpenSCAD
✓ geometry matches on every hidden scenario

How it works

Generate aggressively. Then try to break every task.

Generation is the cheap part. The value is in the gates: a task reaches you only after it survives all of them.

  1. 1

    Seed

    Licensed or authored seed tasks and real software workflows. Never benchmark test sets.

  2. 2

    Extend

    Rewrite operators add tools, services, data, and failure modes. Solution, environment, verifier, and instruction change together.

  3. 3

    Verify

    Fresh-sandbox build. The reference must pass, a do-nothing agent must fail, and every verifier check must be stated in the instruction.

  4. 4

    Calibrate

    Measure pass rates with reference solvers or your checkpoint, and assign every task a difficulty tier.

  5. 5

    Deliver

    Environments, verified trajectories, and a calibration and integrity report for every batch.

  • Reference solution passes
  • No-op agent fails
  • Replay shortcuts fail
  • Instruction ↔ verifier consistency
  • Leakage & overlap audit
  • Sampled human review

Trajectory scaling

More passing trajectories. One consistent harness.

Hard tasks get solved in different ways under different harnesses, and raw trajectories carry each harness's scaffolding with them. We pool successes across harnesses, then rewrite each one into the single harness your model trains and runs under, so it learns the skill, not the scaffolding.

  1. Discover. Several harnesses attack the same tasks and solve ones no single harness can.
  2. Plan. Each success becomes a runbook: end state, milestones, checks, and recovery steps, never the finished answer.
  3. Screen. A critic rejects runbooks that leak the verifier or the solution and asks for revisions.
  4. Re-execute. The task is solved again in a fresh sandbox under one consistent harness. Only passing, leak-free runs are kept.
Discover Rewrite Train General loop Stateful workflow Self-reflect & retry Runbookplanner Criticleak screen Re-executefresh sandbox revise verified runs, one harness
+34%
more tasks solved by pooling harnesses than by the best single harness
5.5×
verified training trajectories per successful source run
57% → 74%
Terminal-Bench 2 pass@3 after training on rewritten trajectories
+20.8 pts
over fine-tuning directly on the raw successes (Terminal-Bench 2)

Calibrated difficulty

A curriculum, not a pile.

Each synthesis round makes tasks measurably harder: longer solutions, more tools, more checks. In our runs, a fixed strong solver went from solving 90% of first-round tasks to 2.5% by round fifteen, with no ceiling in sight.

We measure your checkpoint on every tier and ship the band where it succeeds sometimes, but not yet reliably. As your model improves, the band moves up with it.

WeakerReference solverStronger

Learning zone (10–50% success) Your model (modeled) Reference solver Measured reference results (pass@4) Illustrative: the curve interpolates measured reference-solver results. Moving the slider shows how the recommended band shifts.

Watch one task grow

Same goal. More to get right at every stage.

Successive stages from one task family: build a Git commit using only low-level plumbing commands. Each stage keeps the goal and adds requirements, and the reference solution and verifier grow with it. New requirements are highlighted.

Reference solution –
Instruction –
Reference solution length by stage

Illustrative family: stages are related tasks from successive synthesis rounds, shown in order of depth.

Deliverables

Everything you need to train on it tomorrow.

Delivered in a reproducible, container-first format that drops into SFT and RL pipelines, with adapters for your stack on request.

RL environments

Self-contained Docker environments with instruction, executable verifier, reference solution, and metadata, in Harbor task format.

Verified trajectories

Agent runs kept only when they pass the hidden checks, rewritten into one consistent harness, and exported as chat-format JSONL for fine-tuning.

Calibration report

Tier and measured pass rate for every task, against reference solvers or your own checkpoint.

Integrity report & data card

Overlap audit against public benchmark test sets, shortcut controls, and sampled human review, plus the domains, tools, licenses, and known limitations of every batch.

task/
├── instruction.md      # what the agent sees
├── task.toml           # tier, tags, timeouts
├── environment/
│   └── Dockerfile      # fresh, pinned sandbox
├── solution/
│   └── solve.sh        # reference solution
└── tests/
    └── test_state.py   # executable verifier
trajectories.jsonl      # verified agent runs
report.json             # calibration + audit

Alignment

Built for the skills frontier evaluations measure.

We train skills, not test items. No benchmark test data is used, copied, or derived, and every batch ships with an overlap audit.

Terminal & long-horizon agents

  • Terminal-Bench 2
  • Terminal-Bench 4
  • Terminal-Bench Hard
  • Long-Horizon Terminal-Bench

Multimodal SWE & web

  • SWE-bench Multimodal
  • Vision2Web
  • Interaction2Code
  • Design2Code

CAD & visual-to-code

  • BenchCAD
  • Chart2Code
  • Image2Struct
  • MMSVGBench

Research

Proven in our research.

Each method is measured across multiple seeds and against matched-budget baselines.

Scientific software

Software-in-the-loop reconstruction

Existing scientific software becomes both the reference solution and the source of hidden test cases, so verification scales with the software rather than with authoring effort.

+5.6 pts
Terminal-Bench 2 after fine-tuning (47.9% → 53.6%, three seeds)
#1 of 5
training corpora at a matched token budget, on all four evaluations
98%
verifier agreement with blinded human review

Task difficulty

Recursive verified synthesis

Accepted tasks seed the next round. Each round extends the solution, verifier, and instruction together and re-validates the whole task in a fresh sandbox, with no human authoring in the loop.

+10 pts
best fine-tuning gain across Terminal-Bench 2, Terminal-Bench Hard, and LHTB
+41%
relative gain on Terminal-Bench Hard with agentic PPO
36×
drop in fixed-solver success from round 1 to 15 (90% → 2.5%), no ceiling observed

Trajectory scaling

Recursive self-rewrite

Successes discovered under diverse harnesses are rewritten by the same base model into one general harness, with leakage screening and fresh-sandbox re-execution.

74.2%
Terminal-Bench 2 pass@3, up from 57.0%
6×
Terminal-Bench 4 pass@3, from 1.5% to 9.1%
+20.8 pts
over direct fine-tuning on the same successes

Working together

Start small. Scale every month.

01

Pilot

A calibrated sample batch in your format, measured on your checkpoint, for the task families that matter most to you.

02

Monthly refresh

Fresh, never-published environments every month at the tiers you choose. Difficulty moves up as your model improves.

03

Custom & exclusive

New domains, software stacks, and modalities built for your team alone, on your timeline.

Tell us what your next checkpoint needs to learn.

We'll come back with a pilot plan: task families, difficulty tiers, delivery format, and a calibration run on your model.