Amazon AGI’s AutoGym generated 350 productivity tasks; 57% fell into its hard band for Claude Opus 4.6. After three repair passes, 280 remained, and 39% were hard. Defects had contributed to the initial failures. Repair also changed the task population, so this is not a fixed-sample difficulty comparison.
Generating an agent’s practice tasks creates another engineering problem: the tools must work, the requested outcome must be reachable, and the evaluator must recognize it. Otherwise, training can reward an incorrect result or penalize a valid one.
Three published research systems approach that problem differently: Snowflake and UNC’s Agent World Model, Amazon AGI’s AutoGym, and Alibaba’s Qwen team with Georgia Tech’s VHD-Play. Their papers describe environment-generation methods and experiments, rather than each company’s production training pipeline.
On Thursday, October 1, at 11 am Eastern, the free lightning lesson Can AI agents improve themselves? examines what changes persist between runs and how to evaluate them.
On Saturday, October 3, 10 am-3:30 pm Eastern, Engineering a Multi-Agent Forecasting System teaches agent engineering principles by replicating Bridgewater’s AIA Forecasting Agent.
What happens in one training episode
A task begins from a saved initial state: documents, database records, inventory, or other objects the agent can inspect and change. The model issues tool calls. Executable software returns observations and records their consequences. The agent’s sequence of actions and observations forms a trajectory; an evaluator scores the attempt using that record and the resulting state.
In reinforcement learning, the training procedure collects multiple attempts and uses their rewards to update model weights. The next batch runs with the updated model. Environment generation is the separate process that supplies the tasks, tools, starting states, and evaluators.
Recorded demonstrations can teach a successful tool-use sequence. An executable environment also lets the model try alternative actions and experience their consequences.
Environment generation produces agent-facing tasks and tools plus a separate evaluator. Training resets the task, collects tool interactions, scores the attempt, and updates model weights.
MLCommons’ September 24 announcement describes this integrated workload for MLPerf Training v6.1’s October submission round: inference, executable software-repair tasks and reinforcement-learning updates, measured by time to a quality target. Results from that round have not yet been reported.
Agent World Model builds database-backed software
Agent World Model, released in February, constructs scenarios and tasks, then SQLite databases, tool interfaces and verification components. Its 1,000 environments execute state changes in code. The released implementation supports resetting the database for another attempt.
One Spotify-like task asks the agent to add a track to a playlist. The verifier can inspect the relevant records before and after the attempt. A successful tool call that changes the wrong playlist leaves evidence in the database.
The reported training setup uses code-augmented model judgment: verification code extracts database evidence, and an LLM judge assesses completion using that evidence and the trajectory. Step-level checks also reward valid tool-call format. The repository offers a code-only verification option, but executable state transitions alone do not make the reported outcome judgment deterministic.
AutoGym derives answers from the constructed environment
AutoGym starts with a solution blueprint, materializes the environment, then executes the blueprint’s derivation against it to compute ground truth. Deterministic checks handle explicit constraints; open-ended outcomes also use process, outcome, and meta-judges. All retained productivity tasks received human validation.
Here, hard means a mean verifier score below 0.30 across eight attempts. In the authors’ seven-tool productivity setup, 217 of 300 AWM-generated tasks survived review; 4% were hard and 90% easy, versus 39% hard for AutoGym’s retained harder configuration. This compares generated task difficulty under one protocol, not overall system performance. Tables 1-2.
Reported generation costs are $2-8 per task, excluding human validation. Learning evidence remains limited to one 8B model trained for 500 GRPO steps. Appendix F and limitations.
VHD-Play separates solving from tool-mediated execution
VHD-Play derives dynamics and a scoring reference from a mathematical mechanism, then generates the scenario and tools. Its outcome reward normalizes achieved utility between a baseline and a solver-derived upper reference, clips it to 0-1, and does so without an LLM judge.
Its five-family diagnostic compares Qwen3.6-35B-A3B before and after 34 training steps:
Written problem, complete information: 0.962 to 0.992.
Stateful tools, parameters supplied: 0.231 to 0.875.
Stateful tools, parameters discovered through interaction: 0.204 to 0.815.
These are mean normalized outcome scores, not success percentages. The starting model struggled with execution even when the parameters were supplied. Table 2.
Of 3,300 admitted environments, 2,200 were assigned to training and 1,100 to evaluation. Marginal generation costs are reconstructed at roughly $0.01-0.03 per admitted environment using public model rates. The upper reference can be unattainable under partial observation, and replay checks vary by family. Sections 3-4 and appendices.
For research systems, the diagnostic suggests an evaluation distinction worth preserving. A model might calculate an optimal allocation from a complete specification yet fail to retrieve the relevant data, maintain state, and commit the right action through an interface. A test that supplies everything in one prompt misses those execution demands.
How construction affects the experiment
Each method places a different requirement on the generated implementation:
Database-backed generation: state changes and verifier evidence must refer to the same records and task.
Blueprint-first generation: the materialized documents and tools must support the solution derivation.
Mechanism-first generation: the tool interface must preserve the dynamics and information restrictions used to compute the scoring reference.
Generation cost also depends on the task contract. The two cost estimates above cover different task families and exclude different parts of the work. A cheap generated record can still require substantial validation before it belongs in an evaluation suite.
A point-in-time data task becomes an agent environment
Consider a proposed retrospective reconstruction task: produce a feature table containing the financial information publicly available on a specified decision date. The agent can inspect the full revision history, including later restatements, but must select only eligible versions. This differs from simulating a historical decision, where future documents should be hidden entirely.
The agent’s tools would let it discover filings, retrieve dated versions, resolve identifiers, inspect revision metadata, and write a feature-table artifact. Define availability in the fixture as the public SEC filing date; a fiscal period’s end and a vendor’s ingestion date are separate timestamps.
The Chapter 4 SEC XBRL fundamentals notebook already illustrates an as-of query using filing dates. Turning it into this agent test would require a versioned fixture, restricted tools, and an independently checked reference table.
The evaluator would inspect values, entity mapping, decision dates, and source-version lineage in the output artifact. Before testing an agent, challenge the evaluator with a later restatement substituted for an original value, a missing observation, and a correct value assigned to the wrong security. Equivalent row ordering should pass.
This checks information reconstruction. Whether those features predict returns requires a separate market-data experiment with held-out periods and implementation costs.
Start with a small collection of such executable tasks. Preserve their initial states and expected outcomes, test the checkers, and use the collection to compare frozen models or configurations. It can support a training decision later, after the evaluation itself is dependable.
Free lesson Thursday; hands-on workshop Saturday
On Thursday, October 1, at 11 am Eastern, the free Can AI agents improve themselves? lightning lesson examines what changes persist between runs and how to evaluate them.
On Saturday, October 3, 10 am-3:30 pm Eastern, Engineering a Multi-Agent Forecasting System applies the evaluation side to a working implementation. We’ll inspect traces, compare agent-count and aggregation choices in the ablation lab, and calculate Brier scores and calibration curves on resolved questions, keeping model weights fixed.



