A visual map of the paper’s claim: small open-weight agent models can get materially better when two data flywheels keep turning at once, one for verifiable reasoning failures and one for branching tool-use workflows with recovery.
The paper’s center of gravity is not a single dataset. It is a loop where training failures are recycled into harder reasoning tasks, while simple workflows are expanded into branchy, executable agent trajectories.
This mock matrix visualizes the paper’s logic: as rounds progress, easy linear failures should cool off, while rare recovery branches become the new frontier the curriculum keeps generating.
The sales pitch is operational: smaller agentic models should close enough of the performance gap on search and data-analysis workloads that massive frontier models stop being the default for every repetitive workflow.
The transcript frames behavior trees as the bridge from polite demos to real contingencies. A straight trace says “step 1, step 2, step 3.” A tree adds sold-out flights, contradictory constraints, verification, fallback tools, and the option to ask the user to clarify.
Compact links for the papers and episodes shaping this visual argument.