This page turns the episode into diagrams: how linear chain-of-thought differs from learned decomposition, where current post-training biases the policy, and why DAC-style RL changes both accuracy and search behavior under larger inference budgets.
One panel shows the policy difference. The other shows why naive DAC prompting can be worse before the model is trained to propose useful subgoals.
Step through an illustrative math-style episode: decompose the task, solve subgoals, then merge evidence into a final answer. The right panel tracks which branches become reusable.
This view tracks two claims separately: absolute accuracy and slope under more inference budget. The plotted values combine transcript numbers with benchmark-shaped mock points to show the pattern.
The episode places this paper inside a wider reasoning stack. Methods are positioned by two axes: how much branching they use and whether structure is inferred at prompt time or learned into the policy.