arXiv:2602.02477 Interactive SVG Companion Reasoning · RL · Test-Time Scaling

Training LLMs for Divide-and-Conquer Reasoning

This page turns the episode into diagrams: how linear chain-of-thought differs from learned decomposition, where current post-training biases the policy, and why DAC-style RL changes both accuracy and search behavior under larger inference budgets.

Paper: Xiao Liang et al. (2026) Models: Qwen2.5-7B, Qwen3-4B Benchmarks: AIME24, AIME25, Beyond-AIME, HMMT-25
Core claim
Train the split
Naive branching often fails because standard post-training favors one long reasoning lane.
Observed gain
+8.6 / +6.3
Headline lift reported for Pass@1 and Pass@32 in the Qwen3-4B setup.
Search shape
Tree > trace
Useful extra compute comes from split, solve, and recombine, not only longer diaries.
Theme
Policy mismatch
Prompting for decomposition under a linear-trained policy can underperform even before cost dominates.

Linear Reasoning vs Structured Search

One panel shows the policy difference. The other shows why naive DAC prompting can be worse before the model is trained to propose useful subgoals.

Reasoning Pipeline

trace topology
productive split linear reasoning collapse / dead branch

Policy Bias Heatmap

rows = behaviors, cols = prompting regimes
Hover cells to see where the policy is concentrated. Blue means absent, orange means moderate, red means dominant.

Inference Budget Allocation

token budget share

How Divide-and-Conquer Actually Unfolds

Step through an illustrative math-style episode: decompose the task, solve subgoals, then merge evidence into a final answer. The right panel tracks which branches become reusable.

Progressive Task Graph

click stages

Subproblem Utility Matrix

subgoals × downstream value

Merge Quality

composition pressure
Stage 1 shows the model before useful branching exists.

What Extra Test-Time Compute Buys You

This view tracks two claims separately: absolute accuracy and slope under more inference budget. The plotted values combine transcript numbers with benchmark-shaped mock points to show the pattern.

Accuracy vs Inference Samples

k = 1, 2, 4, 8, 16, 32
CoT-style RL DAC-style RL direct DAC prompt on standard model

Benchmark Delta Bars

DAC-RL minus CoT-RL

Elasticity Gauge

gain from k=1 to k=32
The point is not just a higher starting point. DAC training also makes additional samples less repetitive and more structurally diverse.

Method Map: Prompting Tricks, Search Wrappers, and Trained Policies

The episode places this paper inside a wider reasoning stack. Methods are positioned by two axes: how much branching they use and whether structure is inferred at prompt time or learned into the policy.

Reasoning Strategy Map

x = linear → branching, y = prompt-only → trained

Method Family Share

episode references by emphasis

Tension Matrix

latency · controllability · exploration · data need
This is the design trade space implied by the conversation, not a leaderboard.

References

Liang et al., 2026 · main paper · arXiv:2602.02477
Yao et al., 2023 · branching search baseline
Zhou et al., 2022 · decomposition via prompting
Wang et al., 2022 · sample diversity over linear traces
Zelikman et al., 2023 · decomposition as helper-function structure
OpenAI, 2024 · reasoning-focused post-training context
Guo et al., 2025 · RL for reasoning