nabla-Reasoner · Peihao Wang et al. · UT Austin, UC San Diego, TTIC, Georgia Tech · ICLR 2026 arXiv:2603.04948First-Order InferenceTest-Time Compute
Zeroth-Order vs First-Order Inference
The paradigm shift from sampling to optimization at decode time
Zeroth-Order Methods
Generate candidates, evaluate, pick the best. No directional signal.
CoT · Self-Consistency · ToT · RAP · Best-of-N
First-Order: ∇-Reasoner
Compute gradient of reward w.r.t. logits — follow the slope toward better answers.
Gradient descent on token logits at each decode step
Core Insight: Logits as Optimization Variables
Decoding-Time Optimization (DTO) Pipeline
Four stages repeated at every decoding position — click each step
1 Generate Rollout
→
2 Gradient Descent
→
3 Sample Token
→
4 Rejection Sampling
Step 1: Generate Full Rollout
The base LLM generates a complete candidate continuation from the current position, producing all per-token logit vectors along the way. These logits are retained for optimization rather than being discarded after sampling.
Reward Landscape: Probing vs Navigating
Toggle between zeroth-order sampling and first-order gradient optimization
Process vs Outcome Reward Models
Benchmark Results
MATH benchmark — AMC, AIME, and Olympiad-level competition problems
Accuracy on MATH Benchmark
∇-Reasoner achieves 20%+ improvement over baseline
Model Calls Reduction
Fewer forward passes, but each includes a backward pass (2-3x cost)
Token Skipping Heatmap
DTO skips high-confidence positions — hover cells to inspect
Critical Analysis
Key concerns raised in the discussion
Issue Severity Matrix
Intellectual Lineage
Verdict
The paradigm shift from test-time search-as-sampling to search-as-optimization is genuinely new. The KL-RL duality provides theoretical grounding, and the benchmark results are compelling on math.
But the gap between these results and production deployment is real: backward passes conflict with memory-optimized inference stacks, the reward model dependency limits generalization, and the evaluation perimeter is narrow. The constraints are tractable, not fundamental — the direction is worth following seriously.
References
[1]Wang et al. "∇-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space" ICLR 2026 — arXiv:2603.04948