∇-Reasoner: Gradient Descent at Inference Time

nabla-Reasoner · Peihao Wang et al. · UT Austin, UC San Diego, TTIC, Georgia Tech · ICLR 2026
arXiv:2603.04948 First-Order Inference Test-Time Compute

Zeroth-Order vs First-Order Inference

The paradigm shift from sampling to optimization at decode time

Zeroth-Order Methods

Generate candidates, evaluate, pick the best. No directional signal.

Random probes — no direction

CoT · Self-Consistency · ToT · RAP · Best-of-N

First-Order: ∇-Reasoner

Compute gradient of reward w.r.t. logits — follow the slope toward better answers.

optimum ∇ reward — directed optimization

Gradient descent on token logits at each decode step

Core Insight: Logits as Optimization Variables

Base LLM (weights frozen) Token Logits ℝ^|V| continuous vector ← optimization target ∇ reward / ∂ logits Reward Model (PRM: per-step scores) Token sampled

Decoding-Time Optimization (DTO) Pipeline

Four stages repeated at every decoding position — click each step

1 Generate Rollout
→
2 Gradient Descent
→
3 Sample Token
→
4 Rejection Sampling

Step 1: Generate Full Rollout

The base LLM generates a complete candidate continuation from the current position, producing all per-token logit vectors along the way. These logits are retained for optimization rather than being discarded after sampling.

1. GENERATE ROLLOUT The answer is logit vectors retained 2. GRADIENT DESCENT ∇ on logits R(reward) + KL(regularize) STE ⚠ 3. SAMPLE TOKEN sample from updated π → "42" 4. REJECT / ACCEPT R(new) 0.83 ≥ R(old) 0.61 ✓ ACCEPT advance position repeat at next decoding position Straight-Through Estimator (STE) Forward: argmax(logits) → one-hot token Backward: pretend argmax = identity ⚠ Biased gradient — not the true derivative Bengio, Léonard & Courville (2013) KL-Regularized RL Duality Theorem: DTO at inference ≡ KL-regularized RL Same output distribution as PPO/RLHF alignment → Online alignment without weight updates * Equivalence holds for true gradients, not STE approximation

Reward Landscape: Probing vs Navigating

Toggle between zeroth-order sampling and first-order gradient optimization

Global Optimum Local Optimum Low Reward Low Reward ★ best sample 8 samples, no directional signal → best-of-N picks the lucky one

Process vs Outcome Reward Models

Outcome Reward (ORM) step 1 step 2 step 3 step 4 final ✓ ← single scalar Gradient decays severely before reaching early steps Process Reward (PRM) 0.72 0.85 0.61 0.91 0.94 ← per-step scores Dense gradient signal — DTO's hard dependency (Lightman et al., 2023)

Benchmark Results

MATH benchmark — AMC, AIME, and Olympiad-level competition problems

Accuracy on MATH Benchmark

∇-Reasoner achieves 20%+ improvement over baseline

0% 20% 40% 60% 80% 40% Greedy 53% Best-of-N 50% SC 58% RAP 72% GRPO* *trains weights 69% ∇-Reasoner

Model Calls Reduction

Fewer forward passes, but each includes a backward pass (2-3x cost)

Best-of-N 64 calls (forward only) RAP ~50 calls (forward only) ∇-Reasoner ~38 forward + ~38 backward (2-3x each) ⚠ "10-40% fewer model calls" ≠ fewer FLOPs — backward passes cost 2-3x a forward pass

Token Skipping Heatmap

DTO skips high-confidence positions — hover cells to inspect

Token Position → Confidence

Critical Analysis

Key concerns raised in the discussion

Issue Severity Matrix

Impact on Claims Tractability Low Impact + Tractable High Impact + Tractable Low Impact + Hard High Impact + Hard STE Bias: biased gradient estimator, no ablation vs Gumbel-Softmax or REINFORCE STE Bias PRM sees soft logit inputs it was never trained on — gradient direction may be unreliable Dist. Mismatch Backward passes require 2-4x peak memory — breaks production inference stacks Memory 2-4x Only validated on math benchmarks — code, QA, commonsense untested Narrow Eval Method ceiling is determined by PRM quality — PRMs are scarce outside math PRM Depend. "Fewer model calls" ≠ fewer FLOPs — compute comparison is incomplete FLOP Acctg Token skipping rate unreported — efficiency claim depends on this number Skip Rate

Intellectual Lineage

2013 STE Bengio et al. 2020 PPLM Dathathri et al. 2022 CoT / SC Wei / Wang 2023 ToT / PRM Yao / Lightman 2024 Scaling TTC Snell et al. 2025 GRPO Guo et al. 2026 ∇-Reasoner Wang et al. zeroth-order │ first-order →

Verdict

The paradigm shift from test-time search-as-sampling to search-as-optimization is genuinely new. The KL-RL duality provides theoretical grounding, and the benchmark results are compelling on math.

But the gap between these results and production deployment is real: backward passes conflict with memory-optimized inference stacks, the reward model dependency limits generalization, and the evaluation perimeter is narrow. The constraints are tractable, not fundamental — the direction is worth following seriously.

References