Volatility Optimization Is Actually Bayesian Inference

A visual companion to the episode on Kohei Honda's "Model Predictive Control via Probabilistic Inference: A Tutorial and Survey" — unifying path integral control, RL theory, and variational inference into one framework: PI-MPC.

arXiv:2511.08019 AI Post Transformers

Receding-Horizon Control Loop

The basic MPC cycle — expensive thinking at every tick

MPC predicts forward over a finite horizon, solves for the best action sequence, executes only the first action, throws the rest away, and re-solves from scratch next tick with fresh state — often dozens of times a second.

Why Random Shooting Collapses

Success rate vs. control dimensionality

Random shooting Distribution-based (PI-MPC)

Not a hard wall — a probability collapse. Fixed sample budget, exponentially worse odds of landing near the optimum as dimensions stack up.

Search vs. Inference

Same budget, different sampling strategy

Three Steps: Cost Function → Probability Distribution

Boltzmann posterior over control sequences

Temperature Controls Commitment

Toy cost J(u) = 0.6u² + sin(5πu) — drag λ

λ → 0.30
Cost J(u) Boltzmann posterior (renormalized per λ)

λ → 0: distribution sharpens hard around the global minimum, near-deterministic. λ → large: smooths back toward the prior, local minima gain mass.

Prior Mismatch Matters

Four Gaussian priors × resulting posterior quality

Best case: small variance, mean already near the mode. Worst case: mean far from the mode with small variance — nothing to rescue it. Wide variance on a mismatched mean partially recovers; wide variance on a good mean just dilutes the peak.

MPPI: The Closed-Form Special Case

Sample → rollout → score → softmax-weight → average, one pass

Why Fixing Covariance Matters

KL objective landscape over the mean parameter

GPU Parallelism Pays Off

Latency per control tick: CPU-serial vs. GPU-parallel rollouts

CPU serial GPU parallel

Cost is O(K·T) per tick; independent samples parallelize trivially once rollouts move onto GPU physics simulators.

The Hardware Ladder

Same algorithm, four very different substrates

Dimensionality Reduction

Raw joint-space vs. spline control points

Forward vs. Reverse KL

Mode-covering vs. mode-seeking under a bimodal cost

The Evidence Base Problem

What the illustrative figures actually run on

Every illustrative figure in the paper — temperature, prior mismatch, mode collapse — runs the identical 1-D toy cost. Section 3.4 then claims this scales to 30+-joint humanoids on GPU physics simulators. Nobody has shown the same knob interactions hold outside one dimension.

Self-Citation Density

Honda's own work across the four-way taxonomy

Honda-authored External method

Roughly a third of the "representative methods" column is one lab's output — a calibration note on reading the taxonomy as a neutral field map, not an accusation of rigging it.

What This Survey Actually Adds

Connective tissue, not new math

Honda's Eqs. 2–13 re-derive Williams' 2018 MPPI result; the VI framing makes Levine's RL-flavored inference tutorial legible to control-theory readers. The strongest evidence it matters: Pan et al.'s 2024 diffusion-trajectory paper landed at NeurIPS citing this bridge — outside robotics venues entirely.

References