AI Post Transformers • Visual Companion

Backpropagation Through Time Explained

A visual walk through how an evolving state machine becomes a time-indexed computation graph, why reverse-mode differentiation can traverse both depth and time, and where exact credit assignment collides with vanishing gradients, memory cost, and online alternatives.

Time DepthBackprop over layers and recurrent steps
Shared WeightsOne parameter set reused across the entire unrolled chain
Exact GradientsReverse-mode AD on a trajectory, not finite differences
Main TensionClean math vs memory, latency, and long-horizon instability

From Recurrence to a Movie Reel

Slide the horizon to stretch the same recurrent cell across time. The hidden state passes forward; the loss can attach at every frame or only the end.

Interactive horizon + hover state traces

State Memory Heatmap

Mock hidden activations for a recurrent state vector. Warmer cells indicate stronger carry-over influence from prior time steps.

hover cells for value + timestep
low mid high

Credit Flow Summary

Forward arrows move state. Backward arrows aggregate sensitivity from future losses into earlier states.

Where the Gradient Lives

Every reused parameter copy contributes a local gradient term. BPTT sums them after the reverse sweep.

Backward Through Time

Pick a time slice to see its local loss contribution, state-to-state Jacobian, and cumulative co-state signal coming from the future.

step explorer

Jacobian Product Attenuation

As repeated state transitions multiply, gradient mass can decay or spike. This mock matrix shows how sensitivity survives along time offsets.

Gradient Formula Map

Read the same BPTT computation as reverse-mode AD: local output gradients plus future-state recursion.

Activation Checkpoint Budget

Full exact gradients need stored states. Checkpointing trades recomputation for less memory.

BPTT vs Truncated BPTT vs RTRL

Toggle the metric. The shapes are illustrative: exact reverse sweeps stay accurate but become memory-heavy; online methods flip the cost profile.

mode switcher

Vanishing Gradient Profile

Long horizons make early events hard to credit. LSTM and gated variants were built to flatten this decay curve.

Approximation Design Space

Modern methods bend Werbos’s exact recipe toward streaming, synthetic, hybrid, or checkpointed variants.

Control and Differentiability Boundary

Smooth trajectories fit BPTT cleanly. Switching, saturation, and mode jumps carve out rough regions.

Method Lineage

Reverse accumulation predates deep learning. BPTT sits where automatic differentiation, backpropagation, recurrent learning, and control meet.

Concept Constellation

Hover nodes to see which neighboring ideas BPTT inherits, competes with, or motivates.

hover nodes

References

Core lineage and follow-on work mentioned in the episode. Links prefer the source list supplied for the episode.

Rumelhart, Hinton, Williams, 1986 — Backpropagating Errors
Williams, Zipser, 1989 — RTRL
Pineda, 1987 — Generalization to RNNs
Linnainmaa, 1976 — Reverse Accumulation Roots
Bengio, Simard, Frasconi, 1994 — Long-Term Dependencies
Hochreiter, Schmidhuber, 1997 — LSTM
Baydin et al., 2018 — AD Survey
Margossian, 2019 — Efficient AD Review
Pearlmutter, 1994 — Hessian-Vector Products
AI Post Transformers, 2026 — LSTM and Vanishing Gradients