A visual walk through how an evolving state machine becomes a time-indexed computation graph, why reverse-mode differentiation can traverse both depth and time, and where exact credit assignment collides with vanishing gradients, memory cost, and online alternatives.
Slide the horizon to stretch the same recurrent cell across time. The hidden state passes forward; the loss can attach at every frame or only the end.
Mock hidden activations for a recurrent state vector. Warmer cells indicate stronger carry-over influence from prior time steps.
Forward arrows move state. Backward arrows aggregate sensitivity from future losses into earlier states.
Every reused parameter copy contributes a local gradient term. BPTT sums them after the reverse sweep.
Pick a time slice to see its local loss contribution, state-to-state Jacobian, and cumulative co-state signal coming from the future.
As repeated state transitions multiply, gradient mass can decay or spike. This mock matrix shows how sensitivity survives along time offsets.
Read the same BPTT computation as reverse-mode AD: local output gradients plus future-state recursion.
Full exact gradients need stored states. Checkpointing trades recomputation for less memory.
Toggle the metric. The shapes are illustrative: exact reverse sweeps stay accurate but become memory-heavy; online methods flip the cost profile.
Long horizons make early events hard to credit. LSTM and gated variants were built to flatten this decay curve.
Modern methods bend Werbos’s exact recipe toward streaming, synthetic, hybrid, or checkpointed variants.
Smooth trajectories fit BPTT cleanly. Switching, saturation, and mode jumps carve out rough regions.
Reverse accumulation predates deep learning. BPTT sits where automatic differentiation, backpropagation, recurrent learning, and control meet.
Hover nodes to see which neighboring ideas BPTT inherits, competes with, or motivates.
Core lineage and follow-on work mentioned in the episode. Links prefer the source list supplied for the episode.