Hal Turing and Dr. Ada Shannon open by situating the Dao-Gu paper — "Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality" (arXiv 2405.21060, ICML 2024) — within the bifurcated landscape of sequence modeling. For years, Transformer researchers and SSM researchers developed in parallel, unable to borrow optimizations across the divide. The episode traces the lineage from HiPPO and S4 through Mamba's selective state spaces, explaining why SSMs' linear-time theoretical advantage never translated to wall-clock wins: the GPU ecosystem was built for dense matrix multiply, and SSMs lacked the tooling that FlashAttention brought to attention. The hosts also credit the direct intellectual ancestor — Katharopoulos et al.'s 2020 "Transformers are RNNs" — which showed that softmax attention with a kernel approximation reduces to a linear recurrence, establishing the conceptual template Dao and Gu would later formalize. The core of the episode is a careful unpacking of structured semiseparable matrices, a class of objects from numerical linear algebra — Kalman filter theory, PDE solvers — entirely unknown to the ML community until Dao and Gu made the connection. Every entry of a causal SSM input-output matrix has the form C[i] times a chain of transition matrices A times B[j], which is precisely the generator representation of a rank-d semiseparable matrix. Shannon walks through the O(n) factored form — M[i,j] equals u[i] times v[j] below the diagonal — and explains how this structure encodes both the SSM recurrence and the masked attention computation as two views of the same algebraic object. The canonical reference is Vandebril, Van Barel, and Mastronardi's two-volume work from Johns Hopkins, 2008, a body of theory the ML community had never encountered. Once the connection is made, hardware-efficient algorithms from one domain port directly to the other. But the episode frames this mathematical achievement against a harder question raised by subsequent theoretical work: the L²M Condition of Chen et al. and the bipartite mutual information scaling law. The duality shows that SSM computation and attention computation are equivalent representations — but equivalence of computation does not imply equivalence of information retention. SSMs compress sequence history into a fixed-size state regardless of context length; the Transformer's KV-cache grows linearly, retaining more as context expands. The mutual information scaling law formalizes this gap: capturing the multi-token dependencies present in natural language requires a history state that grows with context length. The episode closes on what this implies for hybrid architectures — systems that combine SSM efficiency with selective attention — and whether the theoretical unification Dao and Gu achieved changes how practitioners should think about where compressed state fails.