1 00:00:01,000 --> 00:00:35,950 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is mHC: Manifold-Constrained Hyper-Connections, by Zhenda Xie and 19 co-authors at DeepSeek-AI, arXiv 2512.24880, version two, dated January 5th, 2026. One number in it is stuck in my head: 3000. That's how far the signal gain reportedly climbed in a 27-billion-parameter model when someone got creative with the residual connection. 2 00:00:35,950 --> 00:01:07,325 [Dr. Ada Shannon] That number is the stakes. The residual connection is the one piece of the transformer nobody has touched in a decade. This paper says you can improve it, but the improvement blows up unless you constrain it. DeepSeek builds on ByteDance's Hyper-Connections, from Zhu and colleagues in 2024. The test models are DeepSeek-V3-style mixture-of-experts at 3B, 9B and 27B. The constraint is a 1967 matrix algorithm. Making it affordable took new GPU kernels and pipeline changes, for a reported 6.7% extra training time. 3 00:01:07,325 --> 00:01:25,450 [Hal Turing] A decade unchanged? Attention has been rewritten a dozen times. That sounds like the field got lazy about the skip path. Remind me what the residual connection even is, and why it earned this immunity. I'd rather hear it from you than pretend I remember the equation. 4 00:01:25,450 --> 00:02:07,700 [Dr. Ada Shannon] I'd push back on lazy, Hal. It's a load-bearing wall. The layer computes x plus F of x. Unroll that over layers and the deep input is the shallow input plus a sum of layer outputs. The shallow signal arrives unmodified, and gradients get an additive path straight back. Kaiming He and colleagues at Microsoft Research argued in 'Identity Mappings in Deep Residual Networks,' 2016, that anything other than identity on that path degrades very deep training. So the field improved the insides, like multi-head latent attention and mixture-of-experts. The paper calls that micro-design. Macro-design, how blocks connect, was tried by DenseNet from Huang at Cornell in 2017 and FractalNet from Larsson at the University of Chicago in 2016. Neither displaced the plain skip. 5 00:02:07,700 --> 00:02:24,175 [Hal Turing] But if the plain skip beat every challenger, that reads like inertia with a good story attached. Losing fights doesn't prove the ring was fair. Still, say the bar is real. What does Hyper-Connections do, and why widen the stream instead of just the layer? 6 00:02:24,175 --> 00:03:07,100 [Dr. Ada Shannon] Inertia is not testing and losing, Hal. Those designs paid in memory or parameters. The real bar was no extra FLOPs and stability at depth, and that's what HC claims to clear. Instead of one C-wide stream, it keeps n of them, with n equal to 4 here. Widening the layer costs FLOPs, and widening the stream doesn't. Three learnable maps: H-pre reads the streams into the layer, H-post writes the output back, and H-res, an n-by-n matrix, mixes the streams together. Each is a static bias plus an input-dependent part gated by an alpha initialized at 0.01. The paper's Table 1 ablation: H-res alone gives a 0.022 absolute loss improvement, adding pre gives 0.025, and adding post gives 0.027. 7 00:03:07,100 --> 00:03:26,725 [Hal Turing] Hold on, that's actually— so 0.022 of the 0.027 comes from H-res alone? That's a clean ablation, three rows and the story is unambiguous. But H-res is a matrix applied at every single layer. Multiply enough of those together and something has to give. 8 00:03:26,725 --> 00:04:04,825 [Dr. Ada Shannon] That is the paper's argument. Across layers, the skip path becomes a product of H-res matrices, the composite mapping. Nothing bounds it, so signal can grow or shrink with depth, and the feature mean isn't conserved. Their diagnostic is Amax Gain Magnitude: the largest absolute row sum for the forward signal, the largest column sum for backward gradients, where 1 is ideal. At 27B they report a loss surge near step 12,000 and the gain peaking near 3000. The second problem is the memory wall. At n equal 4, HC reads about 21C values per token versus 2C for a plain residual. It also sends n times more activations between pipeline stages. 9 00:04:04,825 --> 00:04:17,775 [Hal Turing] So an exploding product plus a bandwidth bill. I'm parking my skepticism until later. What's the one-line fix, and where does a doubly stochastic matrix come in? I need that term unpacked. 10 00:04:17,775 --> 00:04:53,200 [Dr. Ada Shannon] Force H-res to be doubly stochastic: non-negative entries, every row and column summing to 1. That set is the Birkhoff polytope, the convex hull of permutation matrices, so H-res applied to the streams is a convex combination. The spectral norm is at most 1, and products of such matrices stay in the set, so the composite stays bounded. The projection is Sinkhorn-Knopp, from Richard Sinkhorn and Paul Knopp at the University of Houston in 1967. Exponentiate the entries, then alternately normalize rows and columns, 20 iterations here. The bill for all this is fused kernels, recomputation and DualPipe changes. 11 00:04:53,200 --> 00:05:18,950 [Hal Turing] A 1967 algorithm holding up a 27-billion-parameter run. Sinkhorn and Knopp never had a GPU budget, and they're still getting overtime. Here's what nags me, though. He et al. needed the skip path to be exactly identity. A doubly stochastic matrix conserves the mean and bounds the norm, but it isn't identity unless it's a permutation. Does the clean gradient-flow argument survive that swap? 12 00:05:18,950 --> 00:05:50,000 [Dr. Ada Shannon] Hold that question, Hal, because the answer depends on the machinery. mHC changes how every coefficient is computed, not just the constraint. Flatten the n-by-C stream into one vector of length nC, RMSNorm it, and multiply by linear maps phi. Scale by alpha and add a static bias. H-pre becomes a sigmoid, H-post becomes two times a sigmoid, and H-res goes through Sinkhorn-Knopp. Original HC used tanh gating on each stream separately. Here every coefficient sees the full context, and pre and post are forced non-negative so their terms can't cancel. 13 00:05:50,000 --> 00:05:59,600 [Hal Turing] So the constraint arrives with a few extra changes attached. Noted. Which guarantees actually come out of the projection, and how exact are they? 14 00:05:59,600 --> 00:06:27,700 [Dr. Ada Shannon] Twenty rounds is a finite approximation, so H-res is only approximately doubly stochastic, and the paper admits it: in Fig. 7a the backward gain drifts slightly off 1. Three guarantees are claimed. The spectral norm is at most 1, so the map never expands. The set is closed under multiplication, so the sixty-layer product stays inside it. And it's a convex combination of permutations, so repeated application only mixes streams more. At n equals 1 it collapses to the scalar 1, the ordinary residual. 15 00:06:27,700 --> 00:06:40,875 [Hal Turing] That's the math. Now the bandwidth bill from the memory wall. Four streams means roughly four times the traffic, and I'm guessing nobody wrote this in plain PyTorch and hoped. What happened to the kernels? 16 00:06:40,875 --> 00:06:57,800 [Dr. Ada Shannon] Kernel fusion means merging operations that touch the same memory into one GPU kernel, so you read the data once instead of five times. RMSNorm on a length-nC vector is slow, so first they move the divide-by-norm to after the matmul— 17 00:06:57,800 --> 00:07:03,600 [Hal Turing] Oh wait wait wait— that's legal? Dividing after the matmul gives the same answer? 18 00:07:03,600 --> 00:07:36,900 [Dr. Ada Shannon] Identical. The norm is one scalar per token, so it factors straight out of the product. Then mixed precision: tf32 for phi, bf16 for x, fp32 for the coefficients. The two scans over x fuse into one kernel with a single backward kernel. Sinkhorn runs in one kernel, with a custom backward that recomputes the iterations on-chip. The apply kernel merges the residual add, cutting reads from (3n+1)C to (n+1)C and writes from 3nC to nC. Most of it is written in TileLang. 19 00:07:36,900 --> 00:07:44,675 [Hal Turing] Fine, but four streams also means four times the activations kept for backprop. Where does that memory go? 20 00:07:44,675 --> 00:08:26,150 [Dr. Ada Shannon] They throw it away. Recomputation: discard the mHC kernels' intermediates after the forward pass and re-run those cheap kernels in backward, never the layer function F. Only the block input is stored, once per L-r layers, plus a transient (n+2)C times L-r. Minimize the total and L-r star is about the square root of nL over n+2. For n=4 and 61 layers, that's about 6. The pipeline traffic from earlier is handled by extending DualPipe: the F-post,res kernels of MLP layers run on a high-priority compute stream so they never block communication, attention avoids persistent kernels so it can be preempted, and recomputation needs no communication because the block input is cached locally. 21 00:08:26,150 --> 00:08:30,376 [Hal Turing] Systems tour done. What did they train, and what came out? 22 00:08:30,376 --> 00:09:12,726 [Dr. Ada Shannon] DeepSeek-V3-style MoE at 3B, 9B and 27B, n=4, AdamW with betas 0.9 and 0.95, weight decay 0.1, 2000 warmup steps, step decay. Epsilon is 1e-20 for both layer norm and AdamW. The main run is 27B for 50k steps on 262 billion tokens, plus a separate 3B on a trillion. Baseline, HC and mHC share settings. Final loss lands 0.021 below baseline. On BBH, mHC hits 51.0 against 48.9 for HC and 43.8 for baseline. On MMLU it's 63.4, 63.0, 59.0. HC edges it on MATH, 26.4 to 26.0. 23 00:09:12,726 --> 00:09:22,101 [Hal Turing] I'll say it plainly: gradient norm tracking the baseline at 27B is a clean result. Does the gain diagnostic back it up? 24 00:09:22,101 --> 00:09:52,026 [Dr. Ada Shannon] The mHC composite gain peaks near 1.6, versus about 3000 for HC. In Fig. 8, HC's single-layer entries include -6.81 and 18.73, and its thirty-layer products run into the hundreds. mHC's deep composites look nearly uniform. In HC, when one gain is large the others tend to be too. Scaling across 3B, 9B and 27B keeps the advantage with only marginal attenuation. The overhead claim is 6.7% at n=4 in their in-house training. 25 00:09:52,026 --> 00:09:59,801 [Hal Turing] Marginal? 6.7% on a run this size is real GPU-days. I don't buy calling that cheap. 26 00:09:59,801 --> 00:10:11,501 [Dr. Ada Shannon] I disagree, Hal. Table 2 puts naive HC at roughly n times the memory traffic. Getting from that down to 6.7% is the point of the fusion work. 27 00:10:11,501 --> 00:10:21,051 [Hal Turing] Against a naive implementation nobody would ship. But you're right that it's the paper's own yardstick. What I can't find is a breakdown. 28 00:10:21,051 --> 00:10:28,501 [Dr. Ada Shannon] There isn't one. No per-component split, no throughput table. One in-house number, reported, not decomposed. 29 00:10:28,501 --> 00:10:53,201 [Hal Turing] Ada, we've heard the cure, so let me audit the disease. The instability case is one 27B run: a loss surge near step 12k, and gain peaking near 3000. No seed count, no repeated HC runs. And in Figure 2 the HC-versus-mHC loss gap tops out around 0.01. Is that a reproducible failure mode, or one unlucky trajectory? 30 00:10:53,201 --> 00:11:15,826 [Dr. Ada Shannon] The paper can't tell you. The mechanism is sound, since an unbounded matrix product can blow up. But it shows the gain explosion and the loss spike are correlated, not that one causes the other. The 3B and 9B runs only appear in scaling, never in stability. And Amax Gain is averaged over one selected sequence, while H-res depends on the input. So 3000 could be typical or a worst case. Sound mechanism, thin demonstration. 31 00:11:15,826 --> 00:11:36,576 [Hal Turing] Here's something I genuinely don't know. Table 5 lists the AdamW epsilon at 1e-20, where most people run 1e-8. What does that do to a small mixing parameter with a tiny gradient? Could it be manufacturing HC's instability? HC only ever fought at the baseline's optimizer settings. 32 00:11:36,576 --> 00:12:12,301 [Dr. Ada Shannon] Nobody can say, because it's never swept. Epsilon that small lets Adam take near-full steps on parameters whose gradients are mostly noise. Also untested: a lower learning rate or longer warmup for the HC parameters, tighter clipping, or bounding H-res with a row-softmax or fixed-scale tanh. DeepNet, from Hongyu Wang at Microsoft Research in 2022, stabilized thousand-layer stacks with scaling and initialization at near-zero cost. The HC authors at ByteDance have their own hyperparameter guidance, and Frac-Connections attacks the memory cost by fractionating the stream. Was any of it tried? The paper doesn't say. 33 00:12:12,301 --> 00:12:47,501 [Hal Turing] Then Table 4. mHC over HC is GSM8K 53.8 to 53.2, MMLU 63.4 to 63.0, HellaSwag 74.7 to 74.3, and HC wins MATH. Single run, no intervals. I say everything between mHC and HC is noise. Even baseline to HC, BBH 43.8 to 48.9 for a 0.02 loss gap, smells like few-shot variance on 4.14 billion active parameters. 34 00:12:47,501 --> 00:13:05,201 [Dr. Ada Shannon] I actually disagree, Hal. Noise doesn't usually line up. mHC beats the baseline on all eight, and BBH and DROP lead HC by over two points. Big benchmark jumps on small loss gaps are normal for reasoning tasks with thresholds. Eight for eight is not a coin flip. 35 00:13:05,201 --> 00:13:18,951 [Hal Turing] Sorry to cut you off but— eight benchmarks from one checkpoint aren't eight trials. Same weights, correlated errors. That's one coin, flipped once, reported eight ways. Give me three seeds and I'll believe the eight. 36 00:13:18,951 --> 00:13:40,026 [Dr. Ada Shannon] That lands. Correlated tasks, one run, so I'll say probably real against the baseline, unproven against HC. There's also a cost they skip. Table 1 has full HC at minus 0.027 loss, while mHC beats the baseline by 0.021 at 27B. Table 1's scale isn't stated, so we can't tell whether stability cost some of HC's raw gain. 37 00:13:40,026 --> 00:14:01,501 [Hal Turing] Now the part I can't shake. In Figure 8 the sixty-layer mHC composite is nearly 0.25 everywhere. That's four streams averaged into one. Does the widened residual collapse toward a single stream? The layer-30 matrices look near-diagonal, so where does the identity behavior actually live, and did anyone measure it? 38 00:14:01,501 --> 00:14:34,501 [Dr. Ada Shannon] Nobody measured stream rank or cosine similarity by depth. Bounded is not unimpeded: a doubly stochastic product drifts toward the uniform matrix, so the mean survives and everything else contracts. Dong at Google showed in 2021, in Attention is Not All You Need, that row-stochastic products lose rank doubly exponentially. Oono and Suzuki at Tokyo showed graph networks oversmooth the same way in 2020. And Sinkhorn convergence is only linear. Figure 7's backward gain drifts to about 1.6 over 61 layers. No iteration ablation, nothing at 200 layers. 39 00:14:34,501 --> 00:14:49,401 [Hal Turing] Systems side. Setting the 6.7% aside, the hardware and parallelism setup are unspecified too. Is this effort better spent on a bigger model? DenseFormer and MUDDFormer are cited but never run. 40 00:14:49,401 --> 00:15:11,701 [Dr. Ada Shannon] That's the missing comparison. DenseFormer, from Pagliardini at EPFL in 2024, and MUDDFormer, from Xiao and colleagues in 2025, mix across layers without four-fold stream traffic. Figure 6a compares at equal FLOPs, which ignores wall-clock and memory. But the TileLang kernels, stage-aligned recomputation and DualPipe changes are what make this usable, arguably as much a contribution as the math. 41 00:15:11,701 --> 00:15:21,676 [Hal Turing] The I/O accounting in Table 2 is clean, and I'll say that plainly. But it covers training only. Who should adopt this, and what about serving? 42 00:15:21,676 --> 00:15:51,901 [Dr. Ada Shannon] Large MoE teams with pipeline stacks and kernel engineers should watch it. Everyone else can run small HC versus mHC with several seeds, an epsilon sweep, and softmax-bounded H-res as baselines. Log composite Amax Gain as an early warning either way. Serving is untouched: four streams add traffic on memory-bound decode layers, and logit lens, steering, probes and layer skipping all assume a single residual stream. Open questions: other manifolds, larger n, dense models, and the plasticity cost of convex mixing. 43 00:15:51,901 --> 00:16:14,951 [Hal Turing] So mHC trains stably at 27B, keeps composite gain bounded, beats the baseline on eight benchmarks, and reports 6.7% overhead. What it asserts without testing is that HC's instability is general, that gain explosion causes the loss spike, and that HC got a fair fight. Thanks for listening, everyone. Goodbye.