1 00:00:01,000 --> 00:00:45,850 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're covering Hyper-Connections, by Defa Zhu and seven co-authors from the Seed-Foundation-Model Team at ByteDance, published at ICLR 2025. The headline claim is a big one. On OLMoE-1B-7B, trained on 500 billion tokens, they report 1.8 times faster convergence and about six extra points on ARC-Challenge. All of it comes from letting the network learn how strongly each layer connects to its own input. Ada, a number like 1.8x should make people put their coffee down. But it's also the kind of number we should check before we cheer. 2 00:00:45,850 --> 00:01:24,875 [Dr. Ada Shannon] Agreed, and we'll check it properly later. Not yet. The core question is this: can a network learn its own connection strengths, so the residual trade-off stops being a fixed design choice? Here's why that matters. Every modern transformer hard-wires how much a layer keeps of what it had versus how much it takes from what it just computed. Nobody trains that number. It's set by hand, and it's been the same since Kaiming He's Deep Residual Learning paper out of Microsoft Research in 2016. There, a layer computes x plus F of x. The output is added to an unchanged input, with the skip weight fixed at one. That's what let 100-plus-layer networks train at all, because gradients get a clean path straight down the stack. 3 00:01:24,875 --> 00:01:31,825 [Hal Turing] Okay, but we already have variants. Pre-Norm, Post-Norm. Aren't those the knobs people already turn? 4 00:01:31,825 --> 00:02:12,075 [Dr. Ada Shannon] Those are two hand-picked recipes, not knobs. Post-Norm normalizes after the residual sum, so the normalization sits on the skip path and gradients tend to fade with depth. Xiong and colleagues at Microsoft Research Asia showed in 2020 that it needs careful warmup. Pre-Norm normalizes the input to each block and leaves the residual stream alone. Training is stable, so nearly every modern LLM uses it. The catch is that deep-layer hidden states grow in norm and start looking like their neighbours. That's representation collapse, and it means extra layers barely add anything. The authors call the whole situation a seesaw. Pre-Norm buys healthy gradients and pays with collapse. Post-Norm does the reverse. 5 00:02:12,075 --> 00:02:28,625 [Hal Turing] So it's a playground seesaw where the only way to get one kid off the ground is to slam the other one into it. Here's my genuine question, though. If Pre-Norm is that stable, how do we know collapse is real and not a story? What does the paper actually show? 6 00:02:28,625 --> 00:03:10,925 [Dr. Ada Shannon] They show one picture, and I'll flag now that it's one picture. In Figure 3, they measure the cosine similarity between adjacent layers' inputs in OLMo-1B. Pre-Norm layers look very alike. The hyper-connection model's layers are noticeably less similar. That's suggestive, and Part 3 will ask what it does and doesn't prove. As for the idea itself, it isn't new in spirit. Highway Networks, by Srivastava and colleagues at IDSIA in 2015, put input-dependent gates on the carry-versus-transform choice. DenseNet, from Gao Huang at Cornell in 2017, connected each layer to all earlier layers. DenseFormer, by Pagliardini at EPFL in 2024, brought learned depth averaging to transformers. Keep those three names in your pocket, because they matter later. 7 00:03:10,925 --> 00:03:17,950 [Hal Turing] Got them. So what does hyper-connection actually do? Give me the intuition, and save the matrix for later. 8 00:03:17,950 --> 00:04:01,125 [Dr. Ada Shannon] Instead of one residual vector flowing through the network, you carry n copies, the expansion rate, typically two or four. Copies are replicated at the input and summed at the output. This is not a wider network inside the layers. The attention and FFN blocks are unchanged. What's learned is the wiring. Depth-connections weight a layer's output against a stream's own input, which generalizes the residual. Width-connections swap information between the n streams. In the dynamic version, those weights are predicted from the current token's hidden state, so the wiring changes per token. And the cost claim is tiny: roughly 0.03 percent more parameters and 0.2 percent more FLOPs. One setup fact: n has to be above one. The paper says at n equals one, the seesaw persists. I'll hold the numbers for later. 9 00:04:01,125 --> 00:04:06,150 [Hal Turing] And the headline model is a mixture-of-experts, right? Why test there? 10 00:04:06,150 --> 00:04:29,525 [Dr. Ada Shannon] OLMoE-1B-7B, from Muennighoff and colleagues at the Allen Institute for AI in 2024. A mixture of experts routes each token to a few feed-forward experts. This model holds seven billion parameters in total but activates only about 1.3 billion per token. It's a demanding test, because sparse models are already compute-efficient. The paper also runs dense OLMo-1B and OLMo-7B. So next, the mechanics and the results. 11 00:04:29,525 --> 00:04:38,625 [Hal Turing] Alright, matrix time. Ada, give me the block matrix in plain speech, and I'll try not to flinch at n-plus-one by n-plus-one. 12 00:04:38,625 --> 00:05:17,700 [Dr. Ada Shannon] One small matrix governs everything a layer reads and writes. A_m, a column, takes a weighted sum of the streams to form the layer's input. B, a row, decides how much of the layer's output lands in each stream. A_r, n-by-n, carries and mixes the streams themselves. Split it up and B plus the diagonal of A_r gives the depth-connections. A_m and A_r together give the width-connections. Each transformer block gets two of these modules, one around attention and one around the FFN. Set n to one and fill in ones, and Pre-Norm falls out. Post-Norm is the same shape with weights set by the input and output variances. Both are non-trainable special cases, so HC strictly generalizes them. 13 00:05:17,700 --> 00:05:25,800 [Hal Turing] Genuine question. Hand a network a fresh matrix per layer and it could start somewhere awful. What keeps day one sane? 14 00:05:25,800 --> 00:06:06,550 [Dr. Ada Shannon] It starts as exactly Pre-Norm. A_r is identity, B is ones, and A_m selects stream k mod n. Dynamic weights start at zero, and output projections are scaled by root n so the summed output keeps the baseline's standard deviation. Static parameters skip weight decay; dynamic ones get it. The dynamic version normalizes H, applies a linear map and a tanh, scales by a small learnable factor, and adds the static matrix. For n equals two, one matrix gives plain sequential layers, while an alternating odd-even pair gives parallel pairs, like the parallel transformer block in Ben Wang's 2021 Mesh-Transformer-JAX. Static HC locks in one arrangement after training, and dynamic— 15 00:06:06,550 --> 00:06:11,525 [Hal Turing] Oh wait wait wait— so every token could get its own layer ordering? 16 00:06:11,525 --> 00:06:34,325 [Dr. Ada Shannon] A soft mixture, yes. That's the claim. Now the bill. OLMo-1B with DHC times four adds 394 thousand parameters. But measured training memory rises 26.1 percent for that model, 28.3 percent for OLMo-7B, and 9.7 percent for OLMoE, because n streams of activations get stored. Their recomputation fix is proposed, not measured. The inference KV cache is untouched. 17 00:06:34,325 --> 00:06:37,275 [Hal Turing] Okay, so what was the testbed? 18 00:06:37,275 --> 00:07:26,400 [Dr. Ada Shannon] OLMo and OLMoE recipes, a Dolma sample for the dense models, OLMoE-MIX for the MoE, everything at 500 billion tokens. They report V2 and V3 validation loss plus average zero-shot accuracy, and they disclose dropping benchmarks that swing more than 20 percent between neighbouring checkpoints. In Table 1, baseline OLMo-1B has V2 loss 2.811 and accuracy 62.5. DHC at n=1 is worse: 2.819 and 62.3. Then x2 gives 2.802 and 63.0, x4 gives 2.781 and 63.8, x8 gives 2.778 and 62.8. The no-tanh variants score 62.3, 63.8, 64.4, 63.8. Loss improves through four and saturates at eight, while the accuracy average isn't monotone. No DHC run had loss spikes. 19 00:07:26,400 --> 00:07:36,800 [Hal Turing] Loss dropping in order across four expansion rates is a tidy result, and I'll say so plainly. What about static versus dynamic, and the rivals? 20 00:07:36,800 --> 00:08:19,426 [Dr. Ada Shannon] Static x2 gets 63.4 against dynamic's 63.0. At x4 it's static 63.6, dynamic 63.8, no-tanh 64.4, with dynamic's edge mostly in V2 loss, 2.781 versus 2.791. Freezing the width-connections costs about 0.02 V2 and 0.017 V3, and freezing B hurts less. Against rivals at n=2, AltUp from Baykal and colleagues at Google and ResiDual from Xie and colleagues at Microsoft Research, both 2023, start ahead but get overtaken by the baseline: V2 2.827 and 2.825 versus 2.811, with DHC x2 at 2.802. Those are the only competitors in the paper. 21 00:08:19,426 --> 00:08:25,101 [Hal Turing] Now scale it up. Seven billion dense first, then the MoE headline. 22 00:08:25,101 --> 00:09:27,026 [Dr. Ada Shannon] OLMo-7B-DHC x4 gets V2 loss 2.559 against 2.581, V3 2.304 against 2.322, and task accuracy 71.0 against 70.1, with no spikes where the baseline had frequent ones. There's no token-savings factor for 7B. On OLMoE, training loss is lower by about 0.027 and C4 validation loss by 0.028, both labelled 1.8x in the figure. Two different quantities hide in that label. A fixed-token gap is the loss difference at the same token count. A token-equivalence speedup is how many fewer tokens DHC needs to reach the baseline's loss. Downstream, ARC-Challenge goes from 41.8 to 47.8, ARC-Easy 72.8 to 76.7, BoolQ 65.4 to 68.5, and MMLU-Var gains 1.2. HellaSwag gains 0.7, PIQA 0.6, WinoGrande 0.2. The paper says that in many metrics DHC needs only half the training tokens. 23 00:09:27,026 --> 00:09:30,276 [Hal Turing] And did the learned weights show any structure? 24 00:09:30,276 --> 00:10:27,776 [Dr. Ada Shannon] Unfolding the connection matrix for OLMo-1B-DHC x4 at 500B gives a dense lower-triangular map shaped like a lambda. There's decay over distance, which is Post-Norm-like, plus heavy reuse of the earliest layers, which is Pre-Norm-like. The input embedding feeds most layers and is removed before the output. Parallel pairs emerge, layers 11 and 12 for instance. Attention outputs have few long-range connections while FFN outputs are larger, resembling the two-hop residual from Ma and colleagues at Meta, 2024. Vision, briefly: on the Peebles and Xie diffusion transformer, DiT-XL/2 with static HC x2 gets FID 2.18 against 2.36, near the 50-percent-larger DiT-1B/2 at 2.13. ViT-Large rises from 77.25 to 79.94, but on ViT-Base static 77.60 beats dynamic 77.26, and the gain shrinks over epochs. 25 00:10:27,776 --> 00:10:46,851 [Hal Turing] Ada, I want to understand one number before anything else. The gap at 500 billion tokens is 0.027 in loss. Here's my genuine question: how does a gap that small turn into nearly twofold savings? I went looking in Figure 1 for the matched loss level and couldn't find it. 26 00:10:46,851 --> 00:11:24,951 [Dr. Ada Shannon] You won't find it. The paper never says which loss was matched or at what token count. My guess is 500 divided by 1.8, so about 278 billion tokens of the hyper-connection model reaching the baseline's 500-billion-token loss. That's my inference, not their statement. It's fragile because late in training the curve is nearly flat, so a tiny vertical gap becomes a huge horizontal one. At 200 or 300 billion tokens, where the curves are steeper, the same gap would give a much smaller factor, and we're never shown that. And 500 billion is well short of a full OLMoE budget. Then there's the bigger problem, which is that every one of those— 27 00:11:24,951 --> 00:11:49,351 [Hal Turing] Sorry to cut you off, but every one of those is a single run, right? The downstream column goes 62.5, 62.3, 63.0, 63.8, 62.8, which doesn't track the loss ordering. The authors also admit they dropped benchmarks that swung more than twenty percent between neighbouring checkpoints, after they saw the noise. How much of that ordering is seed? 28 00:11:49,351 --> 00:12:31,401 [Dr. Ada Shannon] Nobody can say. There are no seeds and no intervals. The no-tanh variants beat the tanh ones at both two and four streams, and the paper ships the variant that lost. The loss gaps of 0.01 to 0.03 are a steadier signal than accuracy averages, but they're also unreplicated. Same-named rows disagree across tables too. DHC times two without tanh has a V3 loss of 2.537 in Table 1 and 2.529 in Table 2. Probably transcription, but it's why per-run logs matter. The bigger issue is attribution. DHC at one stream loses to the baseline. The gain only appears once you carry n streams, at that extra training memory. Is that learned connection strength, or just a wider residual? 29 00:12:31,401 --> 00:12:42,551 [Hal Turing] So they widened the highway and credited the toll gates. And the toll gates themselves are a 2015 idea, if I remember right. Which controls should they have run? 30 00:12:42,551 --> 00:13:27,451 [Dr. Ada Shannon] Yes, Highway. DHC is essentially that plus cross-stream mixing. Neither Highway nor DenseNet is cited, even though the unfolded matrix in Section 4.5 is a dense connection pattern. The closest single-stream rival is DenseFormer, which learns scalar weights over all earlier layer outputs. The paper's DenseFormer citation is a person re-identification model, a different work. ReZero, by Bachlechner at UC San Diego in 2021, learns one scalar gate. DeepNet, by Hongyu Wang at Microsoft in 2022, fixes the scaling for a thousand layers. The only rivals they run are AltUp and ResiDual. The baselines also lack the square-root-n init and the weight-decay exemption. The 7B baseline spikes, but we're not told whether it has QK-norm or a lowered learning rate. That may be a fragile baseline rather than HC stability. 31 00:13:27,451 --> 00:13:31,701 [Hal Turing] And the systems side? They call the cost negligible. 32 00:13:31,701 --> 00:14:13,926 [Dr. Ada Shannon] Parameters and FLOPs, yes. But the activation memory hit is real, and there's no throughput number anywhere. Norms, tanh and tiny mixes are memory-bound, so step time could eat much of the token savings. Tokens aren't wall-clock. Serving is silent too: every layer moves n times the hidden state. Logit lens, early exit and layer skipping all assume the single residual stream from the Mathematical Framework paper by Elhage at Anthropic, 2021. Follow-ups exist. DeepSeek's mHC reportedly constrains the mixing matrices for stability at scale, and ByteDance Seed has pursued lower-memory variants. I'm naming both from memory, so check them. If you're planning a pretrain, ablate it against DenseFormer and Peri-LN with several seeds. It isn't a drop-in yet. 33 00:14:13,926 --> 00:14:38,101 [Hal Turing] So what the paper shows is consistent small validation-loss gains, near-zero parameters and FLOPs, and a learned lambda-shaped wiring pattern worth studying. What it asserts is 1.8 times, a resolved seesaw, and generality, and those need seeds, matched-loss curves and stronger baselines. The attribution question stays open. Thanks for listening, everyone. Goodbye.