← All episodes Hyper-Connections: Learning Residual Strengths for Faster Transformer Training

Hyper-Connections: Learning Residual Strengths for Faster Transformer Training

Sep 28, 2026
This episode examines Hyper-Connections (ICLR 2025, ByteDance Seed), which replaces the fixed residual connection in transformers with learned connection strengths. It traces the "seesaw" between Pre-Norm and Post-Norm: Pre-Norm gives stable gradients but suffers representation collapse in deep layers, while Post-Norm does the reverse. The method carries n parallel residual streams and learns depth-connections and width-connections, optionally predicted per token. It initializes as exactly Pre-Norm, so it strictly generalizes both variants. The paper's headline claim is 1.8x faster convergence and about six extra points on ARC-Challenge for OLMoE-1B-7B, for roughly 0.03% more parameters and 0.2% more FLOPs. The hosts place the idea against Highway Networks, DenseNet, and DenseFormer, and they flag that the collapse evidence rests largely on a single cosine-similarity figure. Listeners get a clear view of how a small learned wiring matrix can change a core piece of transformer design, along with a skeptical read of the reported gains.
Sources:
1. Hyper-Connections — Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, Xun Zhou, 2024
http://arxiv.org/abs/2409.19606
2. Deep Residual Learning for Image Recognition — Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun, 2016
https://scholar.google.com/scholar?q=Deep+Residual+Learning+for+Image+Recognition
3. On Layer Normalization in the Transformer Architecture — Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, Tie-Yan Liu, 2020
https://scholar.google.com/scholar?q=On+Layer+Normalization+in+the+Transformer+Architecture
4. Understanding the Difficulty of Training Transformers — Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, Jiawei Han, 2020
https://scholar.google.com/scholar?q=Understanding+the+Difficulty+of+Training+Transformers
5. Densely Connected Convolutional Networks (DenseNet), with DenseFormer as the transformer analogue — Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Q. Weinberger (DenseNet, 2017); Matteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin Jaggi (DenseFormer, 2024), 2017 / 2024
https://scholar.google.com/scholar?q=Densely+Connected+Convolutional+Networks+%28DenseNet%29%2C+with+DenseFormer+as+the+transformer+analogue
6. DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging — Matteo Pagliardini, Amirkeivan Mohtashami, Francois Fleuret, Martin Jaggi, 2024
https://scholar.google.com/scholar?q=DenseFormer%3A+Enhancing+Information+Flow+in+Transformers+via+Depth+Weighted+Averaging
7. Highway Networks / Training Very Deep Networks — Rupesh Srivastava, Klaus Greff, Jürgen Schmidhuber, 2015
https://scholar.google.com/scholar?q=Highway+Networks+%2F+Training+Very+Deep+Networks
8. Densely Connected Convolutional Networks (DenseNet) — Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Weinberger, 2017
https://scholar.google.com/scholar?q=Densely+Connected+Convolutional+Networks+%28DenseNet%29
9. DeepNet: Scaling Transformers to 1,000 Layers (DeepNorm) — Hongyu Wang et al., 2022
https://scholar.google.com/scholar?q=DeepNet%3A+Scaling+Transformers+to+1%2C000+Layers+%28DeepNorm%29
10. ReZero is All You Need: Fast Convergence at Large Depth — Thomas Bachlechner et al., 2021
https://scholar.google.com/scholar?q=ReZero+is+All+You+Need%3A+Fast+Convergence+at+Large+Depth
11. Peri-LN: Revisiting Layer Normalization in the Transformer Architecture / Sandwich-norm variants — Jeonghoon Kim et al., 2025
https://scholar.google.com/scholar?q=Peri-LN%3A+Revisiting+Layer+Normalization+in+the+Transformer+Architecture+%2F+Sandwich-norm+variants
12. The Curse of Depth in Large Language Models — Wenfang Sun et al., 2025
https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models
13. Transformer Layers as Painters — Qi Sun, Marc Pickett, et al., 2024
https://scholar.google.com/scholar?q=Transformer+Layers+as+Painters
14. AltUp: Alternating Updates for Efficient Transformers — Cenk Baykal et al., 2023/2024
https://scholar.google.com/scholar?q=AltUp%3A+Alternating+Updates+for+Efficient+Transformers
15. ResiDual: Transformer with Dual Residual Connections — Shufang Xie et al., 2023
https://scholar.google.com/scholar?q=ResiDual%3A+Transformer+with+Dual+Residual+Connections
16. Value Residual Learning / ResFormer — Zhanchao Zhou et al., 2024
https://scholar.google.com/scholar?q=Value+Residual+Learning+%2F+ResFormer
17. A Mathematical Framework for Transformer Circuits (residual stream view) — Nelson Elhage et al., 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits+%28residual+stream+view%29
18. Reducing Activation Recomputation in Large Transformer Models — Vijay Korthikanti et al., 2022
https://scholar.google.com/scholar?q=Reducing+Activation+Recomputation+in+Large+Transformer+Models
19. mHC: Manifold-Constrained Hyper-Connections — DeepSeek-AI, 2025/2026
https://scholar.google.com/scholar?q=mHC%3A+Manifold-Constrained+Hyper-Connections
20. Frac-Connections / Virtual Width Networks (ByteDance Seed follow-ups) — ByteDance Seed team, 2025
https://scholar.google.com/scholar?q=Frac-Connections+%2F+Virtual+Width+Networks+%28ByteDance+Seed+follow-ups%29
Interactive Visualization: Hyper-Connections: Learning Residual Strengths for Faster Transformer Training