AI Post Transformers · Episode Companion

Hyper-Connections: Learning Residual Strengths for Faster Transformer Training

ByteDance Seed · ICLR 2025 OLMoE-1B-7B · 500B tokens arXiv:2409.19606 ↗

Hyper-Connections replace the fixed residual connection with a learned wiring matrix across n parallel streams. It initializes as exactly Pre-Norm, so it strictly generalizes the Pre-Norm / Post-Norm seesaw — at the cost of real activation memory the paper doesn't headline. This page visualizes the mechanism, the reported results, and where the evidence gets thin.

Headline numbers

OLMoE-1B-7B, dynamic Hyper-Connections, expansion rate n=4, 500B tokens
1.8×
claimed convergence speedup
+6.0
pts on ARC-Challenge
0.03%
extra parameters
0.2%
extra FLOPs
+26%
train memory, OLMo-1B (unheadlined)

From one residual stream to n streams

Toggle the expansion rate to see how many parallel streams a layer carries.
residual stream
attention / FFN block (unchanged)
learned connection matrix
Key point: the block itself (attention/FFN) is never touched. Only the wiring around it — how streams are read in and written back — is learned.

Pre-Norm vs Post-Norm vs Hyper-Connections

Same trade-off, three ways. Toggle to compare gradient stability against representation collapse.

Gradient magnitude by depth

Adjacent-layer cosine similarity (collapse signal, ~Fig. 3)

The seesaw: Pre-Norm keeps gradients flat and healthy but deep-layer states converge (high cosine similarity = collapse). Post-Norm avoids collapse but gradients fade toward the input without careful warmup. HC starts identical to Pre-Norm, then learns wiring that lowers similarity — the paper's evidence for this is one figure (Section on collapse, OLMo-1B).

The block connection matrix

Each block gets an (n+1)×(n+1) matrix: Am (read-in column), B (write-out row), Ar (n×n stream mixer).
Am (read-in)
B (write-out)
Ar (stream mix)

Unfolded wiring across depth — the “lambda” shape

Illustrative reconstruction of Section 4.5's OLMo-1B-DHC×4 (500B tokens) unfolded connection map: dense, lower-triangular, decaying with distance, with heavy reuse of the earliest layers.
Decay-with-distance looks Post-Norm-like; heavy early-layer reuse looks Pre-Norm-like. Parallel-pair patterns (e.g. layers 11/12) also emerge without being specified up front.

Table 1: OLMo-1B, expansion rate sweep

500B tokens. Toggle metric and tanh variant.
Note: no-tanh beats tanh at n=2 and n=4 accuracy, yet the paper ships the tanh variant. Single runs, no seeds reported.

How fragile is “1.8×”?

The paper never states which loss value was matched or at what token count. This is an illustrative reconstruction of why a flat late-training curve makes small vertical gaps look like huge horizontal (token) speedups. Pick a point on the curve.

Prior art the paper leans on — and doesn't cite

cited in the paper
not cited / miscited

Claim vs. evidence strength

Hover a cell for the reasoning.

References