Attention Residuals: Rethinking Depth-Wise Aggregation

Moonshot AI (Kimi Team) • March 2026 arXiv:2603.15031
Overview
PreNorm Dilution
Architecture
Performance
Analysis

The Asymmetry Problem

Transformers use learned attention for sequence modeling but fixed summation for depth aggregation. This visualization shows the architectural inconsistency.

Sequence Dimension (2017) RNN (Before) h_t = f(h_{t-1}, x_t) Fixed recurrence Uniform weights Learned Attention (Now) α = softmax(QK^T) Content-based Adaptive weights Depth Dimension (2015-2026) Standard ResNet h_l = Σ v_i Fixed summation Uniform weights (1.0) Learned AttnRes (2026) h_l = Σ α_i v_i Softmax attention Learned weights Expert Routing (MoE) Fixed Routing All experts equal Uniform selection Learned Learned Routing Gating network Content-based Evolution Timeline 2015 ResNet 2017 Attention 2020 PreNorm 2026 AttnRes

Key Insight

Transformers revolutionized sequence modeling by replacing fixed RNN recurrence with learned attention. But depth-wise aggregation still uses the same fixed summation from 2015 ResNet. AttnRes completes the transition to learned weights across all architectural dimensions.

PreNorm Dilution Problem

In PreNorm architectures, hidden-state magnitudes grow linearly with depth, causing early-layer contributions to be progressively buried. This heatmap shows layer contribution percentages across depth.

0-1% contribution
5-10%
20-30%
40-50%
>50%

Magnitude Growth (PreNorm)

Hidden states accumulate unnormalized: ||h_L|| ≈ L × ||v_single||. At layer 100, individual layer output is ~1% of total magnitude.

Bounded Magnitude (AttnRes)

Softmax attention produces weights summing to 1.0, maintaining bounded magnitude: ||h_L|| ≈ ||v_single|| regardless of depth.

Architecture Comparison

Full AttnRes attends over all layer outputs. Block AttnRes partitions layers into blocks for memory efficiency. Toggle between variants to see the difference.

Memory & Compute Overhead

Full AttnRes

Memory: O(L×d) activations. Each layer attends over all previous layer outputs. Free in standard training, expensive with activation recomputation.

Block AttnRes (Production)

Memory: O(N×d) block summaries where N≪L. Attention over N blocks instead of L layers. ~3% compute overhead, <2% inference latency.

Scaling Law & Performance Results

Block AttnRes matches baseline performance with 1.25× less compute. Results from 460M to 7B parameters on up to 100B tokens.

48B Parameter Model: Downstream Task Performance

Kimi Linear architecture (48B params, 3B activated MoE) trained on 1.4T tokens. All benchmarks show improvement over standard residual baseline.

Gradient Distribution Analysis

AttnRes produces more uniform gradient norms across depth compared to standard residuals, indicating all layers contribute meaningfully to training.

Layer Attention Weight Patterns

Heatmap shows which layer outputs are attended to at each depth. Darker cells indicate higher attention weights. Hover to see exact values.

Early Layers (1-20)

Strong attention to immediate predecessors and first few layers. Feature extraction phase maintains short-range dependencies.

Deep Layers (60-100)

Selective attention to mid-network representations (layers 20-40). Late layers retrieve high-level features while downweighting early low-level outputs.

Key References