Transformers use learned attention for sequence modeling but fixed summation for depth aggregation. This visualization shows the architectural inconsistency.
Transformers revolutionized sequence modeling by replacing fixed RNN recurrence with learned attention. But depth-wise aggregation still uses the same fixed summation from 2015 ResNet. AttnRes completes the transition to learned weights across all architectural dimensions.
In PreNorm architectures, hidden-state magnitudes grow linearly with depth, causing early-layer contributions to be progressively buried. This heatmap shows layer contribution percentages across depth.
Hidden states accumulate unnormalized: ||h_L|| ≈ L × ||v_single||. At layer 100, individual layer output is ~1% of total magnitude.
Softmax attention produces weights summing to 1.0, maintaining bounded magnitude: ||h_L|| ≈ ||v_single|| regardless of depth.
Full AttnRes attends over all layer outputs. Block AttnRes partitions layers into blocks for memory efficiency. Toggle between variants to see the difference.
Memory: O(L×d) activations. Each layer attends over all previous layer outputs. Free in standard training, expensive with activation recomputation.
Memory: O(N×d) block summaries where N≪L. Attention over N blocks instead of L layers. ~3% compute overhead, <2% inference latency.
Block AttnRes matches baseline performance with 1.25× less compute. Results from 460M to 7B parameters on up to 100B tokens.
Kimi Linear architecture (48B params, 3B activated MoE) trained on 1.4T tokens. All benchmarks show improvement over standard residual baseline.
AttnRes produces more uniform gradient norms across depth compared to standard residuals, indicating all layers contribute meaningfully to training.
Heatmap shows which layer outputs are attended to at each depth. Darker cells indicate higher attention weights. Hover to see exact values.
Strong attention to immediate predecessors and first few layers. Feature extraction phase maintains short-range dependencies.
Selective attention to mid-network representations (layers 20-40). Late layers retrieve high-level features while downweighting early low-level outputs.