Model pipeline — step through it
Token → embedding → gated causal conv blocks (bottleneck + residual) → adaptive softmax.
Autoregressive factorization (unchanged)
Only the context encoder differs: RNN chain vs conv stack.
Hierarchical causal receptive field
Click a top-layer node to trace what it sees. Front of the sequence is zero-padded, so no future peeking.
Layers Kernel k
Visibility heatmap (output i × input j)
Upper triangle = masked future. Hover a cell.
Hover a cell.
Sequential operations
RNN: O(N) steps. Conv: O(N/k) per the discussion, all positions parallel within a layer.
The Gated Linear Unit
h(X) = (X·W + b) ⊗ σ(X·V + c). One branch stays linear.
Activation and gradient (scalar, W=V=1)
Solid: f(x). Dashed: f′(x).
Gate explorer
Drag the linear input a and gate logit z.
a z
Convergence by gating (illustrative curves; GLU converges faster and lower)
Ablation: remove the gate
Test perplexity, 100-hour training budget. Linear < KN-5 < bilinear < GLU.
Results
Context size returns (illustrative shape of Figure 4)
Gains flatten past ~20–40 words, even on WikiText-103 (~4000 tokens per doc).
One idea moving between labs
Click a node. The architecture faded; the gate survived.
Where the claim stands
Hover a bar for the framing caveat.
References
- Dauphin et al. — Language Modeling with Gated Convolutional Networks (1612.08083)
- Jozefowicz et al. — Exploring the Limits of Language Modeling (1602.02410)
- van den Oord et al. — Conditional Image Generation with PixelCNN Decoders (1606.05328)
- Kalchbrenner et al. — Neural Machine Translation in Linear Time (1610.10099)
- Grave et al. — Improving Neural Language Models with a Continuous Cache (1612.04426)
- Grave et al. — Efficient softmax approximation for GPUs (1609.04309)
- Shazeer et al. — Outrageously Large Neural Networks: Sparsely-Gated MoE (1701.06538)