Gated Convolutions Challenge RNNs in Language Modeling

Dauphin, Fan, Auli, Grangier · Facebook AI Research · ICML 2017 · “Language Modeling with Gated Convolutional Networks” arXiv 1612.08083 Interactive viz

Model pipeline — step through it

Token → embedding → gated causal conv blocks (bottleneck + residual) → adaptive softmax.

Autoregressive factorization (unchanged)

Only the context encoder differs: RNN chain vs conv stack.

Hierarchical causal receptive field

Click a top-layer node to trace what it sees. Front of the sequence is zero-padded, so no future peeking.

Layers Kernel k

Visibility heatmap (output i × input j)

Upper triangle = masked future. Hover a cell.

Hover a cell.

Sequential operations

RNN: O(N) steps. Conv: O(N/k) per the discussion, all positions parallel within a layer.

The Gated Linear Unit

h(X) = (X·W + b) ⊗ σ(X·V + c). One branch stays linear.

Activation and gradient (scalar, W=V=1)

Solid: f(x). Dashed: f′(x).

Gate explorer

Drag the linear input a and gate logit z.

a z

Convergence by gating (illustrative curves; GLU converges faster and lower)

Ablation: remove the gate

Test perplexity, 100-hour training budget. Linear < KN-5 < bilinear < GLU.

Results

Context size returns (illustrative shape of Figure 4)

Gains flatten past ~20–40 words, even on WikiText-103 (~4000 tokens per doc).

One idea moving between labs

Click a node. The architecture faded; the gate survived.

Where the claim stands

Hover a bar for the framing caveat.

References