This episode examines "Language Modeling with Gated Convolutional Networks" (Dauphin et al., Facebook AI Research, ICML 2017), which replaces recurrent architectures with stacked causal convolutions for language modeling. The discussion covers why LSTMs dominated the field due to vanishing-gradient mitigation through gating, and contrasts that with the paper's Gated Linear Unit — a sigmoid-gated linear projection that preserves an undiminished gradient path, unlike the tanh-based gating in DeepMind's PixelCNN. A central debate weighs the RNN's theoretically unbounded context against the convolutional model's finite but better-conditioned receptive field, testing whether structured, hierarchical context outperforms raw sequential memory. The hosts walk through the model's pipeline — embeddings, bottlenecked causal convolution blocks with residual connections, and an adaptive softmax — and preview results on the Google Billion Word benchmark and the long-range WikiText-103 test. Listeners interested in the shift toward parallelizable, non-recurrent sequence models will find the tension between theoretical and practical memory capacity especially compelling.
Sources:
1. Language Modeling with Gated Convolutional Networks — Yann N. Dauphin, Angela Fan, Michael Auli, David Grangier, 2016
http://arxiv.org/abs/1612.080832. Exploring the Limits of Language Modeling — Jozefowicz, Vinyals, Schuster, Shazeer, Wu, 2016
https://scholar.google.com/scholar?q=Exploring+the+Limits+of+Language+Modeling3. Conditional Image Generation with PixelCNN Decoders — van den Oord, Kalchbrenner, Vinyals, Espeholt, Graves, Kavukcuoglu, 2016
https://scholar.google.com/scholar?q=Conditional+Image+Generation+with+PixelCNN+Decoders4. Neural Machine Translation in Linear Time — Kalchbrenner, Espeholt, Simonyan, van den Oord, Graves, Kavukcuoglu, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+in+Linear+Time5. Improving Neural Language Models with a Continuous Cache — Grave, Joulin, Usunier, 2016
https://scholar.google.com/scholar?q=Improving+Neural+Language+Models+with+a+Continuous+Cache6. Efficient softmax approximation for GPUs — Grave, Joulin, Cissé, Grangier, Jégou, 2016
https://scholar.google.com/scholar?q=Efficient+softmax+approximation+for+GPUs7. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Shazeer, Mirhoseini, Maziarz, Davis, Le, Hinton, Dean, 2017
https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+LayerInteractive Visualization: Gated Convolutions Challenge RNNs in Language Modeling