This episode revisits Teuvo Kohonen's 1972 paper "Correlation Matrix Memories," which reframes associative memory as a hardware fault-tolerance problem rather than a representation-learning one. Kohonen builds a memory from outer-product sums of key and data vectors, then shows mathematically how much recall quality degrades when connections are randomly dropped (an "incomplete" correlation matrix memory) rather than fully wired. The discussion traces the paper's lineage against optical holography models and Steinbuch's Lernmatrix, and unpacks concepts like crosstalk and graceful degradation as information gets smeared additively across the matrix instead of stored in one fragile spot. A tangent draws — and partly disputes — a comparison between Kohonen's outer-product accumulation and the mechanics underlying modern attention, debating whether the resemblance is structural or purely coincidental given the total absence of learning or gradients in the original scheme. Listeners interested in the deep history of neural memory models and how old hardware constraints shaped ideas that echo in today's architectures will find plenty to chew on.
Sources:
1. Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention
https://lucidar.me/fr/neural-networks/files/1972-correlation-matrix-memories.pdf2. Neural networks and physical systems with emergent collective computational abilities — John J. Hopfield, 1982
https://scholar.google.com/scholar?q=Neural+networks+and+physical+systems+with+emergent+collective+computational+abilities3. Non-Holographic Associative Memory — David Willshaw, O. P. Buneman, H. Christopher Longuet-Higgins, 1969
https://scholar.google.com/scholar?q=Non-Holographic+Associative+Memory4. Hopfield Networks is All You Need — Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, et al., 2020
https://scholar.google.com/scholar?q=Hopfield+Networks+is+All+You+Need5. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021
https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers6. A Simple Neural Network Generating an Interactive Memory — James A. Anderson, 1972
https://scholar.google.com/scholar?q=A+Simple+Neural+Network+Generating+an+Interactive+Memory7. Representation of Associated Data by Matrix Operators — Teuvo Kohonen, Matti Ruohonen, 1973
https://scholar.google.com/scholar?q=Representation+of+Associated+Data+by+Matrix+Operators8. Sparse Distributed Memory — Pentti Kanerva, 1988
https://scholar.google.com/scholar?q=Sparse+Distributed+Memory9. Die Lernmatrix — K. Steinbuch, 1961
https://scholar.google.com/scholar?q=Die+Lernmatrix10. Associative holographic memories — D. Gabor, 1969
https://scholar.google.com/scholar?q=Associative+holographic+memories11. A class of randomly organized associative memories — T. Kohonen, 1971
https://scholar.google.com/scholar?q=A+class+of+randomly+organized+associative+memoriesInteractive Visualization: Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention