This episode explores Structural Language Models of Code, a 2020 paper from Technion, Tel Aviv University and Facebook AI Research that tackles "any-code completion": predicting a missing piece of a program with no restriction on vocabulary or structure. It traces the history from 1969-era automatic programming through domain-specific systems like FlashFill and DeepCoder, and prior general-language work that limited APIs, types or syntax. Then it contrasts a subtoken sequence-to-sequence baseline with the paper's approach of modeling code as an abstract syntax tree. That approach applies the language-model chain rule over a depth-first tree traversal, using partial AST paths that end at the node being generated, which extends the authors' earlier code2vec and code2seq work from reading code to writing it. The discussion covers the headline exact-match gains, Java accuracy@1 of 18.04 versus 16.93 and C# 37.61 versus 26.42. It also raises the question of whether a one-point win is meaningful, given that exact match undercounts logically equivalent code. Listeners interested in how syntax-aware models can generate code with unseen identifiers will find the mechanics and the skeptical read of the metrics useful.
Sources:
1. Structural Language Models of Code: Any-Code Completion with Trees
https://arxiv.org/pdf/1910.005772. Structured Generative Models of Natural Source Code — Chris J. Maddison, Daniel Tarlow, 2014
https://scholar.google.com/scholar?q=Structured+Generative+Models+of+Natural+Source+Code3. On the Naturalness of Software — Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, Premkumar Devanbu, 2012
https://scholar.google.com/scholar?q=On+the+Naturalness+of+Software4. A Syntactic Neural Model for General-Purpose Code Generation — Pengcheng Yin, Graham Neubig, 2017
https://scholar.google.com/scholar?q=A+Syntactic+Neural+Model+for+General-Purpose+Code+Generation5. Generative Code Modeling with Graphs — Marc Brockschmidt, Miltiadis Allamanis, Alexander L. Gaunt, Oleksandr Polozov, 2019
https://scholar.google.com/scholar?q=Generative+Code+Modeling+with+Graphs6. Code Completion with Statistical Language Models — Veselin Raychev, Martin Vechev, Eran Yahav, 2014
https://scholar.google.com/scholar?q=Code+Completion+with+Statistical+Language+Models7. IntelliCode Compose: Code Generation Using Transformer — Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, Neel Sundaresan, 2020
https://scholar.google.com/scholar?q=IntelliCode+Compose%3A+Code+Generation+Using+Transformer8. Efficient Training of Language Models to Fill in the Middle — Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen, 2022
https://scholar.google.com/scholar?q=Efficient+Training+of+Language+Models+to+Fill+in+the+Middle9. A General Path-Based Representation for Predicting Program Properties — Uri Alon, Meital Zilberstein, Omer Levy, Eran Yahav, 2018
https://scholar.google.com/scholar?q=A+General+Path-Based+Representation+for+Predicting+Program+Properties10. code2vec: Learning Distributed Representations of Code — Uri Alon, Meital Zilberstein, Omer Levy, Eran Yahav, 2019
https://scholar.google.com/scholar?q=code2vec%3A+Learning+Distributed+Representations+of+Code11. code2seq: Generating Sequences from Structured Representations of Code — Uri Alon, Shaked Brody, Omer Levy, Eran Yahav, 2019
https://scholar.google.com/scholar?q=code2seq%3A+Generating+Sequences+from+Structured+Representations+of+Code12. Learning to Represent Programs with Graphs — Miltiadis Allamanis, Marc Brockschmidt, Mahmoud Khademi, 2018
https://scholar.google.com/scholar?q=Learning+to+Represent+Programs+with+Graphs13. Abstract Syntax Networks for Code Generation and Semantic Parsing — Maxim Rabinovich, Mitchell Stern, Dan Klein, 2017
https://scholar.google.com/scholar?q=Abstract+Syntax+Networks+for+Code+Generation+and+Semantic+Parsing14. PHOG: Probabilistic Model for Code — Pavol Bielik, Veselin Raychev, Martin Vechev, 2016
https://scholar.google.com/scholar?q=PHOG%3A+Probabilistic+Model+for+Code15. Incorporating Copying Mechanism in Sequence-to-Sequence Learning — Jiatao Gu, Zhengdong Lu, Hang Li, Victor O.K. Li, 2016
https://scholar.google.com/scholar?q=Incorporating+Copying+Mechanism+in+Sequence-to-Sequence+Learning16. On the Bottleneck of Graph Neural Networks and its Practical Implications — Uri Alon, Eran Yahav, 2020
https://scholar.google.com/scholar?q=On+the+Bottleneck+of+Graph+Neural+Networks+and+its+Practical+Implications17. InCoder: A Generative Model for Code Infilling and Synthesis — Daniel Fried et al., 2022
https://scholar.google.com/scholar?q=InCoder%3A+A+Generative+Model+for+Code+Infilling+and+Synthesis18. Evaluating Large Language Models Trained on Code (Codex / HumanEval) — Mark Chen et al., 2021
https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code+%28Codex+%2F+HumanEval%2919. RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation — Fengji Zhang et al., 2023
https://scholar.google.com/scholar?q=RepoCoder%3A+Repository-Level+Code+Completion+Through+Iterative+Retrieval+and+Generation20. UniXcoder: Unified Cross-Modal Pre-training for Code Representation — Daya Guo et al., 2022
https://scholar.google.com/scholar?q=UniXcoder%3A+Unified+Cross-Modal+Pre-training+for+Code+RepresentationInteractive Visualization: Structural Language Models of Code: Any-Code Completion with Trees