A visual companion to the episode on the 2021 Google Research paper that tracked the shift from code-like text generation to executable short-program synthesis in Python, emphasizing run-and-pass-tests evaluation over string similarity.
Natural language, visible asserts, model generation, and an unforgiving execution gate. Hover the stages to see where synthesis fails or succeeds.
The key distinction from autocomplete is the rightmost box: code must execute and satisfy tests.
A compact contrast between constrained search-based synthesis and LM-based emission in general-purpose Python.
MBPP emphasizes short utility functions and imperative Python patterns; MathQA-Python is larger and more regular, with arithmetic word-problem structure. Hover the cells.
Rows are problem archetypes, columns are coding features. Color intensity indicates how common each pattern is in the selected benchmark.
Visible asserts provide both instruction signal and evaluation target. This is why prompt formatting mattered so much.
Model size tracks substantial gains. Toggle benchmark and regime. The plotted values are realistic mock reconstructions anchored to numbers discussed in the episode.
The central scientific signal: larger models become materially better synthesizers under execution-based scoring.
MBPP few-shot peaks at 59.6%. MathQA-Python few-shot is much lower, but fine-tuning climbs much higher on the larger, more regular dataset.
The transcript highlights two things at once: iterative repair helps a lot, while direct “mental execution” remains weak. Use the slider and toggles.
The paper reports improvement from roughly 30% to 65% over four turns on a dialog subset. Drag the turn slider.
A model may emit code that passes tests yet still perform poorly when asked to predict program outputs directly.
Compact citations for the paper and nearby work discussed around synthesis, code models, execution feedback, and evaluation.