Program Synthesis with Large Language Models

A visual companion to the episode on the 2021 Google Research paper that tracked the shift from code-like text generation to executable short-program synthesis in Python, emphasizing run-and-pass-tests evaluation over string similarity.

Google Research · 2021 Benchmarks: MBPP · MathQA-Python Few-shot · Fine-tuning · Execution-based eval arXiv:2108.07732

Overview: From English Spec to Executed Python

Natural language, visible asserts, model generation, and an unforgiving execution gate. Hover the stages to see where synthesis fails or succeeds.

Synthesis pipeline

The key distinction from autocomplete is the rightmost box: code must execute and satisfy tests.

specification / successful path transformer generation execution and test harness failure modes

Classical search vs transformer generation

A compact contrast between constrained search-based synthesis and LM-based emission in general-purpose Python.

Benchmark anatomy

MBPP emphasizes short utility functions and imperative Python patterns; MathQA-Python is larger and more regular, with arithmetic word-problem structure. Hover the cells.

Task-feature heatmap

Rows are problem archetypes, columns are coding features. Color intensity indicates how common each pattern is in the selected benchmark.

Prompt template microscope

Visible asserts provide both instruction signal and evaluation target. This is why prompt formatting mattered so much.

Scaling results

Model size tracks substantial gains. Toggle benchmark and regime. The plotted values are realistic mock reconstructions anchored to numbers discussed in the episode.

Accuracy vs model scale

The central scientific signal: larger models become materially better synthesizers under execution-based scoring.

Benchmark headline deltas

MBPP few-shot peaks at 59.6%. MathQA-Python few-shot is much lower, but fine-tuning climbs much higher on the larger, more regular dataset.

Interaction, execution feedback, and semantic gaps

The transcript highlights two things at once: iterative repair helps a lot, while direct “mental execution” remains weak. Use the slider and toggles.

Multi-turn dialog improvement

The paper reports improvement from roughly 30% to 65% over four turns on a dialog subset. Drag the turn slider.

turn: 4success: 65%remaining error: 35%

Generate code ≠ reliably simulate execution

A model may emit code that passes tests yet still perform poorly when asked to predict program outputs directly.

References

Compact citations for the paper and nearby work discussed around synthesis, code models, execution feedback, and evaluation.