Scaling Laws for Multilingual Code Pretraining

A visual companion to the episode’s central claim: code tokens are not equal. The page treats multilingual pretraining as a token-allocation control problem shaped by language-specific scaling slopes, bilingual transfer, and translation-style alignment.

Episode Signals
4core experiment blocks
7languages in allocation story
1.5Bheadline visible allocation scale
128Bfixed-budget bilingual control
Extracted arXiv IDs

Monolingual Scaling Landscape

Different languages keep paying off at different rates. The chart compresses the paper’s story into a single view: slope encodes how much more performance arrives with more tokens, while the lower bars hint at language-specific floors.

higher scaling utility cross-language hub earlier saturation
Highest slopePython
Tightest pairJS ↔ TS
Fast saturatorRust
Visible riskOvergeneralizing 14B

Bilingual Transfer Heatmap

Each cell compares training a target language with itself versus replacing half that budget with another language. Positive values mean useful transfer. Toggle between synergy, syntax overlap, and ecosystem overlap to see what the matrix seems to be rewarding.

Hover any cell for the target/source pair. The strongest patches are intentionally clustered around Java/C#, JavaScript/TypeScript, and Python-adjacent hubs, mirroring the episode’s interpretation.

Token Allocation Lab

Uniform shares are the control. The optimized mixes use mock scaling slopes plus pairwise transfer bonuses to reallocate the same total budget toward higher-utility languages and tighter bilingual neighborhoods.

Total tokens 128B
The right panel turns token shares into estimated downstream gains. The point is not exact forecasting. It is to show how non-uniform budgets can lift the average without needing every language to receive an equal share.

Parallel Pairing Flow

Translation-style concatenation tests a different mechanism from raw multilingual mixing. Instead of just adding another language, it injects explicit correspondence between snippets. Step through the pipeline to watch alignment become transfer.

The episode’s caution still applies: Python-centered pair construction can discover that Python is a strong hub without proving universal gains for every non-Python pair.

References

Compact source map for the episode and the visual framing.

Scaling Laws for Multilingual Code Pretraining Primary paper discussed in the episode. arXiv:2512.13472
Unsupervised Translation of Programming Languages Translation-style background for cross-language alignment. Scholar lookup
XLCoST / MultiPL-E / CRUXEval-X Benchmarks anchoring multilingual code evaluation beyond Python-only setups. XLCoST  ·  MultiPL-E  ·  CRUXEval-X
Code Scaling and Transfer Context Luo 2025 on code being data-hungry, plus later transfer and repository-level translation work. Code scaling  ·  Transfer study  ·  RepoTransBench