1 00:00:01,000 --> 00:00:36,990 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "Post-Training Language Models for Gold-Medal Performance in Coding Competitions," first author Aleksander Ficek, et al. — four more co-authors join him, including Boris Ginsburg — out of NVIDIA, posted to arXiv on September 2nd of this year. And Ada, I'm just going to read you one number before we do anything else: 535.4 out of 600. 2 00:00:36,990 --> 00:01:16,279 [Dr. Ada Shannon] And the top human contestant at that same event scored 498.27. Gold medal cutoff was 361.12. So yes, an NVIDIA model didn't just clear gold, it beat the single best human on the planet who showed up to compete that day. That's the headline of this paper — at the 2026 International Olympiad in Informatics, under the actual contest clock, actual submission limits, no internet beyond what a human contestant gets, their system called Ultra-CC put up a score no AI has ever posted against a live human field. Not a retrospective benchmark. Live, before the problems were even public. 3 00:01:16,279 --> 00:01:34,065 [Hal Turing] Okay, so before we get into how they pulled that off, I want to make the case for why anyone should care about a programming olympiad in the first place. Because on the surface it sounds like a flex — 'look, our model can solve puzzle problems.' What's actually hard about IOI-level problems? 4 00:01:34,065 --> 00:02:29,468 [Dr. Ada Shannon] It's the gap between recognizing an algorithm and inventing one. Most coding benchmarks people cite — HumanEval, MBPP — are pattern-completion tasks, you've probably seen something structurally similar in training data. IOI and ICPC problems are adversarially designed to require synthesizing a novel algorithmic idea under tight time and memory constraints, then implementing it correctly against test cases you never see. And it's not pass-fail — each problem is split into subtasks with different input constraints, so a brute-force solution that only handles small inputs still earns partial credit. Six problems, up to 100 points each, 600 max. That partial-credit structure is actually what makes the benchmark so informative — it separates 'has no idea' from 'has the right idea but the wrong complexity' from 'nailed it.' 5 00:02:29,468 --> 00:02:40,892 [Hal Turing] Wait wait wait — so the gold threshold you quoted earlier, 361.12, that's specific to 2026. Why does that number move around instead of being fixed? 6 00:02:40,892 --> 00:03:17,626 [Dr. Ada Shannon] Because it's set from the score distribution of that year's actual human field, not a fixed bar — roughly the top one-twelfth of contestants get gold, the next sixth silver, the next quarter bronze. In IOI 2025, gold was 438.3 on the same 600-point scale. So the threshold is a moving target tied to how hard that year's problem set turned out to be and how the human contestants performed against it, which is actually relevant later — but for now, just know 2025 and 2026 aren't directly comparable numbers, they're each calibrated to their own competition. 7 00:03:17,626 --> 00:03:24,778 [Hal Turing] Got it. So walk me through the pipeline at a high level — what's actually happening before you get a model that can do this? 8 00:03:24,778 --> 00:04:09,174 [Dr. Ada Shannon] Four stages. They start with 22,000 curated competitive programming problems pulled from real competition archives. Those get turned into synthetic reasoning traces — essentially, a strong teacher model works through each problem step by step, and that becomes training data. Then supervised fine-tuning on those traces. Then, for one of their two models, reinforcement learning on top of that. And then at inference time, a test-time compute strategy called GenCorrect that generates and refines candidate solutions before submitting — we'll get into exactly how that works later, but that's the four-stage shape: curate, synthesize, fine-tune, and optionally reinforce, then spend extra compute at answer time. 9 00:04:09,174 --> 00:04:11,821 [Hal Turing] And there are two models here, right? Nano and Ultra? 10 00:04:11,821 --> 00:05:12,240 [Dr. Ada Shannon] Right — Nemotron-3-Nano-CC, a 30-billion-parameter mixture-of-experts model with only 3 billion parameters active per token, which gets both SFT and RL. And Nemotron-3-Ultra-CC, 550 billion total parameters with 55 billion active, which only gets SFT, no RL — we'll get into why they skipped RL on the bigger model later, but it's basically a compute decision. Mixture-of-experts, for anyone who hasn't hit that term before, is an architecture where instead of every parameter firing for every token like a dense transformer, you route each token to a small subset of specialized sub-networks, called experts. Shazeer and colleagues at Google Brain introduced the modern version of this back in 2017 with their 'Outrageously Large Neural Networks' paper, and Google's Switch Transformer work in 2021 showed you could scale it to trillion-parameter territory while keeping inference cost tied to the much smaller active parameter count, not the total. 11 00:05:12,240 --> 00:05:18,927 [Hal Turing] So it's like having a huge team of specialists on staff, but only paying the ones who actually show up to that particular meeting. 12 00:05:18,927 --> 00:06:11,033 [Dr. Ada Shannon] That's basically it. Now, the training loop itself leans on a few ideas worth defining. Supervised fine-tuning, SFT, is just training the base model directly on curated input-output examples — here, those reasoning traces distilled from a teacher model, DeepSeek-V4-Flash, solving competition problems. That's a well-established recipe going back to instruction-tuning work like Ouyang and colleagues' InstructGPT paper out of OpenAI in 2022. Reinforcement learning from verifiable rewards, RLVR, is the newer piece — instead of a learned reward model guessing what's good, you actually compile and execute the generated code against real test cases and get a binary pass or fail signal. That verifiable, executable reward is the same spirit behind DeepSeek's R1 work in 2025, and it's what they apply to Nano-CC via a method called GRPO. 13 00:06:11,033 --> 00:06:18,231 [Hal Turing] And 'synthetic data generation' — that's the traces you mentioned, using a teacher model instead of humans writing solutions by hand? 14 00:06:18,231 --> 00:07:11,683 [Dr. Ada Shannon] Exactly — 1.2 million traces for Nano, 477,000-plus for Ultra, all generated rather than hand-authored, which is how you get training volume at this scale without an army of competitive programmers. It's the same instinct behind Microsoft's 'Textbooks Are All You Need' line of work with the Phi models a couple years back — synthetic, teacher-generated data can substitute for scarce human-authored examples if you curate it well. And the last piece, test-time compute scaling, is the idea that you don't have to bake all your gains into training — you can spend extra compute at inference by sampling many candidates and picking or refining the best one. Snell and colleagues at Berkeley formalized a lot of the tradeoffs there in 2024. GenCorrect is this paper's specific implementation of that idea, and it's where a huge chunk of the score jump actually comes from — but that's Part 2. 15 00:07:11,683 --> 00:07:28,866 [Hal Turing] Right, exactly that instinct. So before the traces even exist, though, somebody has to build the raw problem set they're generated from. Where does that 22,000 number actually come from, and how do you turn a scraped competition problem into something a model can actually train on? 16 00:07:28,866 --> 00:08:18,324 [Dr. Ada Shannon] They pull from 16 regional and international competition families going back roughly two decades, plus some online judge platforms, then run an agentic pipeline that packages each one into an executable evaluation environment — statement, constraints, test cases, auxiliary files, reference solutions, the whole harness. Then they validate: does a known-correct solution actually get full credit, does a known-wrong one actually fail? They even use gpt-oss-120b to generate throwaway candidate solutions just to stress-test the harness for broken statements or bad compiler configs. And critically, IOI 2025, ICPC 2025, and LiveCodeBench Pro are excluded and deduplicated out of that corpus entirely — otherwise you're training on your own exam. 17 00:08:18,324 --> 00:08:33,743 [Hal Turing] Okay, that contamination point matters a lot given the headline numbers later. So once you've got clean problems, the traces themselves — DeepSeek-V4-Flash is doing the generating. What's actually different between how Nano and Ultra get fed? 18 00:08:33,743 --> 00:09:15,678 [Dr. Ada Shannon] Nano gets 1.2 million traces and trains for three epochs; Ultra gets 477,642 traces and trains for just one. That's a huge asymmetry, and it's deliberate — Ultra starts from a much stronger base, an RLVR-teacher checkpoint already distilled for reasoning, so it needs less. There's also a detail that matters a lot later: some of those traces are self-improvement traces, where the teacher is handed a previous flawed solution and told to produce a better one. That's not incidental — it's explicitly there to expose the student to iterative-refinement behavior before it ever sees GenCorrect at inference time. They're pre-training the habit, not just the skill. 19 00:09:15,678 --> 00:09:29,656 [Hal Turing] So the model is basically getting a preview of its own future test-time strategy baked into the weights. That's a neat trick. What about the RL stage — you said earlier it's Nano-only. Walk me through why. 20 00:09:29,656 --> 00:10:16,700 [Dr. Ada Shannon] GRPO, applied only to Nano-CC. They start from four thousand problems with reliable executable environments, trim anything that takes too long to run, and land on 3,219 — split 2,847 for training and 372 for validation, split at the parent-problem level so subtasks from the same problem don't leak across the split. Reward is binary: full credit compiles and passes, or it doesn't, no partial-credit shaping, no reference-policy KL penalty at all. And Ultra just doesn't get this stage — running GRPO at 550-billion-parameter scale was outside their compute budget, plain and simple. SFT already does most of the work, so the marginal RL dollar was better spent elsewhere. 21 00:10:16,700 --> 00:10:30,539 [Hal Turing] Oh wait, hold on — before we move past that, I want to make sure I've got GenCorrect right, because this is the part I actually find the most clever. Two hundred candidates down to ten submissions — how does it decide which ten? 22 00:10:30,539 --> 00:11:14,610 [Dr. Ada Shannon] It's not picking the ten best by some score — it can't, it hasn't submitted anything yet. It generates up to two hundred candidates, then does diversity selection using token-shingle similarity: pick a center, then repeatedly pick whichever remaining candidate is farthest from every center already chosen, until you've got ten clusters. That's deliberately score-blind. Then it submits those ten, gets subtask-level feedback back, and accumulates the best score seen so far per subtask across rounds. The next round is conditioned on that accumulated feedback plus three reference solutions chosen to cover different gaps. Five rounds, ten submissions each, which isn't arbitrary — it exactly matches the official 50-submission-per-problem IOI limit. 23 00:11:14,610 --> 00:11:24,223 [Hal Turing] That's a genuinely elegant constraint match. So walk me through what all three stages actually add up to on IOI 2025 — give me the full arc for Nano. 24 00:11:24,223 --> 00:11:58,728 [Dr. Ada Shannon] Base Nano is 130 points. SFT alone takes it to 280, then RL nudges it to 291 — so SFT is doing the heavy lifting, RL is a real but modest add. Then five rounds of GenCorrect take it to 468, which clears the gold threshold of 438.3. Ultra follows a similar shape but skips RL entirely: SFT alone gets it to 304, and GenCorrect carries it to 502 — a bigger jump than Nano's, actually, which tells you Ultra makes better use of the extra sampling and feedback even without RL fine-tuning it toward that behavior. 25 00:11:58,728 --> 00:12:12,289 [Hal Turing] So GenCorrect is the single biggest lever in the whole pipeline, at least in points terms. Now, that's all IOI 2025, retrospective. What changed when they actually had to do this live at IOI 2026? 26 00:12:12,289 --> 00:13:06,113 [Dr. Ada Shannon] Three adaptations. First, they swapped the SFT teacher for Ultra from DeepSeek-V4-Flash to GLM-5.2, because GLM-5.2 hit a higher IOI 2025 score with noticeably shorter outputs — shorter generations mean more candidates fit in a fixed live inference window. Second, they blew up the final GenCorrect round from 200 to 1,000 candidates, using an execution-based selection method adapted from their own prior GenCluster work — generate test-input validators, run every candidate against a hundred generated inputs, then score and rank with a model-written scoring script. Third, they quantized Ultra to NVFP4 with MTP set to 5, which nearly quadruples throughput — 3.7x — for a cost of about 6.6 points of Score@1. And they ran it on up to 760 GB300 GPUs during the live window. 27 00:13:06,113 --> 00:13:16,794 [Hal Turing] Okay, so with all of that stacked together — GLM-5.2 teacher, the thousand-candidate final round, the quantization tradeoff — what actually happened on the day? 28 00:13:16,794 --> 00:14:24,085 [Dr. Ada Shannon] Under real competition constraints — no internet, one submission per minute, 50 submissions per problem — it beat gold by 174 points and the best human by 37. And here's the detail I actually find more interesting than the headline: after the competition, they reran the standard five-round GenCorrect pipeline — not the souped-up live version — independently five times, and got a mean of 521.72, ranging from 495.0 up to 545.8. Nobody tells us what that live run actually cost. Up to 760 GB300 GPUs at peak, and GenCorrect's final round alone generates a thousand candidates per problem across six problems — six thousand compiled, executed solutions just for the last round, on top of four earlier rounds of two hundred each. The paper's own limitations section says the live result should be read as a system-level comparison under matched time and submission limits, not an equal-resource one. That's an honest sentence, but it's carrying a lot of weight. Same submission limits as a human, sure. Same resources? Not close. 29 00:14:24,085 --> 00:14:50,231 [Hal Turing] Right, and 'system-level' quietly forecloses the harder question. So here's the number that actually made me pause: 535.4 live, but the five offline reruns of the general pipeline averaged 521.72, range 495 to 545.8. That's only 13.68 points above the mean. And the bottom of that range, 495, is still below the top human's 498.27. 30 00:14:50,231 --> 00:15:29,983 [Dr. Ada Shannon] Exactly — with a spread that wide from just five runs, you can't rule out a different seed on competition day landing below the human record. We don't get a standard deviation, only min and max, so there's no clean probability to hand you. But qualitatively, one live attempt sitting inside a fifty-point observed range isn't the same claim as 'this system reliably beats humans.' It's 'it beat them once, and offline testing says that's plausible, not guaranteed.' Reporting the distribution at all is more transparent than most of this field bothers with. The abstract still states it as settled fact rather than a probabilistic one, though. 31 00:15:29,983 --> 00:15:57,615 [Hal Turing] That variance question connects to something else — every competition-specific choice, the GLM-5.2 teacher, the thousand-candidate final round, the exact NVFP4 settings, was all tuned against IOI 2025. Then they get exactly one shot at IOI 2026, a genuinely novel problem set. How much of 535.4 is generalization versus just being well-fit to how IOI 2025 happened to be structured? 32 00:15:57,615 --> 00:16:38,668 [Dr. Ada Shannon] Honestly, we don't fully know. IOI problems share format conventions year to year, so some transfer is expected, but it's still a single dev-year driving every hyperparameter for a single live shot. Compare that to OpenAI's o1-ioi and o3 system, from 'Competitive Programming with Large Reasoning Models,' OpenAI et al., 2025 — the direct predecessor, RL plus a hand-engineered inference pipeline reaching gold under the same submission limits. This paper name-checks it but never really contrasts compute budgets or argues why GenCorrect's execution-grounded selection should generalize better than o3's approach. It's positioned as beaten, not seriously argued against. 33 00:16:38,668 --> 00:17:06,114 [Hal Turing] Oh wait, hold on — that's the thing that actually bugs me most about the human comparison generally. The top contestant is typing C++ alone in real time, no ability to fork off a thousand parallel attempts and execution-test them before submitting anything. GenCorrect's entire value is parallelizing cognitive labor a human structurally cannot. Same clock, same internet restriction — but that's not the same task. 34 00:17:06,114 --> 00:17:58,313 [Dr. Ada Shannon] It's not, and the paper doesn't sit with that. Worth putting next to their own prior work, too — GenCorrect's final-round execution-based selection is explicitly adapted from GenCluster, Samadi, Ficek, Narenthiran and colleagues, 2026, same author group, so a real chunk of this win is refinement of a system they already had gold with. On efficiency, Nemotron-Cascade 2, Yang and colleagues, 2026, hit gold-level performance with only 3B active parameters, versus Ultra-CC's 55B here — the paper never engages with what that eighteen-times parameter budget buys beyond peak score. And DeepSeek-V3.2-Speciale, DeepSeek-AI, 2025, claims gold on both IOI 2025 and ICPC 2025, but it's never benchmarked head-to-head in their own table. 35 00:17:58,313 --> 00:18:13,731 [Hal Turing] Same goes for that NVFP4 tradeoff — quantization is a dial, and where you set it changes which problems actually get solved, not just how fast. They pick one operating point and never show whether a gentler setting changes the final score. 36 00:18:13,731 --> 00:18:55,944 [Dr. Ada Shannon] The ordering that survives all this scrutiny is still solid, though — SFT does the heavy lifting, RL adds real but modest gains and only for the smaller model, GenCorrect is the biggest late-stage jump. Genuinely useful map for anyone building this kind of pipeline. What's missing is transparency on cost and reproducibility — nobody outside a hyperscale lab can rerun this live comparison to check it. Worth flagging Huang, Chen, Mishra and colleagues, 2024, who showed self-correction without execution grounding is unreliable — GenCorrect avoids that by grounding in real execution, but the model also writes its own tests in the final round, and that circularity risk never gets investigated. 37 00:18:55,944 --> 00:19:17,632 [Hal Turing] So: a real, gold-beating result, reported honestly enough that you can see the cracks if you look. SFT still wins the day, RL and test-time compute are real but secondary levers, and 'beats the best human' is true once, under conditions almost nobody else can reproduce. That's it for this one. Thanks for listening, and take care.