NVIDIA Ultra-CC Beats Top Human at IOI 2026

Post-Training Language Models for Gold-Medal Performance in Coding Competitions · Ficek, Narenthiran, Samadi, Majumdar, Ginsburg — NVIDIA, 2026
arXiv:2609.02849 IOI 2026 Nemotron-3 Nano-CC / Ultra-CC GRPO · GenCorrect
361.12
2026 Gold Cutoff
498.27
Top Human, IOI 2026
535.4
Ultra-CC, Live, IOI 2026

Score Comparison — IOI 2026

Same contest, same clock, same 600-point scale. Ultra-CC clears gold by 174.28 points and the best human on-site by 37.13.

Gold cutoff Top human Ultra-CC (live)

The Gold Bar Moves Every Year

Gold cutoff is set from that year's human score distribution (~top 1/12), not a fixed number — 2025 and 2026 are not directly comparable.

Partial Credit: Why IOI Is a Harder Benchmark

Each of 6 problems splits into subtasks by input size. Hover a cell — this is what separates "no idea" from "right idea, wrong complexity" from "fully correct."

No credit Partial (wrong complexity) Full credit

Four-Stage Pipeline

Curate → distill teacher traces → supervised fine-tune → (optionally) reinforce → spend extra compute at inference. Toggle to see which stages each model actually gets.

Score Progression on IOI 2025 (retrospective)

SFT does the heavy lifting for both models. RL adds a modest gain — but only exists for Nano. GenCorrect is the single biggest jump either way.

Nano-CC Ultra-CC Gold cutoff (438.3)

Training Data Asymmetry

Ultra starts from a stronger RLVR-teacher checkpoint, so it needs far fewer traces and epochs.

GRPO Problem Funnel (Nano only)

4,000 candidate problems trimmed to 3,219 with reliable, fast executable environments.

GenCorrect, Round by Round

Score-blind diversity selection picks 10 of 200 candidates per round; feedback accumulates across 5 rounds — exactly matching IOI's 50-submission limit. Step through it.

Standard vs. Live-Day Configuration

Three changes shipped for the live 2026 run: a shorter-output teacher, a 5x bigger final round, and NVFP4 quantization for throughput.

Mixture-of-Experts: Only Some Specialists Show Up

Instead of firing every parameter for every token, MoE routes each token to a small subset of experts. Toggle to compare a dense pass with a sparse (MoE) pass.

Total vs. Active Parameters

Ultra-CC pays 18x the active-parameter budget of the most efficient cited alternative for extra peak score — a tradeoff the paper never quantifies.

Total params Active params / token

One Live Run Inside a Wide Offline Range

Five offline reruns of the standard GenCorrect pipeline: mean 521.72, range 495.0–545.8. The live score (535.4) sits inside that range — the bottom of the range sits below the human record (498.27).

Offline reruns (n=5) Live IOI 2026 result Top human (498.27)
With only 5 samples and no reported standard deviation, a different random seed on competition day landing below the human record can't be ruled out.

What the Paper Argues vs. What It Skips

References

1Post-Training Language Models for Gold-Medal Performance in Coding Competitions — Ficek, Narenthiran, Samadi, Majumdar, Ginsburg (NVIDIA), 2026arXiv:2609.02849
2Competitive programming with large reasoning models (o1-ioi, o3) — OpenAI et al., 2025Scholar
3Scaling test-time compute to achieve IOI gold medal with open-weight models (GenCluster) — Samadi, Ficek, Narenthiran et al., 2026Scholar
4Nemotron-cascade 2: Post-training LLMs with cascade RL and multi-domain on-policy distillation — Yang et al., 2026Scholar
5Large Language Models Cannot Self-Correct Reasoning Yet — Huang, Chen, Mishra et al., 2024Scholar
6DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models (DeepSeek-V3.2-Speciale) — DeepSeek-AI, 2025Scholar