Same contest, same clock, same 600-point scale. Ultra-CC clears gold by 174.28 points and the best human on-site by 37.13.
Gold cutoff is set from that year's human score distribution (~top 1/12), not a fixed number — 2025 and 2026 are not directly comparable.
Each of 6 problems splits into subtasks by input size. Hover a cell — this is what separates "no idea" from "right idea, wrong complexity" from "fully correct."
Curate → distill teacher traces → supervised fine-tune → (optionally) reinforce → spend extra compute at inference. Toggle to see which stages each model actually gets.
SFT does the heavy lifting for both models. RL adds a modest gain — but only exists for Nano. GenCorrect is the single biggest jump either way.
Ultra starts from a stronger RLVR-teacher checkpoint, so it needs far fewer traces and epochs.
4,000 candidate problems trimmed to 3,219 with reliable, fast executable environments.
Score-blind diversity selection picks 10 of 200 candidates per round; feedback accumulates across 5 rounds — exactly matching IOI's 50-submission limit. Step through it.
Three changes shipped for the live 2026 run: a shorter-output teacher, a 5x bigger final round, and NVFP4 quantization for throughput.
Instead of firing every parameter for every token, MoE routes each token to a small subset of experts. Toggle to compare a dense pass with a sparse (MoE) pass.
Ultra-CC pays 18x the active-parameter budget of the most efficient cited alternative for extra peak score — a tradeoff the paper never quantifies.
Five offline reruns of the standard GenCorrect pipeline: mean 521.72, range 495.0–545.8. The live score (535.4) sits inside that range — the bottom of the range sits below the human record (498.27).
| 1 | Post-Training Language Models for Gold-Medal Performance in Coding Competitions — Ficek, Narenthiran, Samadi, Majumdar, Ginsburg (NVIDIA), 2026 | arXiv:2609.02849 |
| 2 | Competitive programming with large reasoning models (o1-ioi, o3) — OpenAI et al., 2025 | Scholar |
| 3 | Scaling test-time compute to achieve IOI gold medal with open-weight models (GenCluster) — Samadi, Ficek, Narenthiran et al., 2026 | Scholar |
| 4 | Nemotron-cascade 2: Post-training LLMs with cascade RL and multi-domain on-policy distillation — Yang et al., 2026 | Scholar |
| 5 | Large Language Models Cannot Self-Correct Reasoning Yet — Huang, Chen, Mishra et al., 2024 | Scholar |
| 6 | DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models (DeepSeek-V3.2-Speciale) — DeepSeek-AI, 2025 | Scholar |