IMO-Bench for Robust Mathematical Reasoning

A visual companion focused on what the benchmark measures: not just final answers, but the full triangle of solving, proof writing, and proof grading. The page emphasizes benchmark structure, task difficulty, grading risk, and where Olympiad-style reasoning exposes brittle pattern matching.
arXiv: 2511.01846
2025 • Luong et al.
400 answer problems
60 proof problems
1,000 human gradings
Domains: algebra • combinatorics • geometry • number theory
80.0%
Top reported answer score
65.7%
Top reported advanced proof score
+42.4
Proof gap vs best non-Gemini
3-part
Benchmark decomposition
Benchmark Pressure Map Open paper ↗
1. Benchmark Anatomy
2. Difficulty & Domains
3. Results & Separation
4. Proof Grading Risks

Three-task evaluation pipeline

IMO-Bench splits mathematical competence into distinct tasks. This diagram shows why answer-only evaluation can hide weak reasoning: the benchmark forces exposure of intermediate competence, not just the endpoint.

solving proof generation grading / verification brittleness / Goodhart risk

Answer-only vs proof-aware measurement

Same final answer, very different evidence. Toggle to compare what each evaluation regime “sees.”

Domain × difficulty heatmap

Mocked from the paper’s qualitative description: combinatorics and number theory skew harder; algebra and geometry contain more easy/medium answer items. Hover cells for counts and benchmark pressure.

What Olympiad problems demand

These tasks are hard because they require discovering the right move, not just executing a known template. The radar-style profile contrasts benchmark families.

Reported benchmark separation

Mock bars illustrate the paper’s headline pattern: modest lead on answer solving, much larger separation on advanced proof generation.

Capability decomposition matrix

Models can be strong in one axis and weak in another. This matrix visualizes why “one scalar leaderboard” is a poor summary of mathematical competence.

Autograder alignment vs failure modes

The benchmark needs scalable grading, but proof judges can be style-sensitive. Toggle between optimistic alignment and adversarial stress cases.

Governance checklist

Benchmark trust depends on duplicate audits, development-overlap disclosure, blinded grading, and style-bias checks. Click steps to inspect each control point.

Selected references

Compact source list for the visuals and benchmark context.

Towards Robust Mathematical Reasoning (Luong et al., 2025) — arXiv:2511.01846
Training Verifiers to Solve Math Word Problems (Cobbe et al., 2021) — verifier framing for answer-centric math evals
MATH Dataset (Hendrycks et al., 2021) — competition-style answer benchmark
FrontierMath (Glazer et al., 2024) — advanced math benchmark with strong final-answer emphasis
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al., 2023) — judge reliability caveats
G-Eval (Gao et al., 2023) — model-based evaluation alignment
Automatic Evaluation of Mathematical Proofs in Natural Language — survey context for proof grading
LeanDojo (Yang et al., 2023) — formal theorem proving contrast