IMO-Bench splits mathematical competence into distinct tasks. This diagram shows why answer-only evaluation can hide weak reasoning: the benchmark forces exposure of intermediate competence, not just the endpoint.
Same final answer, very different evidence. Toggle to compare what each evaluation regime “sees.”
Mocked from the paper’s qualitative description: combinatorics and number theory skew harder; algebra and geometry contain more easy/medium answer items. Hover cells for counts and benchmark pressure.
These tasks are hard because they require discovering the right move, not just executing a known template. The radar-style profile contrasts benchmark families.
Mock bars illustrate the paper’s headline pattern: modest lead on answer solving, much larger separation on advanced proof generation.
Models can be strong in one axis and weak in another. This matrix visualizes why “one scalar leaderboard” is a poor summary of mathematical competence.
The benchmark needs scalable grading, but proof judges can be style-sensitive. Toggle between optimistic alignment and adversarial stress cases.
Benchmark trust depends on duplicate audits, development-overlap disclosure, blinded grading, and style-bias checks. Click steps to inspect each control point.
Compact source list for the visuals and benchmark context.