Gold-medal standard at the 2025 International Mathematical Olympiad
Gold-medal standard at the 2025 International Mathematical Olympiad is graded independently checked on whataifound.org, with the AI's role graded ai-led.
An advanced version of Gemini Deep Think read the official problem statements and wrote natural-language proofs for five of the six 2025 IMO problems inside the contest time limits, scoring 35 of 42 — graded and certified by the official IMO coordinators.
- Verification
- Independently checked
- Autonomy
- AI-led
- Lab
- Google DeepMind
- Model
- Gemini Deep Think (advanced version)
- Field
- Mathematics
- Date
- 2025-07-21
What was found
The model worked end to end in natural language within the two 4.5-hour sessions, with no human translating problems into a formal language, and solved problems one through five. Google DeepMind was in the first cohort whose results were officially graded and certified by IMO coordinators against the rubric used for student scripts; 35 points is the gold-medal threshold. OpenAI announced the same 35/42 score two days earlier with an experimental reasoning model, graded internally by three former IMO medallists rather than by the IMO.
Novelty check
IMO problems have published official solutions, so this is a capability milestone rather than a new mathematical result, and the registry records it as one. The prior best was AlphaProof and AlphaGeometry 2's 28/42 in 2024, which required humans to hand-translate each problem into Lean first; removing that step is the substantive change.
Caveats
Benchmark performance on problems with known answers, not a discovery. Google's own account describes training on a curated corpus of high-quality solutions and providing general guidance on how to approach IMO problems, which is human mathematical input beyond posing the problem — hence ai-led rather than autonomous. Problem six went unsolved. The parallel OpenAI claim is self-graded and not IMO-certified, and drew criticism for being published before the closing ceremony, which organisers had asked AI developers to wait for.
Independent checks
Official IMO coordinators (certified grading): 35/42, gold-medal standard · link ↗
Sources
How this is graded
whataifound.org grades every entry on two axes: verification (how solid the result is, from a machine-checked proof down to refuted) and autonomy (how much the AI did versus its human collaborators). This finding is independently checked and ai-led. Full definitions are in the methodology.
Cite this entry
whataifound.org (2025). Gold-medal standard at the 2025 International Mathematical Olympiad. whataifound.org: A Registry of AI Scientific and Mathematical Discoveries. https://whataifound.org/finding/2025-07-21-gemini-deepthink-imo