Independently checked AI-led

Gold-medal standard at the 2025 International Mathematical Olympiad

An advanced version of Gemini Deep Think read the official problem statements and wrote natural-language proofs for five of the six 2025 IMO problems inside the contest time limits, scoring 35 of 42, graded and certified by the official IMO coordinators.

Model
Gemini Deep Think (advanced version)
Field
Mathematics
Date
2025-07-21

What was found

The model worked end to end in natural language within the two 4.5-hour sessions, with no human translating problems into a formal language, and solved problems one through five. Google DeepMind was in the first cohort whose results were officially graded and certified by IMO coordinators against the rubric used for student scripts; 35 points is the gold-medal threshold. OpenAI announced the same 35/42 score two days earlier with an experimental reasoning model, graded internally by three former IMO medallists rather than by the IMO.

Novelty check

IMO problems have published official solutions, so this is a capability milestone rather than a new mathematical result, and the registry records it as one. The prior best was AlphaProof and AlphaGeometry 2's 28/42 in 2024, which required humans to hand-translate each problem into Lean first; removing that step is the substantive change.

Caveats and known objections

Benchmark performance on problems with known answers, not a discovery. Google's own account describes training on a curated corpus of high-quality solutions and providing general guidance on how to approach IMO problems, which is human mathematical input beyond posing the problem, hence ai-led rather than autonomous. Problem six went unsolved. The parallel OpenAI claim is self-graded and not IMO-certified, and drew criticism for being published before the closing ceremony, which organisers had asked AI developers to wait for.

Independent checks

Official IMO coordinators (certified grading): 35/42, gold-medal standard · link ↗

Disagree with these grades?

Bring a citation: a grade moves on evidence, not on argument.

Or on GitHub: submit a check challenge the grade send a correction or send a pull request

Entry history (1 event)
  1. AddedEntered the registry graded Independently checked and AI-led.

Entries are never deleted. A grade that does not hold up is downgraded on the record, with the reason beside it.

Graded independently checked for verification and ai-led for autonomy. What these mean.

Cite this entry

Plain text
whataifound.org. (2025). Gold-medal standard at the 2025 International Mathematical Olympiad. whataifound.org: A Registry of AI Scientific and Mathematical Discoveries. https://whataifound.org/finding/2025-07-21-gemini-deepthink-imo
BibTeX
@misc{whataifound-googledeepmind-2025-imo,
  title        = {Gold-medal standard at the 2025 International Mathematical Olympiad},
  author       = {{whataifound.org}},
  year         = {2025},
  howpublished = {whataifound.org: A Registry of AI Scientific and Mathematical Discoveries},
  note         = {Result by Google DeepMind. Verification: Independently checked. Autonomy: AI-led.},
  url          = {https://whataifound.org/finding/2025-07-21-gemini-deepthink-imo}
}

Related findings

← All mathematics findings in the registry