Independently checked Search scaffold

AlphaEvolve across 67 problems: 20 improvements, 8 regressions

A systematic study with Terence Tao of an evolutionary coding agent on 67 open problems in analysis, combinatorics and geometry, reporting where it beat the literature and where it did not.

Model
AlphaEvolve (Gemini-based)
Field
Mathematics
Date
2025-11-03
Human collaborators
Terence Tao, Javier Gómez-Serrano, Bogdan Georgiev, Adam Zsolt Wagner

What was found

Rather than a single headline result, this is the systematic version: 67 problems attempted, with outcomes recorded either way. Tao reports roughly 20 cases of a new result, most others matching known bounds, and 8 where the tool did worse than the literature. Genuine advances include new three-dimensional constructions for finite-field Nikodym sets, which then prompted a human-refined hybrid construction beating both algebraic and random baselines, and an improved upper bound for Sidon sets, from 1.96365 to about 1.9526. On the well-known conjectures of Sidorenko, Sendov and Crouzeix it found no counterexample, which Tao notes may simply reflect that they are true.

Novelty check

Prior bounds are cited per problem in the 81-page paper, with data and prompts released publicly. Tao flags the central novelty risk himself: for well-known problems the tool often proposed the optimal construction immediately, suggesting recall from training data rather than search, and one supposedly new four-dimensional Kakeya construction turned out to essentially match a prior Bukh–Chao paper. Entries in this registry drawn from the earlier AlphaEvolve release (kissing number, minimum overlap, matrix multiplication) are tracked separately.

Caveats and known objections

Tao's own framing is that the contribution is scale and adaptability, not a fundamental breakthrough, and he describes the LLM as supplying 'educated randomness' inside an evolutionary loop rather than doing symbolic reasoning. Substantial human expertise was required to design non-exploitable verifiers: the system reliably finds loopholes, such as placing points at nearly identical locations to satisfy a distance tolerance. Included here because the negative and null results are reported alongside the wins, which is rare and is what makes it citable.

Independent checks

Terence Tao (co-author, published assessment): reports ~20 new results, 8 worse than literature; cautions on training-data contamination · link ↗

Disagree with these grades?

Bring a citation: a grade moves on evidence, not on argument.

Or on GitHub: submit a check challenge the grade send a correction or send a pull request

Entry history (1 event)
  1. AddedEntered the registry graded Independently checked and Search scaffold.

Entries are never deleted. A grade that does not hold up is downgraded on the record, with the reason beside it.

Graded independently checked for verification and search scaffold for autonomy. What these mean.

Cite this entry

Plain text
whataifound.org. (2025). AlphaEvolve across 67 problems: 20 improvements, 8 regressions. whataifound.org: A Registry of AI Scientific and Mathematical Discoveries. https://whataifound.org/finding/2025-11-03-alphaevolve-at-scale
BibTeX
@misc{whataifound-googledeepmind-2025-scale,
  title        = {AlphaEvolve across 67 problems: 20 improvements, 8 regressions},
  author       = {{whataifound.org}},
  year         = {2025},
  howpublished = {whataifound.org: A Registry of AI Scientific and Mathematical Discoveries},
  note         = {Result by Google DeepMind. Verification: Independently checked. Autonomy: Search scaffold.},
  url          = {https://whataifound.org/finding/2025-11-03-alphaevolve-at-scale}
}

Related findings

← All mathematics findings in the registry