whataifound.org ← Back to the registry

Methodology

We grade every finding as one record on two axes: how solid the evidence is and how much the AI actually did. We set grades conservatively, and only ever downgrade them, never quietly remove one. The schema is the product; everything else renders it.

Verification: how solid is it?

We order these strongest to weakest. When we're unsure, the lower grade wins.

Formally verified

A machine-checked proof (Lean, Coq, Isabelle) or an exact symbolic or computational verification that anyone can rerun. The gold standard.

Independently checked

Multiple qualified people unaffiliated with the announcing lab have checked the result and confirmed it.

Peer reviewed

Published in a venue with real review. Weaker than a formal proof for mathematics, but stronger for empirical claims.

Author verified

The human collaborator checked it, but no independent party has confirmed it yet.

Claimed

Announced but not yet checked by anyone outside the lab. The default for a press release.

Disputed

Substantive technical objections have been raised and remain unresolved.

Already known

The result turned out to already exist in the literature. These entries stay on the record; showing the failure mode is what makes the rest trustworthy.

Refuted

Shown to be wrong.

Autonomy: how much did the AI do?

This axis is the whole point. Most breathless claims collapse here; we grade on the strictest defensible reading.

Autonomous

The AI produced the core idea and the result with no human mathematical input beyond posing the problem.

AI-led

The AI produced the key insight; humans formalised, checked, or cleaned it up.

Collaborative

Genuine back-and-forth. Neither party would have gotten there alone.

AI-assisted

A human drove the research; the AI accelerated search, algebra, or literature review.

Search scaffold

The AI is a component inside a human-designed search loop (FunSearch, AlphaEvolve). The system found it, but the framing was human.

Retrieval

The AI surfaced an existing result humans had overlooked. Valuable, but not new mathematics.

Editorial rules

These are what separate the registry from a press-release aggregator.

  1. We run a novelty check before publishing, no exceptions. What was searched gets recorded even when it comes back clean.
  2. We never delete an entry. We downgrade and annotate it instead; the public history is the credibility mechanism, and quiet edits destroy it.
  3. We cap a claim with no reproducible artifact at claimed. No exceptions for famous labs.
  4. We start lab-announced results at claimed regardless of how confident the announcement sounds.
  5. We grade autonomy on the strictest defensible reading. If a human posed the problem, suggested the approach, and checked the algebra, that is not autonomous.