Verification: how solid is it?
We order these strongest to weakest. When we're unsure, the lower grade wins.
A machine-checked proof (Lean, Coq, Isabelle) or an exact symbolic or computational verification that anyone can rerun. The gold standard.
Multiple qualified people unaffiliated with the announcing lab have checked the result and confirmed it.
Published in a venue with real review. Weaker than a formal proof for mathematics, but stronger for empirical claims.
The human collaborator checked it, but no independent party has confirmed it yet.
Announced but not yet checked by anyone outside the lab. The default for a press release.
Substantive technical objections have been raised and remain unresolved.
The result turned out to already exist in the literature. These entries stay on the record; showing the failure mode is what makes the rest trustworthy.
Shown to be wrong.
Autonomy: how much did the AI do?
This axis is the whole point. Most breathless claims collapse here; we grade on the strictest defensible reading.
The AI produced the core idea and the result with no human mathematical input beyond posing the problem.
The AI produced the key insight; humans formalised, checked, or cleaned it up.
Genuine back-and-forth. Neither party would have gotten there alone.
A human drove the research; the AI accelerated search, algebra, or literature review.
The AI is a component inside a human-designed search loop (FunSearch, AlphaEvolve). The system found it, but the framing was human.
The AI surfaced an existing result humans had overlooked. Valuable, but not new mathematics.
Editorial rules
These are what separate the registry from a press-release aggregator.
- We run a novelty check before publishing, no exceptions. What was searched gets recorded even when it comes back clean.
- We never delete an entry. We downgrade and annotate it instead; the public history is the credibility mechanism, and quiet edits destroy it.
- We cap a claim with no reproducible artifact at
claimed. No exceptions for famous labs. - We start lab-announced results at
claimedregardless of how confident the announcement sounds. - We grade autonomy on the strictest defensible reading. If a human posed the problem, suggested the approach, and checked the algebra, that is not autonomous.