A fully machine-generated paper passed workshop peer review
A fully machine-generated paper passed workshop peer review is graded author verified on whataifound.org, with the AI's role graded ai-led.
One of three end-to-end AI-generated manuscripts submitted to an ICLR 2025 workshop averaged 6.33 from reviewers and would have been accepted; the authors withdrew it, and the bar cleared was a workshop track, not the main conference.
- Verification
- Author verified
- Autonomy
- AI-led
- Lab
- Sakana AI
- Model
- The AI Scientist-v2
- Field
- Computer science
- Date
- 2025-03-12
What was found
The AI Scientist-v2 generates a hypothesis, writes and runs the experiments through agentic tree search, and produces the manuscript, figures and related work without human editing. Sakana submitted three such papers to the ICLR 2025 workshop "I Can't Believe It's Not Better", with the knowledge and cooperation of the workshop organisers and ICLR leadership, and agreed in advance to withdraw anything accepted. One paper — on whether an explicit compositional regularisation term improves compositional generalisation — scored 6.33 and cleared the bar.
Novelty check
Machine-written text had passed review before in the form of nonsense-generator stings (SCIgen, 2005), which tested review rather than produced research. The claim here is different: an autonomous pipeline produced a real experimental paper that reviewers rated acceptable. Sakana's own write-up is the source for the process, and the code and submitted manuscripts are public.
Caveats
Workshop tracks accept a far higher share of submissions than the main conference, and Sakana says so itself; the company also notes the accepted paper contains a citation error and that its own reviewers judged the work below the main-conference bar. The paper's scientific content is a negative result: the proposed regularisation did not help. Reviewers were not told which submissions were machine-generated, but the experiment ran with organiser consent, so this is not a blind test of peer review at scale. Humans chose the venue and selected which generated papers to submit, which caps autonomy at ai-led.
Sources
How this is graded
whataifound.org grades every entry on two axes: verification (how solid the result is, from a machine-checked proof down to refuted) and autonomy (how much the AI did versus its human collaborators). This finding is author verified and ai-led. Full definitions are in the methodology.
Cite this entry
whataifound.org (2025). A fully machine-generated paper passed workshop peer review. whataifound.org: A Registry of AI Scientific and Mathematical Discoveries. https://whataifound.org/finding/2025-03-12-ai-scientist-workshop-paper