← BlogAI Research

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Isaac Kargar3 min read

  • AI
  • LLM
  • Scientific Discovery
  • Automation
The AI Scientist coordinates idea generation, experiments, writing, and simulated review
The framework connects ideation, code changes, experiment execution, paper writing, and an automated review. The AI Scientist paper describes the components and their limits.

What the framework does

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, by Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha, starts from a small research template. It generates ideas, writes and runs experiment code, records results, produces figures, drafts a LaTeX paper, and applies a simulated review. The paper demonstrates this workflow in diffusion modeling, Transformer-based language modeling, and learning dynamics.

The word “automated” describes the pipeline in the paper. A generated manuscript still needs scientific checking, and a simulated review is not peer review or evidence that a result is correct.

Three phases

1. Idea generation

The system proposes research directions with an experiment plan and self-assessed novelty, feasibility, and interest scores. It queries Semantic Scholar and web tools to identify ideas that appear too similar to existing work. Semantic Scholar supplies literature search; it does not validate a proposed idea or its claims.

2. Experiment iteration

The framework uses Aider as a coding assistant to edit the experiment template, run the proposed experiments, and respond to errors or timeouts. The paper’s workflow can retry a failed experiment up to four times and repeat the experiment-planning loop up to five times. Aider supplies code editing and command execution; The AI Scientist supplies the surrounding research prompts and records.

The framework can change plotting code and add metrics. That flexibility can also produce unexpected outcomes, so generated figures and numerical results need to be checked against the executed code.

3. Paper write-up

The system passes experiment notes and plots to Aider to fill a paper template section by section. It queries Semantic Scholar for related-work references, refines the draft, and compiles the LaTeX document. A successful compilation checks document syntax. It does not establish that the citations are valid, the analysis is sound, or the discovery is scientifically important.

Simulated review

The review component uses a GPT-4o-based agent with NeurIPS-style criteria. It reads the generated PDF, returns scores and strengths or weaknesses, and can threshold the score into an accept or reject decision.

The AI Scientist paper evaluates this reviewer on 500 ICLR 2022 papers from OpenReview. In the paper’s balanced-data comparison, the calibrated GPT-4o (1-shot) reviewer with a score-6 threshold reports 0.65 ± 0.04 balanced accuracy, 0.66 ± 0.04 accuracy, 0.57 ± 0.05 F1, and 0.65 ± 0.04 AUC. The table’s human comparison values, including 0.66 balanced accuracy, 0.49 F1, and 0.65 AUC, come from a separate NeurIPS consistency experiment rather than measurements on those same ICLR papers. The reviewer has a higher false-positive rate than the human baseline, and the ICLR dataset is class-imbalanced. These are measurements of agreement with historical review decisions, not evidence that generated papers are ready to publish.

The paper also reports that each idea can be implemented and developed into a full paper for less than $15 in its experimental setup. That is a reported cost for the selected models, templates, and runs, rather than a general price for a valid research paper.

What the result means

The paper demonstrates an end-to-end automation pattern for small machine-learning experiments. It does not guarantee valid citations, successful code repairs, correct conclusions, or reliable scientific discovery. Human researchers still need to inspect the source data, implementation, statistical analysis, and references before treating an output as research.

Resources:

Work with Nazmi

Build your AI system with Nazmi.

Tell us what you are building, what exists today, and where your team needs help.

Start a conversation or book a 20-minute call →