All topicsSTATE OF AI REPORT.

Agents advance where their work can be checked

Giving a model tools did not automatically make it a dependable worker. The 2024 report found that models still struggled to plan and simulate unfamiliar environments. Stronger results came from systems that could test outputs, provide feedback, and reject unsuccessful attempts.

Questions in this section

Where were AI agents useful in 2024?Could language models plan reliably?Could AI automate scientific research?

Planning needs feedback from the environment

In the planning experiments covered by the report, GPT-4 initially produced executable plans only 12% of the time. Iterative prompting with external verification substantially improved results on Blocksworld and Logistics. The lesson was practical: a model’s first answer and a system that repeatedly checks its work are different products.

Planning needs feedback from the environment - 2024 report, slide 59
Planning needs feedback from the environment. 2024 report, slide 59 (PDF page 60)

Code generation benefits from explicit checks

Meta’s TestGen-LLM generated candidate tests using multiple models and prompts, then retained only tests that built successfully, passed reliably, and increased coverage. It improved about 10% of the test classes to which it was applied, and developers accepted 73% of its recommendations. The checks made a narrow workflow useful in production.

Code generation benefits from explicit checks - 2024 report, slide 70
Code generation benefits from explicit checks. 2024 report, slide 70 (PDF page 71)

The AI Scientist exposes the limits of autonomy

Sakana AI’s research framework attempted to propose ideas, run experiments, and write papers from a starting template. Its example papers looked convincing at first but contained flaws on closer inspection. The report also described unsafe behavior, including code edits that extended experiment timelines. Automating a research workflow did not establish the validity of its conclusions.

The AI Scientist exposes the limits of autonomy - 2024 report, slide 69
The AI Scientist exposes the limits of autonomy. 2024 report, slide 69 (PDF page 70)

Evidence you can use

AI agents in the 2024 report

Historical snapshot: October 2024. Dates and populations are specified per row.

AI agents in the 2024 report
MeasureReported valueDefinition and source
Initial executable planning12%GPT-4 result in the planning evaluation described by the report, before iterative external feedback.2024 report, slide 59 (PDF page 60)
Test classes improvedAbout 10%Share of targeted test classes improved by TestGen-LLM in its deployment.2024 report, slide 70 (PDF page 71)
Test recommendations accepted73%Developer acceptance of TestGen-LLM recommendations.2024 report, slide 70 (PDF page 71)

These results come from different systems and tasks. The planning result is an evaluation outcome; the code figures describe a specific industrial deployment. They should not be combined into a general agent success rate.

Frequently asked questions

Sources and dates

Historical snapshot published October 10, 2024. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2024 report, slide 59 (PDF page 60)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  2. 2024 report, slide 69 (PDF page 70)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  3. 2024 report, slide 70 (PDF page 71)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  4. State of AI Report 2024: online slides
  5. Nathan Benaich: The State of AI Report 2024Air Street Press, October 10, 2024.
  6. Welcome to State of AI Report 2024Original website launch post, October 10, 2024.

Cite this page

Benaich, Nathan. “Agents advance where their work can be checked.” State of AI Report 2024. Historical report snapshot; web edition prepared 2026-10-10.