All topicsSTATE OF AI REPORT.

AI-assisted drug discovery is reaching Phase 3

AI-assisted drug discovery is finally reaching late-stage clinical development. Generate:Biomedicines' GB-0895 for severe asthma and Insilico's rentosertib for idiopathic pulmonary fibrosis are recruiting in Phase 3. Enveda’s ENV-294 is in placebo-controlled Phase 2a trials for eczema and asthma. Human trials are of course the ultimate test and will determine whether these medicines are safe and effective.

Questions in this section

Which AI-assisted medicines have reached Phase 3?Does reaching Phase 3 mean an AI-assisted medicine works?Are these trials a complete list of AI-assisted medicines?Can AI drug discovery companies build businesses without owning medicines?Does a predicted protein structure establish a useful medicine?Can scientific agents reliably verify their own work?Has AI solved the Navier-Stokes problem?
Clinical progress of medicines from AI-first drug discovery companies, including two programs in Phase 3.
Clinical progress of medicines from AI-first drug discovery companies, including two programs in Phase 3.

AI is solving problems across science

AI is solving mathematical problems that resisted professional researchers. Epoch AI's open-problem collection records nine solutions, four autonomous and five through human-AI collaboration. In biology, better predictions of molecular interactions are improving which designs reach laboratory testing. I’m seeing this first hand within my portfolio company, Profluent, which builds frontier models for designing proteins such as gene editors.

New businesses are forming around drug discovery

The good news for entrepreneurs in biotech is that new business models are emerging around drug discovery that aren’t about owning drugs. For example, Tempus licenses oncology data for foundation-model development with AstraZeneca and Pathos, while GSK's five-year Noetik agreement includes annual fees for access to virtual-cell models for cancer research.

Better selection means more useful laboratory experiments

Generating a plausible molecule is only the beginning. Researchers need to decide which designs deserve the cost of synthesis and testing. BoltzProt-1 generates protein binders and uses an interaction predictor, BoltzPPI, to rank them. Across ten novel targets, this produced 12 confirmed binders from 150 tested designs, compared with five from 150 when ranking by structure confidence alone.

The hit rate rose from 3.3% to 8.0%. Seven of the twelve confirmed binders passed all the reported developability tests, including stability and unwanted binding. Improving the selection step can increase the useful output of a fixed laboratory budget, while leaving substantial downstream optimization to be done.

Further reading: slide 76.

Binding prediction increases the laboratory yield of designed nanobodies.
Binding prediction increases the laboratory yield of designed nanobodies. Report slide 76.

An antibody needs more than the ability to bind

A candidate medicine must remain stable enough to manufacture and use. Chai’s detailed Chai-2 study evaluated 88 designed antibodies against 28 targets. Eighty-six percent had at most one developability flag, and 24 targets had at least one design with no flags. It also found binders for all six tested GPCRs, a difficult class of membrane receptors.

These experiments extend the evidence beyond a predicted molecular shape or a single binding result. They still do not establish clinical benefit. The report uses the detailed Chai-2 experiments rather than treating Chai-3’s announced improvement as a matched comparison, because the newer announcement does not establish a comparable denominator.

Further reading: slide 77.

A scientific agent needs checks on how it reaches an answer

Scientific prose can look convincing even when its results are fabricated. In one Co-Scientist study, thirty experts reviewed 150 manuscripts across fifty matched topics. Removing the system’s reliability modules increased hallucinations severe enough to invalidate the results from 4% to 46%.

Verification improved the result substantially, but serious problems remained: 24% of the full system’s papers had severe methodology failures and 16% had severe plagiarism. A separate crystal-growth experiment still required humans to set constraints, load samples, and validate the output. Automating more of the workflow increases the value of reliable checks at each stage.

Further reading: slide 70.

Reliability checks reduce fabricated scientific results, with other failures remaining.
Reliability checks reduce fabricated scientific results, with other failures remaining. Report slide 70.

Physical completion also needs independent verification

Mecka tested three frontier models on nine labware-handling tasks across 540 attempts, giving them camera views and robot-arm control without demonstrations or fine-tuning. Human reviewers found that 89 of 192 declarations of completion were incomplete: 79 were partial completions and ten were failures.

This is a test of handling laboratory equipment, not of choosing scientific hypotheses. It exposes a practical requirement for automated labs: the system that proposes or performs an experiment cannot be the only judge of whether it happened. The path from a model-generated idea to a scientific result runs through measured physical outcomes, experimental controls, and human scrutiny.

Further reading: slide 71.

AI-generated mathematics needs proof checking and expert assessment

The report describes Astra resolving three Erdős problems about patterns in networks, with proofs checked by the Lean proof assistant. Formal checking verifies a proof against a precisely stated mathematical claim. Expert assessment is still needed to establish what that claim resolves and how it relates to the original problem.

OpenAI also reported constructing a singularity in three-dimensional fluid flow under external forcing. The report explicitly distinguishes this result from the unforced Navier-Stokes case, which remains open. Describing the work simply as solving Navier-Stokes would erase that distinction.

On October 6, OpenAI released 722 manuscripts in 372 result families from an unreleased model, with partial Lean verification. Those counts describe released material, not 722 independently accepted discoveries. Formal proofs, proposed arguments, and results still undergoing expert assessment should be identified separately when evaluating AI’s mathematical contributions.

Further reading: slide 67.

Evidence you can use

Two recruiting Phase 3 programs

Registry checked 2026-10-07

Two recruiting Phase 3 programs
MedicineConditionRegistry status
GB-0895SourceSevere asthmaPhase 3 · Recruiting
Rentosertib (INS018_055)SourceIdiopathic pulmonary fibrosisPhase 3 · Recruiting

Selected programs from the report, checked against the ClinicalTrials.gov API. GB-0895’s record was updated on 2026-10-07 and rentosertib’s on 2026-09-17. Trial phase and recruitment status do not establish efficacy, regulatory approval, or how much of development was performed using AI.

Sources: ClinicalTrials.gov: GB-0895, NCT07276724 · ClinicalTrials.gov: rentosertib, NCT07687459.

Frequently asked questions

Answers drawn from the report and the sources below.

Does a predicted protein structure establish a useful medicine?

A structural prediction is one step. Binding, stability, unwanted interactions, manufacturability, and human safety and efficacy require further tests. The report separates computational designs, laboratory results, and clinical development.

Source: State of AI Report 2026, slide 76: Using a binding predictor more than doubles the yield of designed nanobodies · State of AI Report 2026, slide 77: Chai's designed antibodies pass laboratory tests beyond binding.

Sources and dates

2026 report snapshot. Preview revised 2026-10-07. Individual data periods and source checks are listed below. This is not a claim that every source was updated on that date.

  1. ClinicalTrials.gov: GB-0895, NCT07276724Registry updated 2026-10-07. Registry API checked 2026-10-07.
  2. ClinicalTrials.gov: rentosertib, NCT07687459Registry updated 2026-09-17. Registry API checked 2026-10-07.
  3. Noetik: GSK licenses OCTO2026. Retained from the launch essay.
  4. State of AI Report 2026, slide 67: OpenAI graduates from Erdős problems to a $1M Millennium Prize problem2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  5. State of AI Report 2026, slide 70: Verification cuts fabricated results, while human scientific oversight remains essential2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  6. State of AI Report 2026, slide 71: Nearly half of frontier models’ “done” claims in lab-handling tasks were incomplete2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  7. State of AI Report 2026, slide 76: Using a binding predictor more than doubles the yield of designed nanobodies2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  8. State of AI Report 2026, slide 77: Chai's designed antibodies pass laboratory tests beyond binding2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  9. State of AI Report 20262026 report snapshot. Report PDF. See individual slide references for the expanded analysis.
  10. open-problem collection - epoch.aiOriginal source link retained from the launch essay.
  11. laboratory testing - boltz.bioOriginal source link retained from the launch essay.
  12. oncology data - investors.tempus.comOriginal source link retained from the launch essay.

Cite this page

Benaich, Nathan. “AI-assisted drug discovery is reaching Phase 3.” State of AI Report 2026. Published 2026-10-08.