
How to read this finding
Model size, training tokens, and dataset size are distinct quantities. These values describe the reported research system; they do not establish clinical efficacy or imply that every generated protein works in an experiment.
The report describes systems whose hypotheses or designs reached experimental validation, including Co-Scientist and Stanford’s Virtual Lab. Human teams and laboratories remained involved.

Model size, training tokens, and dataset size are distinct quantities. These values describe the reported research system; they do not establish clinical efficacy or imply that every generated protein works in an experiment.
Evidence you can use
2025 report snapshot
| Measure | Reported value | Definition and source |
|---|---|---|
| Largest model | 46B parameters | Mixture-of-experts protein language model2025 report, slide 63 (PDF page 64) |
| Training tokens | 1.5 trillion | Total tokens used in training2025 report, slide 63 (PDF page 64) |
| PPA-1 dataset | 3.4 billion full-length proteins | Reported dataset composition2025 report, slide 63 (PDF page 64) |
Model size, training tokens, and dataset size are distinct quantities. These values describe the reported research system; they do not establish clinical efficacy or imply that every generated protein works in an experiment.
Historical snapshot published October 9, 2025. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan. “Could AI generate useful scientific hypotheses in 2025?.” State of AI Report 2025. Historical report snapshot; web edition prepared 2026-10-10.