From a question to a testable hypothesis
DeepMind’s Co-Scientist organized agents around hypothesis generation and experimental planning. The report described drug candidates and biological targets tested in vitro, along with a blind test concerning bacteriophage transfer. Stanford’s Virtual Lab used a principal-investigator agent and specialists to design nanobodies, including experimentally confirmed binders. Human researchers and laboratories remained part of these workflows.

Protein models gain a scaling framework
Profluent’s ProGen3 studied how protein language-model performance changed with compute and scale. Its largest model was a 46B-parameter mixture of experts trained on 1.5 trillion tokens. Larger models generated viable proteins across a broader sequence space, and alignment helped more at larger scale. That supplied a framework for choosing training investments in biological design.

Novelty is the difficult test
The report also examined where molecular predictions failed. On a benchmark of 2,600 protein-ligand pairs, AlphaFold 3 and reproductions performed better when pockets and poses resembled familiar cases. Novel chemistry remained harder. A plausible-looking structure or a good average score was not enough: useful discovery required tests of unfamiliar cases and physical validity.
