
AI is solving problems across science
AI is solving mathematical problems that resisted professional researchers. Epoch AI's open-problem collection records nine solutions, four autonomous and five through human-AI collaboration. In biology, better predictions of molecular interactions are improving which designs reach laboratory testing. I’m seeing this first hand within my portfolio company, Profluent, which builds frontier models for designing proteins such as gene editors.
New businesses are forming around drug discovery
The good news for entrepreneurs in biotech is that new business models are emerging around drug discovery that aren’t about owning drugs. For example, Tempus licenses oncology data for foundation-model development with AstraZeneca and Pathos, while GSK's five-year Noetik agreement includes annual fees for access to virtual-cell models for cancer research.
Better selection means more useful laboratory experiments
Generating a plausible molecule is only the beginning. Researchers need to decide which designs deserve the cost of synthesis and testing. BoltzProt-1 generates protein binders and uses an interaction predictor, BoltzPPI, to rank them. Across ten novel targets, this produced 12 confirmed binders from 150 tested designs, compared with five from 150 when ranking by structure confidence alone.
The hit rate rose from 3.3% to 8.0%. Seven of the twelve confirmed binders passed all the reported developability tests, including stability and unwanted binding. Improving the selection step can increase the useful output of a fixed laboratory budget, while leaving substantial downstream optimization to be done.
Further reading: slide 76.

An antibody needs more than the ability to bind
A candidate medicine must remain stable enough to manufacture and use. Chai’s detailed Chai-2 study evaluated 88 designed antibodies against 28 targets. Eighty-six percent had at most one developability flag, and 24 targets had at least one design with no flags. It also found binders for all six tested GPCRs, a difficult class of membrane receptors.
These experiments extend the evidence beyond a predicted molecular shape or a single binding result. They still do not establish clinical benefit. The report uses the detailed Chai-2 experiments rather than treating Chai-3’s announced improvement as a matched comparison, because the newer announcement does not establish a comparable denominator.
Further reading: slide 77.
A scientific agent needs checks on how it reaches an answer
Scientific prose can look convincing even when its results are fabricated. In one Co-Scientist study, thirty experts reviewed 150 manuscripts across fifty matched topics. Removing the system’s reliability modules increased hallucinations severe enough to invalidate the results from 4% to 46%.
Verification improved the result substantially, but serious problems remained: 24% of the full system’s papers had severe methodology failures and 16% had severe plagiarism. A separate crystal-growth experiment still required humans to set constraints, load samples, and validate the output. Automating more of the workflow increases the value of reliable checks at each stage.
Further reading: slide 70.

Physical completion also needs independent verification
Mecka tested three frontier models on nine labware-handling tasks across 540 attempts, giving them camera views and robot-arm control without demonstrations or fine-tuning. Human reviewers found that 89 of 192 declarations of completion were incomplete: 79 were partial completions and ten were failures.
This is a test of handling laboratory equipment, not of choosing scientific hypotheses. It exposes a practical requirement for automated labs: the system that proposes or performs an experiment cannot be the only judge of whether it happened. The path from a model-generated idea to a scientific result runs through measured physical outcomes, experimental controls, and human scrutiny.
Further reading: slide 71.
AI-generated mathematics needs proof checking and expert assessment
The report describes Astra resolving three Erdős problems about patterns in networks, with proofs checked by the Lean proof assistant. Formal checking verifies a proof against a precisely stated mathematical claim. Expert assessment is still needed to establish what that claim resolves and how it relates to the original problem.
OpenAI also reported constructing a singularity in three-dimensional fluid flow under external forcing. The report explicitly distinguishes this result from the unforced Navier-Stokes case, which remains open. Describing the work simply as solving Navier-Stokes would erase that distinction.
On October 6, OpenAI released 722 manuscripts in 372 result families from an unreleased model, with partial Lean verification. Those counts describe released material, not 722 independently accepted discoveries. Formal proofs, proposed arguments, and results still undergoing expert assessment should be identified separately when evaluating AI’s mathematical contributions.
Further reading: slide 67.