The 2022 report showed AI becoming useful across scientific domains. AlphaFold predictions expanded dramatically, and reinforcement learning controlled plasma in a research tokamak. These advances also increased the need to check whether machine-learning methods supported the scientific claims being made.
Predicted protein structures become widely available
AlphaFold’s database expanded to around 200 million predicted structures. The report emphasized that the downstream benefits would take years to emerge. A predicted structure could guide a research question without being equivalent to an experimentally measured structure or a successful therapeutic intervention.
Reinforcement learning controls plasma in a tokamak
DeepMind trained a system in simulation and deployed it to adjust magnetic coils in Lausanne’s TCV tokamak. The result demonstrated flexible plasma control in that device. It was not a demonstration of commercial fusion power or a solution to every engineering challenge in fusion.
The report warned that information unavailable in a real prediction setting can leak into training or evaluation. This can make a model appear successful while undermining the scientific claim. Strong-looking results therefore required careful attention to how data were collected and separated.
Protein language models were used to assess new viral variants
BioNTech and InstaDeep built an Early Warning System using a protein language model to assess viral spike sequences. In the validation described by the report, it identified all 16 WHO-designated variants an average of more than one and a half months before their official designation. This was evidence for prioritizing potentially risky variants from sequence data, with the result tied to the variants and validation period examined.
Predicted structures and database usage do not measure discoveries or clinical benefits. Plasma control in one device does not establish net-energy fusion. Data leakage can invalidate apparently strong evaluations.
They greatly expanded the set of predicted protein structures available to researchers. They were computational predictions, and the report expected downstream scientific benefits to emerge over time.
No. The reported result was reinforcement-learning control of plasma in the TCV tokamak, an important control demonstration rather than commercial fusion power.
No. They were predicted structures in the expanded database. The scale of the resource was a modeling achievement, not 200 million separate experimental determinations.
It described researchers reported to have used the database. Usage indicated reach, but it did not count discoveries, validated hypotheses, or therapies produced from it.
The report described deployment on the TCV tokamak in Lausanne after training in simulation. The result concerned controlling plasma configurations in that experimental device.
Controlling plasma was one component of a much larger technical problem. The reported experiment did not establish a commercially viable reactor or demonstrate net-energy fusion.
It occurs when information that should remain outside the training process influences the model, directly or indirectly. The resulting evaluation can make the system appear more generalizable than it is.
In the reported validation, the system identified all 16 WHO-designated variants an average of more than one and a half months before official designation. It used a protein language model to assess spike sequences. This historical result did not establish that every future variant would be detected equally early.
Historical snapshot published October 11, 2022. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan, and Ian Hogarth. “AI tackles protein structures and plasma control.” State of AI Report 2022. Historical report snapshot; web edition prepared 2026-10-11.