
Last year, we described reasoning becoming useful at scale with OpenAI’s o-series of models. I have lived through the improvements since then, as the amount of useful work I can iterate on and increasingly delegate to agents keeps growing.
My view is that much of the gap between people getting substantial value from AI and those getting little comes down to knowing how to use it and how to set it up. Put another way, it’s a “human skill issue,” not an “AI technology issue.” I really believe that we need a Genius Bar for AI, where people can get hands-on help choosing tools and setting them up to help them in their everyday tasks.
Benchmarks are approaching their ceilings
Across mathematics, scientific reasoning, and coding, benchmarks intended to challenge models for years are approaching their ceilings within months. We need new evaluations that reveal where models still separate and what they can reliably do.

AI is accelerating the work of building AI
AI is already helping build better AI. According to Anthropic's internal index, Claude led 26% of measured model R&D work in August, up from under 1% in February. Researchers set the tasks and supervise execution, and Anthropic reports that this is speeding up development. This is a branch of AI research that I am most excited about.

For many, Karpathy's autoresearch makes part of that loop tangible: an agent edits training code, runs five-minute experiments, and keeps improvements overnight on one GPU. This automates a useful part of research, although sustained, fully autonomous recursive improvement remains to be demonstrated.
What I want to see next is agents developing scientific taste: choosing experiments, recognizing promising directions, and knowing when to abandon a familiar approach. As I argued in Can AI learn scientific taste?, learning that judgment may require the alternatives, failures, and decisions that papers leave out. The ambition is an AlphaGo-like shift in scientific strategy.
Open models are changing who builds on AI
Our analysis with Zeta Alpha tracks AI papers on arXiv that mention 21 model families. Chinese open-weight families gained ground sharply between 2024 and 2026, and Qwen overtook Llama. Closed US models still accounted for 42% of mentions in the 2026 snapshot.
Paper mentions reveal which models researchers choose to study and build on. They do not measure commercial revenue or establish a universal capability ranking. A model used widely in research can become the starting point for new methods, evaluations, and specialist applications. The research ecosystem can shift even while American labs retain leading closed models.
Further reading: slide 6.
Epoch AI found flaws in nine of 15 benchmarks
A high score is only useful if the test measures what it claims to measure. Epoch AI found substantive flaws in nine of its first 15 benchmark reviews. It examined tasks, grading, prompts, tools, and resource limits, finding problems such as broken scoring and exploitable environments. Four benchmarks met its minimum standards with caveats. Two could not be judged with the available information.
Even a sound evaluation answers a particular question. FrontierSWE gives agents 20 hours to build software in a common harness. MirrorCode asks them to reconstruct programs from documentation and a runnable reference, with up to seven days and ten billion tokens. They produce different leaders. The task, completion criteria, and spending allowance belong beside the model’s name whenever we compare results.
Further reading: slide 36, slide 41.
Progress on a task is different from finishing it
On FrontierChallenge’s 97 scientific workflows, GPT-5.6 Sol earned 87.9 out of 100 for satisfying individual requirements but fully completed only 20.6% of tasks. On OSWorld 2.0’s computer workflows, Opus 5 earned 77.7% partial credit and completed 44.3%. An otherwise correct expense claim can remain unfinished because the agent misses an approval or never submits it.
Task duration needs similar care. METR’s 50% time horizon estimates the human-expert time associated with tasks an agent can complete half the time. It is not a measurement of how long an agent can safely operate unattended. With only five tasks taking humans more than 16 hours, METR warns that estimates at the upper end are unreliable. Longer and better-scored evaluations are necessary to tell whether capability is becoming dependable.
Further reading: slide 42, slide 43.

Research engineering is advancing faster than research judgment
In a shadow evaluation, agents received the central questions from unpublished NeurIPS submissions and the original authors reviewed their work. Each run had six days, GPUs, and $3,000 of API credit. Opus 4.8 completed the engineering without human help, but its two papers received rejection scores of 2/6 and 1/6. The recurring weaknesses included uncreative fixes, ineffective backtracking, and a poor sense of what would merit publication.
Agents can make experiments easier to implement while people still choose the research questions and judge the results. Demonstrating sustained autonomous research requires evidence that the system can make those choices too.
Further reading: slide 17, slide 23.
Video generation is becoming fast enough for interactive streams
Video generation usually returns a completed clip after a prompt. fal adapted MiniMax’s open-weight H3 video and audio model into H3 Max, combining additional training data with a serving engine and optimized GPU kernels. The company reports producing a five-second clip in under three seconds, with about 35 times the throughput of the official H3 endpoint.
Its Director mode carries 39 frames into the next segment and remembers previous prompts. Removing the overlapping frames lets the segments play continuously. On fal.live, viewers vote on what happens next in a 24-frame-per-second stream. Their prompts affect the following segment, so the interaction still has a delay. Faster generation creates room for audience-directed video and other experiences in which the scene changes while someone watches.
Further reading: slide 48.

AI tutoring needs to help students solve problems themselves
A tutor can give the right answer while leaving the student unable to solve the next problem. A trial in Sierra Leone tested a more structured use of AI. Across 48 classrooms in 12 schools, 1,763 students were randomized to teacher-led Gemini activities or standard instruction. After eight weeks, the AI program improved math scores by 0.258 standard deviations, with a 95% confidence interval of 0.027 to 0.488.
The result applies to that classroom program, including its teachers and activities. It does not establish the effect of unrestricted chatbot use or long-term learning. TutorMoments and MathTutorBench also distinguish answering correctly from teaching effectively. Useful tutoring needs to decide when to give a hint, when to ask the student to explain, and when to let them work through a mistake.
Further reading: slide 105.
