All topicsSTATE OF AI REPORT.

AI inference is becoming a vast business

Agents repeatedly call trained models as they work, creating recurring demand across APIs, subscriptions, and infrastructure. Many companies are rushing with their empty buckets to capture water from this inference (revenue) waterfall. In fact, I would argue that the industry is in a state of “either you die trying to get to the frontier, or you live long enough to serve inference.”

Questions in this section

Is AI inference getting cheaper?Does cheaper inference mean lower total AI spending?Do frontier lab revenues measure inference alone?Does faster growth prove that adopting AI causes it?Can a specialized open model beat a frontier model on business work?What should businesses measure beyond token price?Is AI eliminating jobs?What training data are AI labs buying?How concentrated is AI startup funding?

Across five benchmarks, Epoch AI estimates that achieving a fixed score has become 47% cheaper each quarter since 2023, roughly 13-fold cheaper each year. That makes more work economical to delegate.

The minimum inference cost of achieving a given benchmark score falls rapidly.
The minimum inference cost of achieving a given benchmark score falls rapidly.

Revenue is growing across the industry

OpenAI and Anthropic's combined reported annualized revenue run rates reached an astonishing $105B by late summer, up from about $30B at the start of the year. These figures cover their whole businesses across subscriptions, coding products, and APIs. In the words of OpenAI, “we cannot miss this moment because we are distracted by side quests”.

OpenAI and Anthropic's combined reported annualized revenue run rates rise sharply through 2026.
OpenAI and Anthropic's combined reported annualized revenue run rates rise sharply through 2026.

But this rapid revenue growth extends beyond the frontier labs. In Standard Metrics' preliminary Q2 data, AI-native companies with $1M to $20M in annualized revenue grew 256% year over year at the 75th percentile, versus 90% for companies adding AI to existing software. Above $20M, the comparison was 172% versus 53%. I suspect that adopting AI contributes to faster growth, although these comparisons do not establish causality.

AI-native companies grow about three times as fast as AI-enabled companies at the upper quartile.
AI-native companies grow about three times as fast as AI-enabled companies at the upper quartile.

Software investors face a harder judgment

Claude Cowork and its business plugins across legal, finance, marketing and more brought whiplashing uncertainty to software markets. Nearly $285B was wiped from software stocks in two weeks during February's “SaaSpocalypse.” Mythos Preview later prompted fears that cybersecurity companies would become irrelevant, contributing to further losses before a sharp rebound. I thought that was a silly take, given how much demand these capabilities create for better defense. Investors have to judge whether agents will erode an incumbent's pricing power, expand its market, or enable a competitor to replace it. Or whether the incumbent's data moat is still a moat that an agent can build on top of. We have to decide well before the answer appears in revenue.

Software shares sell off around the SaaSpocalypse and recover sharply later in the year.
Software shares sell off around the SaaSpocalypse and recover sharply later in the year.

AI spending is concentrated in a small group of firms

Broad adoption statistics hide enormous differences in intensity. In Ramp’s August 2026 data, the median firm in the top 1% of AI spenders paid $7,205 per employee per month, compared with $12.50 for the median firm. A separate Ramp analysis found that 1% of customers accounted for about 80% of observed spending on OpenAI and Anthropic.

These figures describe Ramp’s customer sample. They do not imply that 1% of all companies account for the same share of global AI spending. They do show why “uses AI” is too coarse a category for evaluating business change. An occasional subscription and an operation built around sustained agent work can both count as adoption while producing very different workflows and bills.

Further reading: slide 95.

Specialized models can improve the economics of real work

Application companies can start with frontier APIs, then use expert feedback to improve models for their own domain. Harvey’s post-trained GLM-5.2 ran in production at a reported 54.8% lower cost per review-table cell than Sonnet 5. On Harvey’s answer-quality measure, it scored 0.903 against 0.867 for Fable 5.

Mercor trained an open Qwen model on 1,928 expert tasks. Its first-attempt pass rate on 480 held-out APEX-Agents tasks rose from 16.11% to 27.29%. These are results on defined specialist evaluations, rather than evidence that the models became generally superior. They demonstrate the value of knowing what a correct answer looks like in a customer’s work and having the data to train and test for it.

Further reading: slide 113.

Specialized application companies improve open models on their own work.
Specialized application companies improve open models on their own work. Report slide 113.

Customer feedback is becoming part of product development

The useful loop starts with production failures. Preserve the context and tool calls, obtain expert corrections, and turn the failure into a test that can be rerun. A product team can then determine whether the remedy is better retrieval, memory, routing, a tool change, or post-training. Taking control of the model becomes worthwhile when the gains in quality, latency, or cost justify the work.

Customer-service vendors are beginning to automate parts of this loop. PolyAI says customers used Wren for 87% of deployed changes in the week of September 21, up from 52% in the week of July 13. Decagon’s Autopilot tests proposed fixes on past conversations before a person approves them. Decagon reports a 93% score versus 83% for certified staff on its own diagnostic benchmark. Those company-reported measures describe building and maintaining agents, rather than customer ROI.

Further reading: slide 114, slide 115.

The business metric is a correctly completed job

Capability can vary sharply even within one profession. Opus 5 scored 100% across twenty attempts on four structured accounting tasks, but passed only 12.3% of ATLAS-Finance’s hundred simulated banking assignments, which require both a correct workbook and proper delivery. In one failed assignment it passed its own checks because those checks relied on the same wrong financial assumption.

Cost per token is similarly incomplete. Reasoning systems consume different quantities of input, cached, reasoning, and output tokens to produce an answer. The relevant comparison includes the cost of retries, review, and delivery at a specified quality level. A cheaper model call is economically useful when it helps produce more correctly completed work.

Further reading: slide 101, slide 110.

AI adoption and employment are growing together at some firms

Early evidence does not support a single story about AI and jobs. Ramp linked AI spending to Revelio headcount records for 21,559 US firms. Heavy spenders added 10.2% to their headcount over two years and 12% at entry level. Almost all of the gains were in technology companies. Light adopters did not separate from the control group. Faster-growing firms may have more money to spend on AI, so the association does not establish that AI caused the hiring.

Other evidence points to pressure on particular jobs. The report describes tentative signs of slower hiring among 22-25-year-olds entering AI-exposed occupations, without a clear relative rise in unemployment. Across US office occupations between May 2023 and May 2025, employment rose 37% for data scientists but fell 18% for data-entry workers and 9% for customer-service staff. Those occupational changes have multiple possible causes.

The infrastructure build-out also creates demand for electricians, construction workers, and engineers. The Economist estimates 320,000 additional US infrastructure jobs and 730,000 additional jobs in AI-related professions relative to broader hiring trends. These estimates and the company-level studies use different populations and methods. Adding them together would not produce a net count of jobs created or eliminated by AI.

Further reading: slide 103, slide 104, slide 106.

Employment is growing in some occupations and shrinking in others.
Employment is growing in some occupations and shrinking in others. Report slide 106.

Labs are buying expert work and company records to train agents

A finished answer leaves out the decisions needed to produce it. An execution trace records the steps, tool calls, and corrections made while completing a task. Labs can use expert-corrected traces to teach agents how to perform work. Private company records supply a different resource: procedures, histories, and domain context that may be scarce on the public internet.

Supplying this material has become a substantial business. The report records Mercor at $2B in gross annualized revenue in June 2026, Handshake AI at nearly $1B in gross annualized revenue in April, and micro1 at more than $500M in annualized revenue in September. Surge reported $1.2B for the full 2024 financial year, while Scale recorded just under $1B for 2025. The dates and accounting bases differ, so these figures should not be summed into a market total.

Evaluation environments help create the next order for data. A test reveals where an agent fails, a lab buys examples or expert corrections for that work, and the improved model is tested again. Suppliers compete on access to useful expertise, realistic tasks, and the ability to judge whether the work was done correctly. More data is valuable when it addresses a failure the lab can measure.

Further reading: slide 116, slide 117.

Training-data suppliers report substantial revenue on different accounting bases.
Training-data suppliers report substantial revenue on different accounting bases. Report slide 117.

Large funding rounds account for most of the capital raised

In the report’s tracked company dataset, rounds of at least $250M accounted for 94% of dollars invested in 2026, compared with 10% in 2022. This is a measure of the capital raised by that sample, not the proportion of companies receiving funding or a census of all venture investment. A small number of large rounds can dominate the dollar total.

Valuations have also risen quickly among the selected private AI companies. Historical fits imply valuation-doubling times of 3.5 to 13.3 months. These are fundraising marks from a selected group, not forecasts of future returns. Revenue growth, the cash required to build infrastructure, and the terms of each financing are needed to assess what a higher valuation represents.

Further reading: slide 152, slide 156.

Evidence you can use

Quarterly decline in cost at a fixed score

2023 onward. Analysis published 2026-09-22

Quarterly decline in cost at a fixed score
BenchmarkCost decline
AIME (OTIS Mock)Source47.1%
Chess puzzlesSource43.0%
FrontierMath, tiers 1-3Source53.1%
GPQA DiamondSource47.0%
Mystery game puzzlesSource44.0%

These are the model-free estimates in Epoch AI’s Table 1. The average quarterly decline is 47.0%. The comparison holds benchmark performance fixed. A forecast of total spending also needs usage, task difficulty, and the number of model calls.

Sources: Epoch AI: The plunging price of thought.

Frequently asked questions

Answers drawn from the report and the sources below.

What should businesses measure beyond token price?

Measure the cost of a correctly completed job, including reasoning tokens, retries, review, and delivery. Structured accounting success and whole-assignment banking success differ sharply in the report, illustrating why a model score alone is insufficient.

Source: State of AI Report 2026, slide 101: AI performance still varies widely across financial work · State of AI Report 2026, slide 110: Reasoning makes token price a poor proxy for the cost of an answer.

Is AI eliminating jobs?

The report finds mixed evidence. Heavy AI spenders in Ramp’s US company sample hired faster, while some exposed occupations and younger entrants showed weaker hiring. Infrastructure construction also added employment. These studies do not establish a single causal estimate of net jobs created or eliminated by AI.

Source: State of AI Report 2026, slide 103: Heavy AI spenders hire faster…except for scientists · State of AI Report 2026, slide 104: Early AI labor studies point to risks for junior workers · State of AI Report 2026, slide 106: The AI build-out is adding jobs even as some office roles shrink.

What training data are AI labs buying?

Labs buy expert work, corrections, execution traces, and domain records. Traces show the steps and decisions behind an answer, while company records supply context that may be scarce publicly. Mercor, Handshake AI, micro1, Surge, and Scale illustrate the scale of the supplier industry, using different revenue periods and accounting bases.

Source: State of AI Report 2026, slide 116: What training data is valuable? Execution traces and in-domain records · State of AI Report 2026, slide 117: Teaching AI is now generating billions of dollars in revenue.

Sources and dates

2026 report snapshot. Preview revised 2026-10-07. Individual data periods and source checks are listed below. This is not a claim that every source was updated on that date.

  1. Epoch AI: The plunging price of thought2026-09-22. Primary source checked 2026-10-07.
  2. State of AI Report 2026, slide 95: The top 1% of firms spend about 580x the median per employee on AI2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  3. State of AI Report 2026, slide 101: AI performance still varies widely across financial work2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  4. State of AI Report 2026, slide 103: Heavy AI spenders hire faster…except for scientists2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  5. State of AI Report 2026, slide 104: Early AI labor studies point to risks for junior workers2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  6. State of AI Report 2026, slide 106: The AI build-out is adding jobs even as some office roles shrink2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  7. State of AI Report 2026, slide 110: Reasoning makes token price a poor proxy for the cost of an answer2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  8. State of AI Report 2026, slide 113: Vertical AI companies post-train open models past the frontier in their own domain2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  9. State of AI Report 2026, slide 114: Production feedback guides improvements across the AI stack2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  10. State of AI Report 2026, slide 115: Agents now build and fix customer service agents, and the customer's staff approve2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  11. State of AI Report 2026, slide 116: What training data is valuable? Execution traces and in-domain records2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  12. State of AI Report 2026, slide 117: Teaching AI is now generating billions of dollars in revenue2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  13. State of AI Report 2026, slide 152: Private AI valuations have risen fast, very fast2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  14. State of AI Report 2026, slide 156: Mega rounds continue to eat the lion’s share of private AI company raises2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  15. State of AI Report 20262026 report snapshot. Report PDF. See individual slide references for the expanded analysis.
  16. $285B - press.airstreet.comOriginal source link retained from the launch essay.

Cite this page

Benaich, Nathan. “AI inference is becoming a vast business.” State of AI Report 2026. Published 2026-10-08.