The AI industry is spending hundreds of billions of dollars a year on chips and data centers, on the bet that the hardware will run autonomous agents doing the work people do today. A new Epoch AI Agent Estimate, published on October 2, 2026, asks the obvious follow-up question: how many AI agents could all of that silicon actually run at the same time? The answer, according to the research group Epoch AI, is somewhere between 30 and 170 million frontier-model agents running continuously — or about 1.9 billion simultaneous sessions if every agent ran an efficient open-weight model instead.

The report treats the hardware buildout as a supply question and lands on a striking conclusion. The binding constraint is not the raw number of chips, but the high-bandwidth memory on them. And the real unknown, Epoch AI argues, is not supply at all, but whether demand for agent services will ever grow large enough to justify the investment. The analysis was led by Epoch AI researcher Jason Li, who published the code and data in a public reproduction repository so that others can check the arithmetic.

Memory, Not Math, Sets the Ceiling

When a model generates text, it has to hold its own weights in fast memory, plus a growing scratchpad for each conversation, known as the KV cache. Producing every new token means streaming much of that data through the chip again. That makes memory capacity and bandwidth the usual limit on how many sessions a chip can serve — not raw arithmetic power.

Epoch AI therefore counts shipments of high-bandwidth memory from 2025 onward, rather than counting chips. The shipments are converted into equivalents of an NVIDIA GB300 accelerator, which carries 288 gigabytes of memory. Under the group's central assumption, the next memory generation serves twice as many sessions per gigabyte, which works out to roughly 60 million GB300-equivalents shipped through 2027.

The final estimate comes from multiplying two numbers together: how much usable memory is shipping, and how many agent sessions each unit of that memory can serve. The Epoch AI Agent Estimate is the first attempt to convert the chip buildout into a concrete agent headcount rather than an abstract spending total.

Two Kinds of Agent, Two Very Different Answers

The gap between tens of millions and billions comes down to what each agent is running. For closed frontier models — the kind behind services like Claude Code or Codex — Epoch AI cannot look inside the providers' servers, so the researchers worked backward from cost.

Real coding-agent traces suggest that continuous agent use costs roughly 30 dollars per agent-hour. The analysis assumes providers charge five to ten times their serving cost and compares that with a GPU rental price of about 5 dollars an hour. That yields roughly one to two agent sessions per GB300-equivalent, which across the shipped hardware produces the 30-to-170-million range. The central hardware assumption narrows that to 50 to 101 million simultaneous agents.

For open-weight models the calculation is direct rather than inferred, according to the report. Epoch AI used measurements from SemiAnalysis's AgentX benchmark, which replays real coding-agent sessions. An efficient open model, DeepSeek V4 Pro in this analysis, managed about 31 sessions per GB300-equivalent at 50 tokens per second per user. That implies roughly 1.9 billion simultaneous sessions. Push the per-user speed to 100 tokens per second and the figure falls to about 870 million.

The two figures are alternative uses of the same memory pool and cannot be added together — a distinction Epoch AI states explicitly. The report offers one scale comparison: 1.9 billion always-on sessions add up to the weekly hours of about eight billion people working 40-hour weeks. It also warns that agents differ in speed and quality, so matching hours does not mean matching work.

The Trillion-Dollar Demand Question

The report's sharpest point is that demand, not hardware, is the big unknown. As Ground Truth's October 3 coverage of the report puts it, demand could fall behind this potential supply and leave the industry with an overabundance of capacity.

To test the size of the gap, Epoch AI modeled a scenario in which only a fifth of the available capacity is effectively used. Even at 20 percent utilization, the capacity would imply between 2.6 trillion and 5.3 trillion dollars a year in API-equivalent spending. That figure sits against roughly 1 trillion dollars in annualized developer revenue by the end of 2027, assuming the industry's recent fivefold growth rate continues.

Readers should treat these numbers as a capacity ceiling, not a forecast. The analysis assumes every shipped chip is installed, powered, and dedicated to a single agent workload around the clock, while data-center construction can lag chip deliveries by a long way. The 2026 and 2027 memory shipments are projections, and the closed-model figures rest on assumed pricing markups. The coverage also notes that a viral discussion of the 1.9 billion figure dropped the distinction between the efficient open-model case and the frontier range — and that no independent expert review of the report had appeared at the time of writing.

What It Means for the Agent Economy

For the people building and deploying AI agents, the report translates an abstract capital-spending number into a unit that is easy to reason about: simultaneous agents. It also shows that the economics of the model being run can swing the answer by a factor of ten or more — one reason efficient open models matter to the industry's math, not just to hobbyists.

That math is only half of the picture. Questions of identity, security, and governance are already shaping how agents get deployed, from open identity standards for agents to tighter operating-system controls on what agents may touch. Hardware, it turns out, may be the least of the industry's worries: the capacity is being built, and the open question is whether the world wants enough agents to fill it. Ongoing coverage of the agent ecosystem is collected on the AI News topic page.