AI/News

Epoch AI: the memory shipped for AI chips through 2027 could run 30 to 170 million agents at once - and using a fifth of it would cost more than the industry earns

Epoch's new estimate works from high-bandwidth memory shipments, not chip counts. Running 20 per cent of that capacity at today's API prices implies $2.6 to 5.3 trillion a year of spending, against roughly $1 trillion of model-developer revenue by end-2027 if revenues keep growing fivefold a year.

By Daily Aletheia · Checked against the primary source · 3 October 2026 · 3 min read


Image: Epoch AI - Figure 1, "Potential agent capacity under our central hardware assumptions", from "How many AI agents could we run?", epoch.ai, 2 October 2026

Epoch AI published "How many AI agents could we run?", by Jason Li, on 2 October - an attempt to size the AI hardware build-out in the unit that matters for agents: how many could run at the same time. Its answer is that the high-bandwidth memory shipped during 2025 to 2027 could support "tens to hundreds of millions of concurrent frontier-model agents" - about 30 to 170 million under its central assumptions - once deployed and fully allocated to that work.

How the number is built

Epoch counts memory rather than chips because HBM is both the component that limits how many requests a GPU can hold in flight and the one in shortest supply; three firms make it - Micron, Samsung and SK hynix - and TrendForce's figures put 2026 shipments above 3.75 billion gigabytes. Memory is converted into "GB300 equivalents" of 288 GB each, then multiplied by how many agent sessions one such unit can serve. For open models that comes from SemiAnalysis's AgentX serving benchmark; for closed models it is inferred from what an hour of continuous agent work costs at API prices, read from TraceLab's logged Codex and Claude Code sessions.

Those hourly costs are a finding in themselves: Codex workloads averaged about $16 to 18 an hour of continuous activity, Claude Code workloads $24 to 50. Epoch uses $30 an hour and a $5 per GPU-hour rental price as its reference case.

What the number implies

An agent can run 168 hours a week, 4.2 times a full-time employee's 40, so 30 to 170 million agents supply as many weekly working hours as 140 to 720 million people - against a US population of 342 million and an estimated 100 million knowledge workers. On a more efficient open model the ceiling is higher still: applying DeepSeek V4 Pro's benchmarked serving figures to the same hardware yields about 1.9 billion concurrent agents, "as many weekly working hours as 8 billion people each working 40 hours".

The sting is on the demand side. Using just 20 per cent of the central capacity estimate "would imply $2.6-5.3 trillion a year in API-equivalent spending, against roughly $1 trillion in developer revenue by end-2027 at fivefold annual growth". Epoch's conclusion: "Demand could fall behind this potential supply, creating an overabundance of capacity. The key uncertainty is whether sustained, rapid growth in demand for AI services will justify the investment."

The estimate assumes full deployment, full allocation to agent workloads, and that HBM4 systems serve twice the agents per gigabyte of HBM3E; Epoch publishes the sensitivities and its code in the appendices.

Sources & further reading

  1. 01Epoch AI - How many AI agents could we run?, 2 Oct 2026 ↗Primary