On April 18, 2026, OpenRouter registered 1.4 trillion tokens processed by Qwen 3.6 Plus in a single day. According to OpenRouter, it was the first model to cross the 1 trillion daily tokens barrier on a multi-model routing platform. In July, according to the Vercel CEO, the AI Gateway was routing more than 1 trillion tokens per day between OpenAI, Anthropic, Google, and DeepSeek. The number impresses. The context terrifies. We analyzed the data and the signal is clear: inference surpassed training, became a digital public utility, and started demanding architecture — not just capacity.
The scale: 3,700x in 30 months
In February 2024, OpenAI processed about 100 billion tokens per day. According to IO Fund, the global estimate in July 2026 is 370 trillion daily. Growth of 3,700x in 30 months. Google alone processed 3.2 quadrillion tokens in May 2026 — the same Google that processed 9.7 trillion per month two years earlier.
This is not linear growth. It is a regime change. Inference volume stopped being a byproduct of AI usage and became the main driver of compute consumption on the planet. And the cost per token plummeted at the same time: according to BenchLM, the frontier price index fell 88% since March 2023, and sufficient-quality models fell 200 to 300 times. The price drop is not just a benefit — it is what made the 3,700x volume increase financially viable. Total spend, however, may keep growing.
Inference surpassed training
In 2026, more than half of global AI compute is dedicated to serving models in production. Training is episodic. According to Gartner, the projection is 65% or more for inference by 2029. Sector studies estimate that inference represents 80 to 90% of the electricity cost in a model's production lifecycle. Inference became the dominant cost — not training.
Reasoning models generate 5 to 10 times more tokens per query. Reasoning works as a latency dial: you trade tokens for quality, but throughput remains stable. This changes the economics of inference. It is no longer one request, one response. It is one request, a chain of reasoning, one response. The volume of tokens per query grows, and the marginal cost of serving each query grows with it.
A bifurcated market: Anthropic, OpenAI, and the Chinese rise
The market bifurcated. According to Menlo Ventures, in December 2025 Anthropic held 40% of corporate spend. OpenAI fell from 50% to 27%. Leadership is not stable — it is contested every quarter.
According to Presenc AI, Chinese open-weight models jumped from 1% to 15% of the inference market in 12 months. DeepSeek already routes 22.6% of tokens in the Vercel AI Gateway, nearly matching Anthropic in volume. This is not a marginal movement. It is a structural vendor shift. Chinese open-weight models stopped being a research curiosity and became a production-scale option.
Extreme concentration and systemic risk
Concentration is extreme. According to StealthCloud Intelligence, five companies control 80% of global training capacity. The US hosts 74.5% of the world's AI compute. Northern Virginia consumes 26% of the state's electricity in data centers. TSMC is the only supplier at scale of CoWoS advanced packaging.
In June 2026, the European ESRB classified frontier AI models as systemic risk to the financial system. The same category used for systemically important banks — and that is not advisory. It is a classification that carries regulatory obligations. NIST published the AI RMF Critical Infrastructure Profile in April. The EU applies AI Act Annex III to critical infrastructure starting December 2027. AI inference is already treated as a digital public utility — with everything that implies in terms of resilience, sovereignty, and obligation of continuous service.
84% of tokens in 2030 will come from autonomous agents
According to Goldman Sachs, 84% of tokens in 2030 will come from autonomous agent workloads. Humans typing prompts will be a minority of traffic. The traffic of the future is machine-to-machine — programmatic, stateful, high-volume, and with an expectation of low latency.
This changes architecture. It is not enough to serve interactive chat. You need to architect for inference workloads that hold state, make chained calls, and consume tokens at industrial volume. The stack that serves 370 trillion tokens per day is not a bigger endpoint — it is architecture.
Conclusion: the stack that serves 370 trillion tokens per day is architecture
We repeat: AI inference is already a digital public utility. And the stack that serves 370 trillion tokens per day is architecture. At Tech86, we design inference infrastructure with sovereignty, resilience, and controlled marginal cost. Prefill/decode disaggregation, KV-cache-aware routing, speculative decoding. Multi-vendor, multi-region strategy to avoid lock-in. Planning for autonomous agent workloads, not for interactive chat.
The 1 trillion daily tokens barrier was crossed. The next barrier is 10 trillion. And whoever gets there will not be the one who buys the most GPUs — it will be the one who architects best.