Every forecast for AI compute demand written in the last two years has turned out to be too conservative within months of publication. Data centers, chips, power grids, and networking fabric are all being redesigned in real time to keep pace with a growth curve that keeps outrunning the plans built around it. Understanding how AI infrastructure is evolving means looking past the headline chip announcements and into the compute economics, the physical constraints, and the new architectures being built specifically for reasoning and agentic AI rather than the chat-based systems of just a couple of years ago.
THE SCALING LAWS DRIVING THE DEMAND CURVE
Much of today’s compute crunch traces back to three distinct scaling laws that each pull in the same direction: more compute, always. Pretraining scaling, the process of building bigger models on bigger datasets, has increased compute requirements by roughly 50 million times over the past five years alone. Post-training scaling, the fine-tuning of existing models for specific real-world tasks, requires around 30 times more compute during inference than pretraining did. And test-time scaling, the iterative reasoning that lets an agent explore multiple possible responses before settling on one, can consume up to 100 times more compute than traditional inference.
Traditional data centers, built for steady, predictable workloads, were never designed to absorb that kind of compounding demand. That mismatch is why the industry has started talking about AI factories rather than data centers: facilities that don’t just store and process data but manufacture intelligence at scale, measured in AI token throughput rather than storage capacity or uptime alone. The distinction matters because it changes what gets optimized. A conventional data center is judged on reliability and utilization. An AI factory is judged on how many tokens of usable intelligence it can produce per dollar and per watt.
THE PHYSICAL BUILD-OUT IS ACCELERATING FASTER THAN ANYONE PLANNED
The scale of physical construction underway is difficult to overstate. Global data center capacity could more than triple by 2030, with demand reaching at least 170 gigawatts, driven overwhelmingly by AI workloads. Hyperscalers and colocation providers have already announced plans for more than 2,600 new data centers, and roughly a quarter of them are headed to cities with no existing data center footprint at all, pushing the AI build-out into entirely new geographic markets. By the early 2030s, the global footprint of data center facilities is expected to approach 11,000.
Power has become the single most binding constraint on that growth. US power demand from AI data centers could grow more than thirtyfold by 2035, reaching 123 gigawatts, up from just 4 gigawatts in 2024. A five-acre facility that once used 5 megawatts of power can now draw 50 megawatts once it’s retrofitted with GPUs alongside its CPUs, and some data center campuses now in early planning stages could eventually consume 5 gigawatts on their own, more power than the largest nuclear or gas plants operating in the United States today. Grid connection requests in some regions now face a wait of up to seven years, and rising residential electricity rates in the top data center markets have already become a political flashpoint in several states.
CHIPS ARE BEING REDESIGNED AROUND INFERENCE, NOT JUST TRAINING
The hardware itself is shifting in step with the workload. Where training dominated compute planning for years, inference, especially the reasoning-heavy inference behind agentic AI, has become the primary driver of AI economics. Chipmakers have responded with architecture built specifically around that shift. Google’s eighth-generation Tensor Processing Units split into two distinct chips for the first time: one built for high-throughput training, packing 9,600 chips into a single superpod capable of 121 exaflops of compute, and a second built purely for low-latency reasoning and inference, tripling on-chip memory and cutting on-chip latency by up to five times to handle the ultra-low latency that agentic workflows demand.
AMD has been making similar generational leaps on the inference side specifically. Independent benchmarking found AMD’s newest Instinct GPU delivering roughly 3.1 times the inference throughput of its immediate predecessor on a standard large language model benchmark, and AMD’s broader submissions crossed one million tokens per second in aggregate for the first time at multinode scale, a milestone the industry increasingly treats as the real marker of production readiness. Nine separate hardware and cloud providers submitted results using that same chip family, underscoring how quickly a competitive, multi-vendor ecosystem has formed around inference-optimized hardware specifically.
NETWORKING AND STORAGE ARE BECOMING JUST AS IMPORTANT AS THE CHIPS
A cluster of the fastest chips in the world is only as useful as the network and storage feeding them, and that layer has needed almost as dramatic a redesign as compute itself. New data center fabrics are being built to connect well over a million accelerator chips into what functions as a single supercomputer spanning multiple physical sites, using collapsed network architectures that eliminate the performance tax older designs imposed as clusters scaled up. Storage systems have had to keep pace too, with newer managed file systems delivering roughly ten times the bandwidth of the prior generation specifically so that expensive accelerator chips aren’t left idle waiting on data. Losing even a few percentage points of accelerator utilization to a storage bottleneck is now treated as a real financial cost, since keeping training clusters running at 95 percent utilization or higher directly protects the return on an enormous capital investment.
Orchestration software has evolved just as fast. Kubernetes-based systems built specifically for agent-native workloads now aim to start compute nodes several times faster than before and cut model loading times dramatically, because in the agentic era a single user request can trigger a cascade of specialized agents working together, and every millisecond of startup latency compounds across that whole chain.
THE TELECOM AND COLOCATION LAYER IS BEING RESHAPED TOO
As centralized compute strains against power and permitting limits, demand for infrastructure closer to end users is growing in parallel. Telecom operators, who largely missed out on the last wave of tech-driven data growth, are now finding a new role to play, given that many already control the fiber networks, real estate, and power access that distributed AI compute requires. Fiber connectivity for new data centers alone represents a global revenue opportunity estimated between $30 billion and $50 billion by 2030, and the broader GPU-as-a-service market, distinct from what hyperscalers offer directly, could be worth $35 billion to $70 billion by the same year. Accelerated compute workloads are growing at more than 30 percent annually and are expected to represent more than two-thirds of all data center demand within five years, which is pulling telecom infrastructure and colocation providers further into the center of the AI build-out than they’ve ever been before.
CLOSING THE GAP WILL REQUIRE MORE THAN JUST BUILDING FASTER
Industry surveys of power and data center executives consistently point to the same handful of bottlenecks: grid capacity and interconnection delays, supply chain disruptions for critical components, permitting timelines that routinely stretch past two years, and a shortage of skilled labor to build and operate these facilities. The strategies gaining the most traction combine several approaches at once. Liquid cooling and even rooftop rainwater collection are being used to offset the enormous water and power demands of cooling AI-dense racks. Chip-level innovations, including moving power delivery to the back of the chip and using light rather than electrical wiring for on-chip data transmission, are squeezing meaningful efficiency gains out of existing silicon. And on the grid side, allowing data centers even a small amount of load flexibility, curtailing consumption by as little as 1 percent during periods of peak demand, could unlock well over a hundred gigawatts of new capacity without requiring proportional new grid investment.
THE OUTLOOK
None of this suggests the compute build-out is anywhere close to finished. Spending across data centers, chips, and grid infrastructure combined is already moving into the trillions of dollars, and the industries involved, chipmakers, hyperscalers, utilities, and telecom operators alike, are all racing to build capacity for a demand curve that has consistently exceeded even their own most aggressive projections. What’s changed is the sophistication of the response. AI infrastructure in 2026 is no longer just about buying more GPUs. It’s about redesigning chips around inference rather than training, rebuilding networking and storage so accelerators are never left waiting, and rethinking the relationship between compute and the power grid itself. The organizations and countries that master that fuller picture, not just the chip race but the whole stack underneath it, will be the ones actually able to convert AI’s raw growth into sustained, deliverable capacity.
References
Grundin, G., Frade, M., Cubela, S., and Lajous, T. “Issue Brief: AI Infrastructure.” McKinsey & Company, February 27, 2026. https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/issue-brief-ai-infrastructure
Vahdat, A. and Lohmeyer, M. “What’s Next in Google AI Infrastructure: Scaling for the Agentic Era.” Google Cloud Blog, April 22, 2026. https://cloud.google.com/blog/products/compute/ai-infrastructure-at-next26
Harris, D. “AI Factories Are Redefining Data Centers and Enabling the Next Era of AI.” NVIDIA Blog, updated December 2025. https://blogs.nvidia.com/blog/ai-factory/
Stansbury, M., Marchese, K., Hardin, K., and Amon, C. “Can US Infrastructure Keep Up With the AI Economy?” Deloitte Insights, June 24, 2025. https://www.deloitte.com/us/en/insights/industry/power-and-utilities/data-center-infrastructure-artificial-intelligence.html
“AMD Instinct GPU MLPerf Inference Results: Performance, Scale, and Reproducibility for AI Deployments.” Principled Technologies, July 8, 2026. https://www.principledtechnologies.com/clients/reports/AMD/Instinct-GPU-MLPerf-0726/index.php