Foundation models get most of the attention, but the infrastructure underneath them increasingly determines who wins the race to build and deploy them. That infrastructure question has stopped being purely technical. It now touches geopolitics, chip architecture, and where intelligence actually gets to run. Understanding how AI infrastructure powers the next generation of foundation models means looking at three distinct layers: the physical buildout racing to get chips online faster than rivals, the smarter matching of compute to workload inside data centers, and the push to bring foundation models out of the cloud entirely and onto the devices where physical action happens.
SPEED HAS BECOME THE DECISIVE VARIABLE IN THE GLOBAL BUILDOUT
A new financial analysis from the Carnegie Endowment for International Peace, built on a detailed model of data center economics across ten countries, arrives at a finding that reframes how the AI infrastructure race should be understood. Time to power, not energy costs or tax incentives, is what separates winning projects from losing ones. For a typical 100 megawatt US data center, each additional year of delay costs roughly $550 million in lifecycle value, about 5.5 percent of the facility’s total worth. A one-year delay costs more than doubling electricity prices, more than losing state tax incentives, and three times as much as moderate tariffs on servers combined.
The reason is straightforward once you look at the mechanics. Every month a data center sits idle is a month of GPU hours that could have been sold, capital that could have been earning a return, and compute that could have gone toward training a stronger model that attracts more customers and funds the next generation of research. That compounding advantage is why a facility in Abilene, Texas took delivery of its first Nvidia chips and was running an estimated 100,000 GB200 chips within six months, while a similarly ambitious European initiative was still finalizing its formal proposal process more than a year after being announced. The United States currently hosts around three-quarters of the world’s advanced AI computing clusters, but that lead is fragile enough that a nine-month improvement in typical project speed would let the US overtake the United Arab Emirates as the most competitive site in the world, while a one-year delay would drop it to fifth place behind the UAE, Finland, Canada, and India.
Grid connection delays sit at the center of the problem. New power sources took an average of five years to connect to the US grid in 2023, up from under two years between 2000 and 2007, and transformer wait times have stretched past two years in some cases. Behind-the-meter power, generating electricity on-site through gas turbines or solar microgrids rather than waiting on grid connections, has become the industry’s most common workaround. It can shave a year or more off time to first operation, though it carries higher operating costs and real environmental tradeoffs. The countries and companies that have found ways to compress this timeline, whether through streamlined permitting, load flexibility programs, or aggressive behind-the-meter deployment, are the ones actually converting their compute investments into usable training and inference capacity fastest.
NOT EVERY WORKLOAD NEEDS A GPU
While the geopolitical race focuses on raw buildout speed, a quieter shift is underway in how enterprises match compute to workload once that infrastructure is actually built. The rush to adopt generative AI pushed many organizations toward GPU-heavy architectures by default, on the assumption that all AI work demands the same level of compute intensity. That assumption doesn’t hold. Generative AI is just one segment of a much broader spectrum of AI use cases that includes traditional machine learning, statistical analysis, and rule-based systems, each with very different performance requirements.
Accelerated CPUs, general-purpose processors enhanced with integrated AI acceleration capabilities like optimized matrix instruction sets and expanded memory bandwidth, are emerging as the pragmatic answer for a meaningful share of that spectrum. They handle model training on smaller datasets, batch inference, and latency-tolerant AI services well, and because they build on familiar x86 architecture, they integrate into existing IT environments without the redesigned cooling and power infrastructure that dense GPU clusters require. According to a recent IDC study sponsored by Intel, more than 40 percent of organizations are already adopting this kind of hybrid strategy, reserving GPUs for genuinely high-demand cases like large-scale model training and high-concurrency, real-time applications while routing everything else to CPU-based systems that cost meaningfully less to run.
This matters for the foundation model conversation specifically because GPU availability remains constrained even as the physical buildout races ahead, and the high power density of GPU-dense systems can strain data center operations that weren’t designed around it. A workload-aware strategy, auditing existing AI workloads to identify where accelerated CPUs can substitute for GPUs without compromising performance, is becoming as much a part of infrastructure planning as the chip procurement itself.
FOUNDATION MODELS ARE MOVING OUT OF THE CLOUD ENTIRELY
The third major infrastructure shift is arguably the most consequential for where foundation models actually get used. Edge computing, processing data close to where it’s generated rather than in a distant centralized data center, has existed for decades in limited forms. What’s changed is the ability to deploy and run full foundation models directly on edge devices themselves, moving intelligence from the cloud onto local hardware.
Two forces have made this possible at the same time. Specialized AI silicon, including neural processing units and other high-performance modules, provides the raw compute. But it’s software techniques, particularly quantization, that do the real enabling work. By reducing the mathematical precision of a model’s weights, quantization allows state-of-the-art foundation models to shrink dramatically in size without a major loss in capability, letting edge devices host reasoning power that used to require a cloud connection entirely.
The applications this unlocks go well beyond faster response times. For autonomous vehicles and industrial robots that need to interact safely with the physical world, edge-first large language models eliminate cloud latency that would otherwise be dangerous in real-time movement, and processing sensitive data locally rather than transmitting it to the cloud creates a smaller attack surface for cyber threats, which matters enormously for regulated sectors like healthcare and defense. It also opens up environments where the cloud was never a realistic option to begin with, since removing the connectivity tether lets sophisticated AI run in deep-sea exploration, remote mining, and other settings where an internet connection simply doesn’t exist or can’t be trusted. These aren’t incremental improvements on cloud-based deployment. They’re use cases that were fundamentally impossible until foundation model intelligence could live at the point where physical action actually happens.
THE THREE LAYERS ARE CONVERGING
None of these three shifts operate in isolation. The speed of the physical buildout determines how much raw training capacity is available to produce the next generation of frontier foundation models in the first place. The workload-aware matching of GPUs and accelerated CPUs determines how efficiently that capacity gets used once it’s online, stretching scarce and expensive accelerator hours further across a growing range of enterprise AI applications. And the move to the edge determines how much of that trained intelligence actually reaches the physical world in a form fast and secure enough to be useful, rather than sitting locked behind a cloud API with too much latency for real-time action.
Together they describe an infrastructure landscape that looks very different from just training bigger models on bigger clusters and calling it progress. Getting compute online faster than competitors, deploying the right processor for each workload instead of defaulting to the most expensive option, and compressing trained intelligence down to something that can run locally on a robot, a vehicle, or a sensor in a location with no connectivity at all: these are now the actual battlegrounds where the next generation of foundation models will be built, deployed, and made useful.
THE OUTLOOK
The lead in any one of these three areas is far more fragile than it appears from the outside. A country with a commanding chip cluster advantage today can fall behind within a year if a rival streamlines its permitting process. An enterprise that GPU-maxes its entire AI stack will find itself paying for idle capacity that a workload-aware competitor is directing more precisely. And a foundation model that only runs in the cloud will simply be unavailable for an entire category of physical, real-time, or connectivity-constrained applications that a properly quantized edge deployment can already serve today. The organizations and countries that treat all three layers, speed, workload matching, and edge deployment, as parts of a single infrastructure strategy rather than separate technical decisions will be the ones that actually convert today’s compute investment into tomorrow’s usable AI.
References
Phillips-Robins, A., Tawil, T., and Winter-Levy, S. “The Compute Coalition: How to Build the Future of AI in the Free World.” Carnegie Endowment for International Peace, June 8, 2026. https://carnegieendowment.org/research/2026/06/the-compute-coalition-how-to-build-the-future-of-ai-in-the-free-world
West, H. “Rethinking AI-Ready Infrastructure: The Strategic Role of Accelerated CPUs.” Intel Community, August 4, 2025. https://community.intel.com/t5/Blogs/Tech-Innovation/Data-Center/Rethinking-AI-Ready-Infrastructure-The-Strategic-Role-of/post/1707348
“AI Foundation Models at the Edge: Enabling Real-Time Physical Intel.” SDG Group, April 6, 2026. https://www.sdggroup.com/es-mx/insights/blog/ai-foundation-models-at-the-edge-enabling-real-time-physical-intel