The Other Half of Compute
Everyone is counting gigawatts and GPUs. The number that decides the return is what each one actually buys.
Everyone is counting gigawatts and GPUs. The number that decides the return is what each one actually buys.
This is an analytical framework, not financial advice. Numerical claims are referenced to their primary sources in the footnotes.
xAI stood up its first 100,000 GPUs in Memphis in 122 days. It doubled that in another 92. By early 2026 the site, Colossus, held around 555,000 of them, building toward two gigawatts of power, for a reported 18 billion dollars.1
Two sophisticated people can look at that number and reach opposite conclusions.
Jensen Huang’s view is that the only real risk is underspending. He puts the buildout at a trillion dollars and counting, and argues the company that holds back capacity loses the decade.2 Dario Amodei and Ray Dalio sit on the other side. Amodei has said it can be rational not to buy unlimited compute, because the revenue to justify it may arrive on a timeline that bankrupts whoever guessed wrong. Dalio keeps making a narrower point: a technology can succeed completely and still ruin the people who financed it.3
Same buildout. Same dollar figure. One camp calls it the obvious move of the decade and the other calls it the setup for a wipeout. They are not disagreeing about the facts. They are reading the same number and the number is the problem.
What 18 billion dollars buys
Every token a model produces runs down a physical path. Electricity has to be generated, moved across a grid, and stepped down through transformers to a voltage a data centre can use. Chips have to be fabricated at advanced nodes, which in practice means TSMC and a single supplier of the lithography machines that make the process possible. The chips have to be wired together with optical interconnect, assembled into racks, and kept cold. None of those layers move at the same speed, and the slowest one always sets the schedule.
For four years the slowest layer kept changing. In 2022 the constraint was GPUs themselves. In 2023 it was the high-bandwidth memory stacked next to them. In 2024 it was the advanced packaging that bonds the two together. By 2025 it was photonics, the lasers and transceivers that move data between racks. By 2026 it had reached power and the grid, where a new high-voltage connection can take longer to approve than the cluster takes to build. Bringing a large new source of power onto that grid now takes a median of more than four years.4
Each layer is real, each one becomes scarce in turn, and the scarcity moves to the next layer as the one before it gets solved.
Call it the capacity stack. It decides one thing: how much raw compute can physically exist. It tells you what you can run. It says nothing about how much useful work comes out the other end.
The binding constraint has moved through the stack for four years straight. Chips, memory, packaging, photonics, power. Each one stayed invisible until the one before it was solved.
The number that never makes the capex debate
Now look at a different figure.
In March 2023, running a million tokens through GPT-4 cost about 30 dollars. By the middle of 2024, the same class of capability through GPT-4o cost 2.50 dollars. By 2025 a GPT-4-grade model was available at roughly 10 cents per million tokens.5 For the rougher GPT-3.5 tier the price fell from 20 dollars per million tokens to about 7 cents in two years, a drop of more than 250 times. Epoch AI, which tracks this carefully, finds inference prices falling somewhere between 10 and 50 times a year depending on the task.6
Almost none of that came from adding watts. The capacity stack was straining the entire time. The cost of intelligence fell by two orders of magnitude anyway. These are list prices, so some of the fall is competition between providers, but most of it is a second stack that lives inside the software layer and does work the hardware never sees.
That second stack has its own layers. At the bottom is the attention kernel. The 2022 FlashAttention paper showed that a transformer was bound by memory traffic, the data shuttling between the fast and slow memory on the chip, and that rewriting the kernel to respect that traffic multiplied throughput without changing a single transistor.7 Above it sits serving. Key-value caching, which means storing a conversation’s intermediate state instead of recomputing it on every new token, turned long contexts from a quadratic expense into something a business could afford to offer. Above that sits the model itself. Mixture-of-experts routing, the design behind Switch Transformers, broke the link between a model’s total size and the compute each token triggers, so a model can hold a trillion parameters and fire only a fraction of them per word.8
Even the hardware gains are mostly architectural rather than brute force. NVIDIA’s GB200 NVL72 rack delivers up to 30 times the inference throughput of the same number of previous-generation H100 chips, at around 25 times less energy for the same work.9 The watts per chip went up. The useful work per watt went up far more.
Each of these is a multiplier on the same physical base. Stack them and you get the hundredfold collapse in the cost of intelligence that the buildout debate never mentions.
The cost of GPT-4-class intelligence fell roughly 99 percent in two years. Almost none of that came from adding power.
Compute is a product
Raw physical capacity, multiplied by how much useful work each unit of that capacity buys. The capacity stack sets the first term. The efficiency stack sets the second. They run on different clocks, they are built by different people, and the one that is currently scarcer sets the ceiling on what you can do.
Once you read compute that way, the contradictions in the capex fight resolve.
Go back to the 18 billion dollars. Jensen Huang is right that physical capacity is scarce today. A grid connection does take longer than a training run, and the firm that waits loses ground it cannot buy back at any price. Amodei is also right that the return on that capacity is uncertain. Both of them are arguing about the first term and treating the second as a constant.
It is not a constant. It is improving 10 to 50 times a year. That cuts in two directions at once. A capex bill that looks insane against today’s efficiency can look cheap against next year’s, because the same site serves far more useful work for the same power. And capacity bought to serve a workload that the efficiency stack is about to make trivially cheap is capacity that strands. The danger in the buildout is owning the wrong term: paying for raw capacity after the binding constraint has moved to the multiplier, or perfecting the multiplier when you cannot get the megawatts to run it on.
Three years ago the next sentence would have sounded like a category error.
A 2-gigawatt site with a mediocre serving stack loses to a smaller site with a better one.
Where the constraint goes after silicon
The migration does not stop at the efficiency stack either. It keeps walking.
Once serving is efficient and the power is online, the slowest layer becomes the one furthest from the metal: whether an organisation can absorb what the stack has made cheap. Jensen Huang’s own example is the sharpest version of it. A 500,000-dollar engineer who consumes only 5,000 dollars of tokens a year shows the failure mode.10 The tokens are nearly free, and the company still cannot route its own work to the capacity it already owns.
This is the layer Amodei and Satya Nadella keep returning to from opposite ends of the argument. The technical stack gets good faster than institutions reorganise around it. The final constraint on compute is organisational. It is how quickly people change what they do.
The tokens are nearly free. The bottleneck is the company.
A test you can run this week
Take any AI bet you hold, whether it is a position, a product, or a career, and do three things.
Write down which term you are actually betting on. A bet on the capacity stack is a bet that the physical scarcity of power, chips, and interconnect holds. A bet on the efficiency stack is a bet on the people and techniques that multiply the work each watt buys. Most bets are quietly one or the other, and most people have never said which out loud.
Then name the layer that binds right now. Power, today, for raw scale. Serving efficiency, today, for cost per task. Write down what has to stay true one layer below for your bet to survive. A capacity bet dies if grid timelines compress and the scarcity premium decays. An efficiency bet dies if the megawatts never arrive to run on.
Then watch the right number. Not GPU count and not gigawatts. Cost per task and useful work per watt. Those are the readings where the second stack shows up, and the second stack is where most of the last two years of progress came from.
The capacity layers, the ones you could photograph from a satellite, are mapped company by company in the companion to this piece. This is the half you cannot photograph and the half that has been compounding faster.
If you run the test, which term turned out to be the one you were quietly betting on the whole time?

New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap. Subscribe for the rest, or start with what survives.
Footnotes
-
xAI’s Colossus (Memphis) reached 100,000 GPUs in 122 days and doubled to roughly 200,000 in 92 more. By early 2026, across multiple GPU generations (H100, H200, GB200), the Memphis site held around 555,000 GPUs and was built out toward ~2 GW of capacity, for a reported ~$18 billion. https://introl.com/blog/xai-colossus-2-gigawatt-expansion-555k-gpus-january-2026 and https://x.ai/colossus ↩
-
Jensen Huang, NVIDIA GTC 2026: he projected at least $1 trillion of AI-infrastructure spending through 2027 and argued that figure “won’t be enough” to meet demand, framing data centres as “AI factories” that convert electricity into tokens at the lowest cost per unit. https://fortune.com/2026/03/17/jensen-huang-ai-infrastructure-buildout-1-trillion-dollars/ ↩
-
Dario Amodei, interview with Dwarkesh Patel (2026): being off on data-centre timing “by a couple of years can be ruinous,” because the revenue to justify a buildout arrives on an uncertain schedule. https://www.dwarkesh.com/p/dario-amodei-2 . Ray Dalio’s recurring point, repeated in June 2026, is that a technology can succeed while most of the companies financing it fail, as the internet did after the dot-com bust. https://finance.yahoo.com/markets/stocks/articles/ray-dalio-says-ai-investors-121700871.html ↩
-
The 2022 to 2026 bottleneck migration sequence (chips, memory, advanced packaging, photonics, power) is the synthesis of my earlier infrastructure work. The grid figure: Lawrence Berkeley National Laboratory, “Queued Up: 2025 Edition,” finds the median time from interconnection request to commercial operation for new generation has passed four years. https://emp.lbl.gov/publications/queued-2025-edition-characteristics ↩
-
OpenAI list pricing: GPT-4 launched at $30 per million input tokens (March 2023); GPT-4o at $2.50 per million input tokens (May 2024); GPT-4-grade capability available near $0.10 per million input tokens by 2025. Pricing history aggregated by Epoch AI and TokenCost. https://epoch.ai/data-insights/llm-inference-price-trends ↩
-
Epoch AI, “LLM inference price trends”: the cost of GPT-3.5-class capability fell from roughly $20 per million tokens (late 2022) to about $0.07 (late 2024), and inference prices decline between 10x and 50x per year depending on task tier. https://epoch.ai/data-insights/llm-inference-price-trends ↩
-
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” 2022. The paper reframes attention as memory-bandwidth-bound rather than compute-bound. https://arxiv.org/abs/2205.14135 ↩
-
William Fedus, Barret Zoph, Noam Shazeer, “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity,” 2021. Mixture-of-experts routing decouples a model’s total parameter count from the compute activated per token. https://arxiv.org/abs/2101.03961 ↩
-
NVIDIA GB200 NVL72: NVIDIA reports up to 30x faster real-time LLM inference and up to 25x lower energy and cost versus the same number of H100 GPUs, driven by rack-scale architecture rather than raw per-chip power. https://www.nvidia.com/en-us/data-center/gb200-nvl72/ ↩
-
Jensen Huang, All-In Podcast (filmed on the final day of NVIDIA GTC 2026): he said he would be “deeply alarmed” if a $500,000 engineer consumed only $5,000 of tokens in a year, expecting elite engineers to spend closer to half their salary on tokens. Low token use reads as a failure to exploit cheap capacity, not thrift. https://www.tomshardware.com/tech-industry/artificial-intelligence/jensen-huang-says-nvidia-engineers-should-use-ai-tokens-worth-half-their-annual-salary-every-year-to-be-fully-productive-compares-not-using-ai-to-using-paper-and-pencil-for-designing-chips ↩