Jevons Paradox in AI Compute Costs: Price the Bottleneck Behind GPUs
Price Jevons two layers down. H100 rents are cyclical and already crashed once. The supply that genuinely cannot bend on a three-year horizon sits in HBM memory, CoWoS packaging, and grid power.
In May 2026, an H100 GPU spot rental hit $2.39 an hour, the highest print in months, while Nvidia (NVDA) was sitting on a 3.6 million unit Blackwell B200 backlog. The binding constraint that month was not GPU silicon. It was high-bandwidth memory.
The same week, AI-data-center vacancy in the United States tightened to roughly 1 to 1.6 percent, with about 92 percent of under-construction capacity pre-committed. That cluster of facts is the right way into Jevons Paradox AI compute costs. The GPU price line is loud and cyclical. The bottlenecks two layers behind it bend much more slowly, and that is where a durable supply story lives.
Why the GPU rental price is the wrong gauge
The popular telling is that efficient AI uses more compute, so chip rents climb forever. The H100 tape disagrees more than most pitches admit.
Eugene Cheah at Latent Space tracked H100 spot rents collapsing from north of $4.70 an hour in 2023 to roughly $1 to $2 an hour by mid-2024. That was well below the $2.85 break-even his model derived from Nvidia’s own four-year, $4 per GPU-hour underwriting. Carmen Li at Medium traced the same descent.
Introl reported a roughly 64 percent peak-to-trough drop by December 2025. By that point IntuitionLabs was cataloguing $1.49 to $6.98 an hour across 15 plus cloud providers.
Then the cycle turned again. Silicon Data’s H100 index flagged a roughly 10 percent surge into the spring of 2026, and a follow-up Silicon Data history reconstructed the full 2023 to 2025 path. SemiAnalysis described the rebound as a one-year rental price index lifting off the floor.
The lesson is that H100 rents are a cyclical futures contract on training capacity, not a Moore’s Law trend. Read across two years and the price is sharply lower than at launch. Read across six months in 2026 and it just jumped 40 percent. The same chart supports both stories, which is why building an investment thesis on the GPU rental line alone is brittle.
Jevons Paradox, defined for the AI cycle
Jevons Paradox is the observation that, when a resource gets cheaper to use per unit, total consumption can rise faster than the per-unit cost falls. The economy ends up spending more on the resource, not less.
For AI compute, the per-unit metric is dollars per token, or dollars per training FLOP. Both fell hard. Andreessen Horowitz catalogued a roughly 280 to 1,000 times collapse in token cost over two to three years, in a write-up they labelled LLMflation. Enterprise AI bills still rose around 320 percent in the same window.
Deloitte’s 2026 TMT predictions argue that AI’s next phase “demands more computational power, not less,” with inference on track to be about two-thirds of all compute in 2026. That is the Jevons rebound, working at the bill level rather than the chip-hour level. Jevons can be perfectly real and an H100 rental price can still drop in any given quarter. The two are not in contradiction.
The Baumol inversion: why the share can shrink
There is a sharper, less comfortable version of this story. On the episode anchoring this piece, Chicago Booth’s Alex Imas argued that the better AI gets, the smaller its share of GDP may become.
The framework is William Baumol’s cost disease, inverted. As AI’s productivity rises, the price of what AI does falls faster than its quantity scales. Economic value migrates to the parts of the economy that do not automate well: energy permitting, regulated services, healthcare judgement, human supervision. AI Frontiers walked through the same disagreement between Phil Trammell and Anton Korinek’s formal Baumol-AI growth models and the headline GDP-share forecasts.
If the Baumol inversion is even half right, the durable capex inflation is not in GPU rents. It sits in the bottlenecks that cannot be relieved by another fab tape-out: high-bandwidth memory, advanced packaging, grid interconnects, gas turbines. That is the picks-and-shovels layer beneath the picks-and-shovels layer.
A worked example: the bottleneck stack behind one B200
To make the abstraction concrete, here is how supply elasticity stacks up behind a single Blackwell B200 GPU as of mid-2026. The shorter the elasticity window, the more pricing power that layer holds during a Jevons rebound.
| Layer | Representative names | Supply-elasticity window | Why it matters for AI compute costs |
|---|---|---|---|
| GPU silicon | Nvidia (NVDA), AMD (AMD) | 12 to 18 months | Cyclical, two up and two down. Rerates but does not stay tight. |
| HBM memory | Micron (MU), SK Hynix | 24 to 36 months | The binding constraint flagged in May 2026. |
| Advanced packaging | TSMC (TSM) | 24 to 36 months | CoWoS capacity gates the Blackwell ramp. |
| Custom accelerators | Broadcom (AVGO) | 18 to 24 months | Hyperscaler ASICs widen the demand pool, not the supply. |
| Data-center power | GE Vernova (GEV), Vertiv (VRT), Eaton (ETN) | 36 to 60 months | Permits, turbines, switchgear lead times. |
The further down the table you go, the less a Jevons rebound in demand can be absorbed by a quick capacity add. That is where pricing power lives across cycles, not just within one. Deloitte’s outlook cites a McKinsey estimate of 130 to 240 gigawatts of additional data-center capacity needed by 2030, which is a power-grid problem long before it is a chip problem.
Common misconceptions
The first misconception is that H100 prices only go up. They have not. A serious version of the Jevons argument has to engage with the 2024 crash and the fact that, even after the 2026 rebound, hourly rents are well below the 2023 launch range.
The second is that DeepSeek-style efficiency gains mean less compute spend in aggregate. The Jevons rebound has been visible in the customer bill, not in the unit price, for two full years.
The third is that the hyperscalers themselves have priced in a Jevons-driven scarcity rent on GPUs. They have done the opposite. The big hyperscalers stretched server-and-network depreciation to roughly six years, a move SemiAnalysis estimated at about 18 billion dollars a year of book savings. That is a bet that the useful life is longer, not that the per-chip-hour rent will compound.
The fourth is that the bottleneck is the chip. The May 2026 tape says the bottleneck is memory and packaging, with power moving up fast. Silicon Data also flags that ACIE-tier demand of roughly 37 billion dollars now rivals the big-four hyperscalers’ roughly 38 billion dollars, broadening the demand pool into sovereigns and enterprises and confirming the rebound is structural rather than narrowly hyperscaler-driven.
What this does not tell you
This piece does not predict the next H100 spot price. The 2024 collapse and the 2026 rebound prove the silicon line is too cyclical to forecast on a quarter-ahead basis. Nothing here tells you whether MU or TSM is priced correctly today.
The bottleneck framework only argues that the supply-elasticity windows on memory, packaging, and power are longer than the GPU’s, which should show up as steadier pricing power across cycles, not necessarily as the best entry price today.
[NEEDS RESEARCH: a dated cross-section of HBM contract prices versus commodity DRAM at multiple points in 2024 to 2026 would let a reader test the bottleneck claim directly. The podcast that anchored this piece does not provide it, and the cited GPU rental sources stop at the silicon layer.]
[NEEDS RESEARCH: a clean comparison of hyperscaler depreciation-schedule changes against actual realized useful life would let a reader judge whether the six-year stretch is conservative or aggressive. The Imas episode does not address it.]
FAQ
Is Jevons Paradox settled economics, or a stretch when applied to GPUs?
The original 19th-century coal observation is uncontroversial. Applying it to a specific input like GPU silicon is newer, and the H100 cycle shows the per-unit price can fall sharply while the rebound shows up at the customer bill level. Treat the framework as well grounded for aggregate compute spend and as a working hypothesis at the chip level.
If H100 rents already crashed once, why expect another tight cycle?
Demand has broadened. Silicon Data reports ACIE buyers nearly matching the big-four hyperscalers in dollar terms, and Blackwell allocation plus the HBM shortage are the proximate triggers. The crash was not the end of the cycle. It was one full turn of it.
Does the Baumol inversion mean AI is a bad investment theme?
No. It means the value may not sit where the cleanest pitch suggests. If AI’s price falls faster than its volume scales, the durable margin sits in the inputs that cannot keep up, which is the bottleneck stack, not the model layer.
What is the cleanest single number to watch?
HBM contract pricing and CoWoS booking lead times. Both are upstream of GPU rents and slower to adjust. When either tightens, the Jevons rebound is real. When either loosens, the cycle is turning.
How quickly can a power-constrained data-center market loosen?
On a three-to-five year window, not less. Grid interconnect queues, turbine OEM order books at GEV, and switchgear lead times at ETN and VRT all run multi-year. That is why power sits at the bottom of the elasticity table.
Disclaimer. This article is for educational and informational purposes only. It is not investment advice, a recommendation to buy or sell any security, and does not take into account any reader’s specific situation, objectives, or constraints.
Past positioning and historical price data are illustrative. Markets, regulation, and supply conditions in chip, memory, packaging, and power markets can change quickly. Verify any number against the cited primary source before acting on it.