The Efficiency Trap
How Better GPUs Make the Electricity Problem Worse, Not Better
Last week at GTC 2026, Jensen Huang introduced Vera Rubin, NVIDIA’s next-generation AI chip architecture. The headline number: 10x more efficient at inference than the current Blackwell generation. Ten times more AI output per watt of electricity.
If you are a policymaker worried about data centers consuming too much electricity, that sounds like great news. Better efficiency means less power needed, right?
Not quite. In fact, the opposite.
Welcome to the efficiency trap. It has a formal name: Jevons’ paradox. And it is about to play out in electricity markets at a scale that would make the original William Stanley Jevons blush.
A 161-Year-Old Prediction
In 1865, an English economist named William Stanley Jevons published “The Coal Question.” His observation was counterintuitive and, at the time, controversial: when James Watt improved the efficiency of the steam engine, England did not use less coal. It used more. A lot more. The improved engine made coal-powered machinery profitable in applications where it had never been economical before. Factories that could not afford to run on the old engines could suddenly afford the new ones. Total coal consumption exploded.
Jevons’ insight was simple: when you make a resource more efficient to use, you do not reduce demand for the resource. You increase it. Efficiency expands the number of profitable applications, and the expansion overwhelms the per-unit savings.
This is not an obscure academic curiosity. It has played out in every major energy technology transition. Cars got more fuel-efficient; people drove more; total gasoline consumption rose. LED bulbs used a fraction of the electricity of incandescent bulbs; people lit up everything; total lighting electricity consumption barely budged. Air conditioning got cheaper to run; people cooled bigger spaces; total cooling demand soared.
Now it is happening with AI and electricity.
The Math That Matters
I maintain something called the Compute Heat Rate™ (CHR)1, which measures the maximum electricity price an AI data center can profitably sustain. You can read the foundational research here. Think of CHR as a thermometer for how much economic pressure AI demand can exert on the power grid.
Here is what Vera Rubin does to that thermometer.
The current generation of chips, Blackwell, generates up to roughly $74,000 in AI revenue for every megawatt-hour of electricity it consumes when running frontier inference workloads. That is not a typo. For context, an aluminum smelter generates about $60-80 of economic value per MWh. A steel mill, $80-120. AI inference on Blackwell: $74,000.
That gap is why data centers do not want to curtail when electricity prices spike. The electricity cost is a rounding error compared to the revenue the computation generates.
Vera Rubin makes each watt roughly three to five times more productive at inference. Conservatively, the same megawatt-hour of electricity that generated $74,000 of AI revenue on Blackwell chips will generate up to $229,000 on Vera Rubin.
Better chips do not reduce the value of electricity. They increase it. Each watt becomes more economically productive, which makes the operator willing to pay even more to keep it flowing.
The blended Compute Heat Rate across all workload types jumps from roughly $6,350/MWh (about 127 times the gas heat rate benchmark) to approximately $25,700/MWh (about 514 times the gas heat rate) when a critical mass of Vera Rubins hit the grid. The demand-side price tolerance ceiling actually accelerating upward with every chip generation.
Check out the full CHR article on Substack.
Jensen Huang Told You This
At GTC, Huang made a statement that most of the tech press treated as a competitive jab at AMD: “Even if competitors gave their chips away for free, this advantage would still matter more, because better efficiency means more growth and more revenue.”
Read that again from an energy market perspective. He is saying: the bottleneck is no longer silicon, it is electricity and the chip is becoming commoditized.
When scarcity shifts from chips to electricity, the entire economic value of AI computation flows through the electricity input. Every efficiency improvement in the chip makes each unit of electricity more valuable, not less.
This is Jevons’ paradox in Jensen Huang’s words: more tokens per watt means more revenue per watt. More revenue per watt means the operator will pay more for the watt. The efficiency improvement is actually an electricity price accelerant.
And They Will Pay Anything
The chip economics explain why operators want more watts. The sunk cost math explains why they will pay anything to keep them.
A Vera Rubin NVL72 rack is expected to cost somewhere between $5 million and $8.4 million. That rack contains 72 GPUs. Each GPU depreciates at roughly $13 per hour whether it is running or not.
Running that GPU at $500/MWh wholesale electricity, which would be considered an extreme price spike in any historical context, costs about $1.15 per hour.
The depreciation cost is eleven times the electricity cost.
If you are an operator with $8 million in sunk GPU hardware, the rational decision at any electricity price the market can produce is: keep running. An idle GPU burns capital at eleven times the rate that an expensive megawatt-hour does. The operator will absorb any electricity price rather than let the hardware sit idle.
So Who Pays?
Businesses and that share the grid with data centers will share in the costs. When electricity prices rise because supply cannot keep up with demand that refuses to curtail, everyone gets squeezed unless there’s regulatory intervention. Some businesses may curtail their operations.
Residential ratepayers also may pay. Not directly through wholesale price exposure right away in most areas since most households are on fixed or regulated rates. But the capacity charges, transmission upgrades, infrastructure investments and ultimately rate increases due to wholesale price increase may hit eventually.
The Trap
Here is how the efficiency trap works:
A new GPU generation delivers dramatically more AI output per watt. Headlines celebrate the efficiency gain. Policymakers feel reassured.
But the efficiency gain makes AI profitable for workloads that were previously uneconomical. Demand for total compute expands. New data centers get built. Total electricity consumption from AI goes up, not down, because the expansion of demand overwhelms the per-unit savings.
Meanwhile, the new GPUs have even higher revenue per watt, which raises the Compute Heat Rate, which deepens demand-side price inelasticity. The new demand will not want to economically curtail at essentially any price the market can reach.
Then repeat with the next GPU generation. NVIDIA has already announced Rubin Ultra for 2027 and Feynman for 2028.
What Comes Next
GPU efficiency improvements are extraordinary engineering achievements. They will produce enormous economic value. The trap is not the efficiency itself; it is the assumption that efficiency solves the energy problem. It does not. It shifts the problem from “not enough compute” to “not enough electricity for the compute that is now even more profitable.”
If we are not tracking how much demand-side price tolerance is increasing with each GPU generation, we cannot plan for its impact on the grid. That is what the Compute Heat Rate measures: GPU economics into electricity market terms. When the number goes up, the pressure on the grid intensifies, even if per-unit efficiency improved.
Grid capacity models need to account for Jevons’ paradox explicitly. Assuming efficiency gains will moderate demand growth is historically wrong in every energy domain, and especially wrong where the economic returns on compute are growing faster than the efficiency of producing it.
Royal, Hans, The Compute Heat Rate: Quantifying AI-Driven Electricity Price Tolerance
and Its Implications for Wholesale Market Repricing (February 28, 2026).
Available at SSRN: http://dx.doi.org/10.2139/ssrn.6322318