AI runs on chips that do simple arithmetic in parallel, packed so densely that power, cooling and water have become the real limits.
- GPUs won because of parallelism. A model's work is enormous grids of numbers multiplied at once, which is what graphics chips happen to be built for.
- Moving the numbers is often the bottleneck. Fast memory sits right next to the chip, which is also why VRAM decides what runs on your own machine.
- Density made cooling a core problem. Racks draw so much power that air can no longer carry the heat away, and liquid cooling has become standard.
- One request is small; the total is not. A single prompt uses little electricity, but the combined buildout is a large new load on the grid, concentrated in a few regions.
- These figures age. Each number on the page carries its source and date, and the page marks which ones were not re-checked in the latest cycle.
How an AI is built shows what a model is. This page descends through what it is made of: sand, gold, water, and roughly a nation's worth of electricity. Five levels down. Mind the heat.
Why this chip and not your laptop's

A CPU is a brilliant generalist: a few powerful cores doing complicated things one after another. A neural network needs the opposite — one simple thing (multiply, add) done billions of times simultaneously, because a model's "thinking" is just enormous grids of numbers being multiplied together. GPUs, born for videogame pixels, happen to be exactly that: thousands of small cores in parallel. That accident of history is why one company's green silicon became the picks-and-shovels of the entire gold rush.
The die's partner is memory bandwidth — how fast weights can be fed to those cores. Modern accelerators stack high-bandwidth memory physically next to the die because the bottleneck is usually not computing the numbers; it is carrying them. When the syllabus tells you VRAM decides what runs on your machine, this level is why.
Memory decides what runs, not just chip speed. When a model will not fit or crawls on your machine, how fast weights can be fed to the chip is usually the reason.
Density is the whole story
Accelerators are ganged eight-plus to a board, boards stacked into racks, racks laced together with interconnect fabric fast enough that thousands of chips can behave like one giant machine — a training run is exactly that: one calculation smeared across a warehouse for weeks. The consequence is heat density with no precedent in computing: the IEA measured AI rack power density rising 11-fold from 2020 to 2025, projecting a further fourfold jump by 2027 — at which point a single refrigerator-sized rack draws power equivalent to about 65 households. Air cannot carry that heat away anymore; liquid cooling — plates and pipes against the silicon — became standard equipment, not exotica.
Treat heat as the constraint. Packing more chips together is what pushed cooling from air to liquid, so the physical limits of AI show up in the plumbing as much as in the silicon.
Where the pour happens

Gather the racks and you get the hall: a large data centre can consume as much electricity as 100,000 households — the IEA’s own phrasing — and the agency states that the largest currently under construction could consume as much as 2 million households. Training happens here (the weeks-long pour The Assembly describes), and then inference happens here forever after — every chat, every image, every agent, each sip small on its own: the IEA's 2026 indicative estimates, counting GPU electricity only, put a request to a large text model around 0.31 Wh, a moderate agentic task around 1.14 Wh, and the same agentic task with reasoning turned on near 50 Wh, while a disclosed median short text prompt measured 0.24 Wh and 0.26 ml of water (Google technical paper, 2025). Small sips; unimaginable frequency.IEA, Energy and AI, April 2025, read at source 17 Sep 2026: “A large data centre can consume as much electricity as 100 000 households. The largest currently under construction could consume as much as 2 million households.” · IEA, Key Questions on Energy and AI, April 2026, read at source 17 Sep 2026: Figure 2.1 shows 0.31 Wh for a large mixture-of-experts text model, 1.14 Wh for an agentic task and 50 Wh for an agentic task with reasoning; its note: “Agentic estimates represent a moderate task profile of four to six sequential LLM calls”. This page previously labelled the 1.1 Wh bar a standard assistant request.
The part everyone forgets is thermodynamics

Every watt in becomes heat out; the only questions are how efficiently (the industry metric PUE — total facility power divided by computing power — with modern halls pushing toward 1.1) and what carries it away. Often: water, through evaporative cooling — which is why hyperscaler water disclosures rose sharply through 2024–25 and why closed-loop liquid designs, which cut direct water use dramatically at higher capital cost, are spreading. Cooling is now its own line on the world's bill: Gartner forecasts cooling electricity alone climbing 22.6% in 2026 to 195 TWh.Gartner forecast as reported by Data Centre Magazine, 15 Jun 2026, read at source 23 Sep 2026: “Data centre cooling infrastructure electricity consumption is forecast to jump by 22.6% this year, reaching a total of 195TWh.” IEA, Key Questions on Energy and AI, April 2026, read at source 23 Sep 2026: best-in-class hyperscale facilities “now achieve power usage effectiveness (PUE) ratios of approximately 1.1 to 1.2”.
Read at source on 17 Sep 2026 in the IEA's own report PDFs: the IEA's projection of ~945 TWh of datacentre electricity demand by 2030 (April 2025 Energy and AI report; the agency has since updated its central projection to about 950 TWh, from 485 TWh in 2025), and the comparison that a large datacentre can consume as much electricity as 100,000 households, with the largest under construction reaching 2 million. The same day, the April 2026 IEA report was also read for rack power density, the 2025 growth rates and the per-request estimates. The Gartner cooling and server forecasts, the Lawrence Berkeley range and the regional shares were not re-checked in that cycle; on 23 Sep 2026 they were read at source (Gartner through trade coverage, because gartner.com refuses automated access) and are quoted under the numbers above.IEA, Energy and AI, April 2025, read at source 17 Sep 2026: “Data centre electricity consumption is set to more than double to around 945 TWh by 2030.” · IEA, Key Questions on Energy and AI, April 2026, read at source 17 Sep 2026: “Our updated projections see electricity consumption from data centres roughly doubling from 485 TWh in 2025 to 950 TWh in 2030, accounting for around 3% of global electricity demand by that date.”
Ask what carries the heat away. Every watt a data centre uses becomes heat, and where evaporative cooling carries it away, the cost is paid partly in water.
The numbers, dated and sourced
IEA, Key Questions on Energy and AI, April 2026, read at source 23 Sep 2026: “The global electricity demand of data centres - the critical infrastructure for training and running AI models - grew by 17% in 2025”; “Electricity consumption from AI-focused data centres grew even faster, surging 50% in 2025.”; “roughly doubling from 485 TWh in 2025 to 950 TWh in 2030, accounting for around 3% of global electricity demand”. Gartner forecast as reported by Network World, 7 Jul 2026, read at source 23 Sep 2026 (gartner.com refuses automated access): “565 terawatt hours (TWh) in 2026, up 26% from 447 TWh in 2025”; “Gartner estimates AI-optimized server adoption will account for 31% of data center power consumption in 2026, and that by 2027 their power consumption will surpass that of conventional servers.” Lawrence Berkeley National Laboratory, 15 Jan 2025, read at source 23 Sep 2026: “Data centers consumed about 4.4% of total U.S. electricity in 2023 and, depending upon how much the rest of the economy grows, are expected to consume between 6.7 and 12% of total U.S. electricity by 2028.” Carbon Brief, 15 Sep 2025, read at source 23 Sep 2026: “In the US state of Virginia, these facilities already consume 26% of electricity”; in Ireland “around 21% of the nation’s electricity is used for data centres”. First-hand: until 23 Sep 2026 the first line also called the 50% “16× the pace of overall electricity demand”; no source for that ratio could be found, so it was replaced with the IEA’s own comparison.
Two honest framings, held together: an individual query costs less electricity than boiling water for tea — and the buildout in aggregate is the largest new load the grid has met in a generation, concentrated in a handful of regions whose bills and build-outs are already reshaped by it. Both sentences are true. Any page that gives you only one of them is selling something.
Keep both scales in view. Use the per-request figure to judge your own use and the aggregate figure to judge the buildout; a source that offers only one is leaving something out.
Ascend
You have seen the body. How the model it runs is made is in how an AI is built; what it costs the internet is in the quarterly record; and whether it should live in these halls or on your desk is weighed, without sides, in Local or Cloud.