AI Data Centers Explained: From Power Plants to ChatGPT
Summary of the video “The Entire AI Data Center Explained — From Electricity to ChatGPT” by Leo Cui, Ph.D., CFA.
AI data centers are factories converting electricity into tokens (words). A single query travels through fiber optics to a building drawing nuclear-plant-scale power, where 10,000 GPU cores process your question in parallel, cooled by liquid systems, and networked by light-speed interconnects. The $725 billion annual spend by four companies is reshaping electricity grids, memory markets, and semiconductor economics—but the revenue side remains a bet that consumer AI fees grow faster than infrastructure costs spiral.
Why AI Suddenly Costs Billions
Search retrieves, AI generates
Google Search is a librarian—it indexes the internet once and retrieves pre-existing answers (costing a fraction of a cent per query). ChatGPT is a writer—it composes answers from scratch, token by token, every single time, costing 10-100 times more compute per query.
Scaling laws made intelligence purchasable
Around 2020, OpenAI discovered that making models bigger, feeding them more data, and spending more compute reliably improves performance in a predictable, nearly linear relationship. This converted AI from a research problem into a capital expenditure race—outspend competitors, get smarter models.
Software became heavy industry
For 50 years, Silicon Valley software escaped the physical world—zero marginal cost, infinite copies, no factories. AI reversed this: frontier software now requires concrete buildings, megawatts of power, and water cooling. Every additional smart answer requires physical machines.
The Two-Second Journey
Your question's complete path
Your question leaves your phone as radio waves, converts to light pulses in fiber optic cable traveling at two-thirds light speed across hundreds of miles, arrives at a data center, gets tokenized (chopped into math chunks), processed in two phases (prefill: reading your entire prompt in parallel; decode: generating one word at a time), and streams back token by token.
Prefill vs. decode: two different phases
Prefill is a burst of parallel computation where the model reads your entire prompt at once and builds a working memory (KV cache) of the conversation. Decode is the assembly line: the model generates one word at a time, running the entire network for each token. The pause before the first word appears is prefill; the streaming text is decode.
Batching: the invisible profit lever
The data center batches hundreds of users' questions onto the same chip simultaneously, like a delivery driver grouping orders on one route. This batching is the difference between a query costing cents versus dollars.
Power: The Root Bottleneck
Power density explosion
Traditional server racks draw 5-10 kilowatts. Nvidia's flagship AI rack draws 120 kilowatts. The next-generation Vera Rubin racks approach 600 kilowatts—a 60-100x jump in power density in under a decade. This is the root cause of nearly everything that follows.
Grid cannot keep up
The US electrical grid was built for 1-2% annual demand growth. AI is asking for tens of gigawatts right now. Data centers currently consume 4-5% of US electricity; credible projections put it at 9-17% by 2030. The interconnection queue to plug into the grid is 4-5 years long, with 410 gigawatts of projects waiting (87% are data centers).
Behind-the-meter power: the workaround
Rather than wait for the grid, hyperscalers generate electricity on-site or next-door (behind-the-meter), bypassing the public queue. This decision by thousands of companies simultaneously ignited an entire forgotten sector: industrial power companies.
Power Sources: The New Energy Race
Nuclear resurrection
Microsoft signed a 20-year deal with Constellation Energy to restart Three Mile Island's undamaged reactor (835 MW) by late 2027, backed by a $1.6 billion investment and $1 billion federal loan. Nuclear is the only carbon-free 24/7 power source. Constellation (22 GW fleet) is now a ~$90 billion company; peer Vistra (37 GW fleet) rode the same wave.
Small modular reactors: lottery tickets
Oklo (backed by Sam Altman, 14+ GW signed pipeline) and NuScale (only SMR certified by US regulators) are pre-revenue companies whose first commercial electrons arrive around 2030 at earliest. Both stocks are down 65-78% from late 2025 peaks. These are options on the 2030s, not businesses today.
Fuel cells: fastest near-term power
Bloom Energy's solid oxide fuel cells convert natural gas to electricity chemically and can be deployed behind-the-meter in 12-18 months (vs. 4 years for grid, 10 years for nuclear). Stock nearly quadrupled in 2025, then doubled again in H1 2026. Bloom announced $7.65 billion in data center contracts in a single 9-day stretch, though backlog-to-binding-contract ratios face scrutiny.
Gas turbines: the realistic near-term winner
GE Vernova (GE's power spin-off) makes giant gas turbines—the only power source realistically buildable at scale before 2030. Turbine slots are sold out through end of decade; backlog is ~$163 billion. In Q1 2026 alone, they booked $2.4 billion in data center electrification orders (more than all of 2025). Stock is now ~$300 billion.
Power distribution: the invisible shovels
Between the substation and chips sit switch gear, busways, and uninterruptible power supplies (UPS batteries). Three companies dominate: Vertiv, Schneider Electric, and Eaton. Eaton's electric backlog grew 48% YoY; Vertiv's more than doubled to $15 billion and joined S&P 500 in March. These sell shovels to every miner regardless of who wins.
Cooling: From Air to Liquid
The heat problem
A single flagship AI chip dissipates over 1,000 watts of heat (like a space heater on a postcard). Stack 72 chips plus memory in one rack: 120 kilowatts of heat. Computation is heat—every watt of electricity becomes thermal energy. Cooling isn't a support function; it's half the job.
Air cooling hit its limit
For 30 years, data centers used air conditioning (cold air up through floor, hot air out the back). Air worked fine up to 30-50 kW per rack. At 120 kW, you'd need hurricane-force winds through servers. The industry is undergoing its biggest plumbing change in history: switching from air to liquid.
Liquid cooling technology ladder
Rear door heat exchangers (water-cooled radiator on rack back) are a transitional patch. Direct-to-chip cooling (metal plate with liquid channels directly on each chip, connected to a coolant distribution unit) is the 2026 mainstream. Nvidia's flagship racks require it. Extreme immersion cooling (dunking entire servers in non-conductive fluid) is the frontier.
PUE: the efficiency metric
Power Usage Effectiveness (PUE) = total power in / power reaching the computer. Perfect score is 1.0. Old data centers run ~2.0 (one watt of cooling per watt of computing). Modern liquid-cooled facilities hit 1.1. That efficiency gap times a gigawatt times electricity prices is real money.
Cooling market: picks and shovels
Vertiv is the market leader, selling both power gear and liquid cooling (revenue up 28% YoY at 20% margin). Eaton paid ~$9.5 billion for Boyd Thermal; Schneider Electric bought Motivair—both electrics signaling their vision of the future. Specialists: nVent (loops, enclosures), CoolIT (cool plates). Liquid cooling market: ~$5 billion in 2025, forecast $15-27 billion by early 2030s.
The Chip: Nvidia's Empire
The machine: Nvidia GB200 NVL72
72 GPUs wired together so tightly they behave as one giant computer. Weighs ~1.5 tons, draws 120 kW, costs ~$3 million. Inside: CPU (head chef, runs OS), 10,000 GPU cores (line cooks doing parallel math), HBM (high-bandwidth memory countertop), SSD (pantry), NIC (waiter between kitchens), power/motherboard (plumbing).
Nvidia's dominance: the numbers
Nvidia finished FY2026 with $215.9 billion revenue (up 65%), of which ~$194 billion was data center. Controls 80-86% of AI accelerator market. Gross margin ~75% (Apple, the most admired hardware company, runs ~46%). Became first $5 trillion company last October. ~7 cents of every S&P 500 dollar is Nvidia.
CUDA: the 20-year moat
In 2006, Nvidia spent billions building a programming platform so scientists could use gaming chips for general math. For a decade, it looked like an expensive hobby. Then deep learning arrived, and every AI researcher learned CUDA because it was the only mature option. 20 years later: millions of developers, thousands of specialized libraries, every framework optimized for CUDA first. When AMD ships a better chip, customers compare chips plus the cost of retraining their entire engineering org and rewriting code with a $100M training run on the line.
The bear case: hyperscalers building their own chips
~40% of Nvidia's revenue comes from just four customers (Microsoft, Amazon, Google, Meta), and all four are building their own chips to replace Nvidia. This is the primary existential risk to Nvidia's dominance.
AMD: the competitive threat
AMD's Instinct GPUs are genuinely competitive on inference—more memory per chip, 25-40% better tokens per dollar by some estimates. Problem: software. ROCm (AMD's CUDA alternative) now hits 90-95% of Nvidia's performance on standard workloads. But 90% as good with more friction is a hard pitch when training costs $100M. AMD holds 5-7% of market; watchable but distant second.
Broadcom: the quiet assassin
Hyperscalers cannot design frontier AI chips alone (takes a decade of specialized IP). They hire Broadcom, which co-designs Google's TPU, Meta's training chip, and reportedly OpenAI's. Broadcom books 60%+ of custom chip market. AI revenue up 106% last quarter; $73 billion backlog. Management sees line of sight to $100 billion AI revenue in 2027. Broadcom wins either way: Nvidia sells finished kitchens; Broadcom helps the biggest chains build their own and takes a cut.
System integrators: the assembly line
Super Micro integrates full liquid-cooled racks faster than anyone (revenue up 123% last quarter, 90%+ AI, 6-10% gross margin). Dell took $64 billion in AI server orders with $43 billion backlog (segment margins under 9%). HPE writes supercomputing and sovereign AI deals. Beneath the brands: Taiwanese ODMs (Foxconn ~40% of world's AI racks, Quanta, Wiwynn, Celestica). Nvidia captures 75% margin on silicon; the company screwing it together keeps 6-10%. Profit lives in scarcity; assembly is not scarce.
Networking: The Nervous System
The memory bandwidth bottleneck
During decode (one word at a time), GPU math cores are often not the bottleneck. For every token, the chip must pull the model's parameters and the conversation's KV cache from memory into its cores. The math is fast; the fetching is slow. Modern inference is memory bandwidth bound—line cooks are lightning, but the countertop cannot feed them ingredients fast enough.
Why chips must talk constantly
A frontier model has 1+ trillion parameters requiring terabytes of ultra-fast memory. The biggest GPU carries a few hundred gigabytes. The model gets sliced across thousands of chips. To produce every token, those chips must exchange intermediate results constantly at unimaginable speed. A network 10% slower can idle billions of dollars of silicon.
Bandwidth and latency: the twin demands
Bandwidth = how much data moves per second (width of the conveyor belt). Latency = delay for one handoff (how long a single pass takes). AI needs both everywhere at once. Networking is 40-60% of spending per dollar of GPU.
Inside the rack: NVLink proprietary web
Nvidia's proprietary NVLink makes 72 GPUs inside one rack behave as a single machine at extreme speed.
Rack-to-rack: InfiniBand vs. Ethernet war
For years, serious AI clusters ran on InfiniBand (specialized ultra-low latency networking, acquired by Nvidia in 2019 with Mellanox). Premium performance, premium price, one vendor. By early 2026, ~two-thirds of new AI cluster networking is Ethernet (open universal standard, historically slower, backed by everyone except Nvidia). Open standards, given time, usually win.
Broadcom's switch silicon dominance
Broadcom's Tomahawk chips are the merchant silicon inside most high-end Ethernet switches. Arista Networks builds the switches themselves (best-in-class boxes and software hyperscalers standardize on). Arista: ~6% gross margin, guiding to $11.5 billion this year, with risk that two biggest customers are 40% of revenue.
The optical transceiver: light conversion at scale
Copper wires carry high speeds only a few meters before signals degrade. Between racks, everything converts to light through glass fiber via optical transceivers (thumb-sized electrical-to-light translators). A single large AI cluster consumes hundreds of thousands at hundreds of thousands of dollars each, replaced every upgrade cycle. It's a razor blade for data centers. Players: Coherent (market leader), Lumentum, Innnolight, Fabrinet (contract manufacturer), Corning (fiber), Amphenol (connectors).
Signal integrity at extreme speeds
Astera Labs makes retimer chips that clean up electrical signals degrading over mere inches of circuit board at this speed. Boring, invisible, 76% gross margins. Revenue up 93% last quarter. When data moves this fast, even the space between two chips becomes a market.
Memory: The Skyscraper Solution
HBM: high bandwidth memory
Instead of laying memory chips flat on a board a few inches from the processor, HBM stacks them vertically (8-12 stories high), drills thousands of microscopic elevator shafts through silicon, and glues the tower directly next to the GPU on the same package. Result: 5-6x the bandwidth of conventional memory at 5-6x the price. Memory is now one of the biggest cost components inside every AI chip.
The three HBM makers: trillion-dollar milestone
Only three companies on Earth can make HBM: SK Hynix (Korean, ~60% market share, out-executed Samsung by betting on HBM years early and shipping first), Samsung (largest memory overall, embarrassingly late), Micron (American, went from afterthought to selling out entire 2026 HBM capacity in advance). In May 2026, all three crossed $1 trillion market value; combined over $4 trillion (~16x a decade ago).
HBM psychology shift: from commodity to luxury
Memory used to be the most brutal commodity business in tech (boom, bust, bankrupt, repeat). HBM changed the psychology: it sells out more than a year in advance, allocated like a scarce resource, priced like a luxury good. Open question: does that price survive when all three giants finish capacity expansions simultaneously? Memory cycles have broken hearts before.
Storage: The Inference Inflection
Storage hierarchy: speed vs. cost
Cache on chip (fastest, priciest) → HBM beside chip → DRAM on motherboard → SSDs for hot data → spinning hard drives for cold bulk data (unbeatable per terabyte). AI turned out to be a data hoarder: training datasets, model checkpoints saved every few hours, and every conversation/log/generated image retained forever.
Hard drives resurrected by AI
Hard drives were declared dead a decade ago. AI turned out to be an endless data producer (training checkpoints, inference logs, generated outputs). Seagate CEO calls it 'the inference inflection.' Hard drives are a literal duopoly: Seagate and Western Digital. After a decade of decline, both sold out entire production into 2027. Western Digital now ships 89% of revenue to cloud customers; was one of S&P 500 top performers. Consumer hard drive prices jumped 50% because AI ate the supply.
Flash storage players
Kioxia and Solidigm are the flash drive makers also sold out into 2027.
Software: The Invisible Moat
The stack: bottom to top
Linux (free OS running every server), Kubernetes (open-source foreman scheduling work across thousands of machines), serving layer, then models. Two stories matter most to investors.
CUDA: the 20-year trap (reprise)
In 2006, Nvidia built a programming platform for scientists to use gaming chips for general math. It looked like an expensive hobby for a decade. Deep learning arrived; every AI researcher learned CUDA because it was the only mature option. 20 years later: millions of developers, thousands of libraries, every framework optimized for CUDA first. When AMD ships a better chip, customers compare chips plus the cost of retraining their entire org and rewriting code with a $100M training run on the line. That's why 70% gross margins survive competition. The moat is not silicon; it's the muscle memory of a million engineers.
Serving engines: the cost collapse layer
Raw inference on a trillion-parameter model would be super expensive. Economics only work because of serving engines (vLLM, Nvidia's TensorRT) doing three tricks: batching (grouping hundreds of users' questions through the chip at once), caching (reusing KV working memory instead of recomputing conversation from scratch for every word), quantization (rounding model numbers to lower precision, like shipping a compressed photo). Together: 3-10x cost reduction from software alone. When OpenAI or Anthropic cuts API price 80% in a year, it's mostly this layer, not new chips. Token manufacturing costs are collapsing on a curve.
RAG: retrieval augmented generation
The model is a brilliant writer with fixed education. RAG hands it your company's documents at question time via vector database (search engine finding text by meaning, not keywords). Players: Pinecone, Databricks, Snowflake. The librarian and writer working together—that's the enterprise AI pitch in one sentence.
The Money: Where It Goes and Whether It Returns
The demand side: three tiers of miners
Hyperscalers (Microsoft, Amazon, Google, Meta, Oracle) spending combined $725 billion this year. Neoclouds (specialized GPU landlords: CoreWeave, Nebius, Lambda, Crusoe). AI labs (OpenAI ~$850 billion valuation, Anthropic ~$965 billion on ~$47 billion annualized revenue, XAI folded into SpaceX). Both major labs filed for IPO in June; private market priced them as two of the most valuable companies on Earth.
The cost breakdown: $35 billion per gigawatt campus
Bernstein estimate: one gigawatt of AI data center costs ~$35 billion to build. Breakdown: ~39% chips (single biggest line, why Nvidia's gross profit alone is ~30% of total industry cost). Chips depreciate fast—most operators write them off over 4-6 years. Skeptics argue even that flatters accounting since a 5-year-old GPU competes against chips 10x better. Every year of depreciation life added or removed swings billions in reported profit.
The circular money flow
Nvidia invests billions directly into CoreWeave, Nebius, OpenAI. Those companies use the money to buy Nvidia chips (revenue for Nvidia). Hyperscalers sign enormous contracts with neoclouds (Meta committed ~$21 billion to CoreWeave, ~$27 billion to Nebius), moving data center spending off their balance sheet into someone else's debt. AI labs sign compute deals with hyperscalers (OpenAI with Microsoft/Oracle, Anthropic with Amazon/Google), paid partly with money those hyperscalers invested into them. This is not fraud, not new—vendor financing built railroads and telephone networks. The loop has exactly one opening where fresh money enters: end users and enterprises (you paying $20/month, companies paying for APIs/copilots). That's the only exit that isn't recycled capital. Today, that corner is the smallest number on the board.
The revenue vs. spend gap
OpenAI + Anthropic combined annualized run rates now top $70 billion (fastest revenue ramp in software history). But revenue they'll actually book this calendar year is a fraction of that, set against $725 billion annual spend. OpenAI still expected to lose ~$14 billion this year. The entire $7 trillion machine is the bet that a small corner grows faster than the big circle spins.
The Verdict
Two seconds, $700 billion a year
Your thumb hits send. The question becomes light in glass fiber crossing state lines, arriving at a building drawing the power of a city (from restarted nuclear plants, sold-out turbines, fuel cells behind the meter). Chopped into tokens, fed into a $3 million rack assembled in Taiwan. Ten thousand line cooks fetch a trillion parameters from memory skyscrapers while liquid coolant carries away the heat of eighty space heaters and light-speed interconnects let ten thousand chips think one thought. Software batches you with a thousand strangers; the answer streams back token by token. Two seconds, $700 billion a year, the largest infrastructure project our species has ever attempted. Whether that's the most important investment in history or the most expensive, honestly, nobody on Earth knows yet.
Notable quotes
Scaling law turned intelligence into something you could purchase. — Leo Cui
Software became heavy industry, measured in megawatts, and cooled with water. — Leo Cui
The largest infrastructure build out in human history. — Jensen Huang (Nvidia CEO)