Waqar Uddin

The Rack Density Race: Inside America's AI Data Center Buildout

July 23, 2026 (3w ago)23 views

In February 2024, xAI leased an abandoned Electrolux factory shell in Memphis. By June, 100,000 liquid-cooled Nvidia H100 GPUs were training Grok inside it, wired into a single RDMA fabric. One hundred and twenty-two days from empty warehouse to the largest AI training cluster running anywhere in the world at the time.

That timeline is the story. Not the GPU count, the speed. A construction schedule that used to run two to three years for a hyperscale facility got compressed into four months, and the thing that made it possible was not clever project management. It was a wholesale change in what a data center rack physically is.

Disclosure: I work at Jazz as Principal Evangelist Cloud & AI. This article is written as personal industry analysis. Views are my own.

From 5kW to 350kW in two decades

Rack power density sat still for most of the internet era, then broke into a sprint the moment GPU training clusters became the workload that mattered.

  1. 2000 – 2010
    2 – 5kW per rack
    Air-cooled, raised-floor CRAC systems, the internet-era default
  2. 2010 – 2020
    5 – 15kW per rack
    Blade servers and hot-aisle containment stretch air cooling further
  3. 2020 – 2023
    15 – 30kW per rack
    Rear-door heat exchangers supplement air as cloud workloads intensify
  4. 2023 – 2024
    30 – 100kW per rack
    Direct-to-chip liquid cooling becomes mandatory for GPU-dense training clusters
  5. 2025 – 2026
    140 – 250kW per rack
    Liquid cooling is now the default for Blackwell-class GPU racks, not the exception

Physics is the reason the line bends this hard around 2023. Air cooling a rack past roughly 30kW needs airflow north of 2,000 to 3,000 CFM, which means fan noise in the 70 to 80 decibel range, exponentially rising pressure drop, and components running hot enough to shorten their own service life. Past that point, air is not a design choice anymore. It is a constraint you have already lost to. Direct-to-chip liquid cooling, cold plates bolted straight onto the GPU package, removing 70 to 80 percent of the heat before air ever has to touch it, is now simply what a serious AI rack looks like.

Two different bets on what to do with that headroom

Once liquid cooling made 100kW-plus racks viable, US builders split into two distinct strategies, and both are visible in facilities running today.

Bet one: make each rack as dense as physically possible, and sell density as the product. This is the neocloud and colocation play.

Bet two: keep individual racks at a merely aggressive density, and multiply the rack count into the hundreds of thousands. This is the hyperscaler and frontier-lab play, and it is the one producing the eye-catching capacity numbers.

The density race does not actually save money

The counterintuitive part of all this is the economics. Facility construction cost per kilowatt rises sharply with density: roughly $2,300 to $3,500 per kW for traditional air-cooled infrastructure, versus $5,000 to $7,700 per kW once you are building for the 100 to 140kW tier, and $7,500 to $12,500 per kW at the 200kW-plus end where immersion cooling enters the picture. A traditional 10-megawatt facility runs $23 to $35 million to build. An equivalent ultra-high-density facility runs $75 to $125 million.

Higher density racks post better Power Usage Effectiveness, typically 1.2 to 1.35 versus 1.8 to 2.0 for traditional air-cooled facilities, and still cost more to run. One widely cited ten-year total cost of ownership comparison puts traditional 10kW racks at roughly $1,500 per kW per year against $2,350 per kW per year for ultra-high-density racks, a 57 percent premium despite the efficiency gain. The actual justification for building this dense is not cost. It is speed to market, land-constrained sites, and locking in GPU allocation ahead of a competitor.

That last point is the one that explains why xAI, Meta, Amazon, and Microsoft are all racing at once rather than waiting for the economics to soften. The GPUs are the scarce input. The building is a solved problem you throw money at to keep pace with a chip allocation, not the other way around.

Where the rack tops out

Industry engineers increasingly treat 350 to 400kW as close to a hard ceiling for a single rack, not a milestone on the way to something higher. A 415V circuit at its 400A maximum delivers roughly 280kW in practice, fault current risk rises exponentially past that, and at the flow rates required to cool a 400kW rack, hot-swapping a server mid-operation stops being a realistic maintenance procedure. Nvidia's own roadmap has GPU thermal design power climbing from 700W today toward 1,500 to 2,000W by 2027 to 2028, and the expected response is not a 500kW rack. It is a shift toward pod-scale systems spanning five to ten racks and, at the extreme end, entire facilities engineered as a single computer, which is close to a literal description of what Prometheus and Hyperion already are.

Why this matters from where I sit

None of the six builds above are hypothetical vendor pitches. They are operating facilities or funded construction, and the density figures translate directly into what it would take to host a serious training or high-throughput inference workload anywhere else in the world. A 100kW-per-rack deployment, the current baseline for an H100-class cluster, needs $5,000 to $7,700 per kW in facility capital before a single liquid-cooling retrofit is counted, plus a coolant distribution architecture that essentially no data center in Pakistan was built to carry.

That is not a reason for pessimism about Garaj's or any local operator's roadmap. It is a reason to be precise about which layer of this stack is realistically ours to build. Hosting a frontier-scale training run at Colossus or Prometheus density is not a five-year target for any facility on the ground here today, and treating it as one invites the wrong investment. What is a realistic target, and closer to what Pakistan's own compute-layer policy is already aimed at, is inference-tier and mid-density hosting: the 15 to 30kW air-cooled tier that already exists, extending toward the 30 to 100kW tier with a defined, financed liquid-cooling retrofit path, sized for regulated in-country workloads rather than frontier model training. Knowing exactly how far 350kW is from where local infrastructure sits today is what makes that a credible plan instead of a slogan.

Where this fits

Part of an ongoing series on cloud and AI infrastructure from a practitioner's perspective. Related reading: Pakistan's Domestic Routing Mandate Closes the Network Layer, PISF Closes the Hosting Layer in Pakistan's Sovereignty Stack, Picking Battles in the AI Stack.