AI & Automation5 min readNetray Engineering Team

AI Datacenter Power and Cooling Planning for GPU Racks

GPU rack power density has outrun what standard air-cooled data center design can handle, and the planning conversation in 2026 is less about whether you need liquid cooling and more about at what rack count you cross the threshold. A single 8x H100 node draws roughly 10 to 10.5 kW, and a fully populated rack of four such nodes plus networking and storage can reach 40 to 45 kW, well past the 15 to 20 kW ceiling where conventional air cooling remains practical. B200-based racks push density further still, and full GB200 NVL72 rack-scale systems can exceed 120 kW per rack, which is not optional-liquid-cooling territory, it is liquid-cooling-only territory.

Rack Density Thresholds by GPU Generation

Air cooling remains viable up to roughly 15 to 20 kW per rack with well-designed hot aisle/cold aisle containment and sufficient computer room air handler capacity, which covers a lightly populated H100 rack or most CPU-only infrastructure. A single fully populated 8x H100 or H200 node draws 10 to 10.5 kW, meaning even 2 nodes per rack (20-21 kW) sits at the edge of air cooling's practical limit, and most enterprise H100/H200 deployments end up at 2 to 3 nodes per rack specifically to stay air-coolable without a facility retrofit. B200 nodes draw 12 to 15 kW each, and any B200 or GB200 deployment beyond a handful of GPUs should be planned for direct liquid cooling from the outset.

  • Air cooling practical limit: roughly 15-20 kW per rack with proper containment
  • 8x H100/H200 node: ~10-10.5 kW, limiting most racks to 2-3 nodes without liquid cooling
  • 8x B200 node: ~12-15 kW, pushing full racks well past air cooling limits
  • GB200 NVL72 rack-scale systems: 120+ kW per rack, liquid cooling is mandatory, not optional

When Liquid Cooling Becomes Necessary, Not Optional

Direct-to-chip liquid cooling, where coolant loops attach directly to cold plates on the GPU and CPU packages, is now the default design for any rack exceeding roughly 30 to 40 kW, and it is required infrastructure for B200 and GB200 deployments at meaningful scale. Retrofitting an existing air-cooled facility for liquid cooling means adding a coolant distribution unit (CDU), redundant piping, leak detection, and often a dedicated facility water loop, which is a multi-month construction project, not a hardware order. Immersion cooling remains a niche alternative for extreme-density deployments but carries higher operational complexity and is far less common in enterprise deployments than direct-to-chip liquid cooling as of 2026.

  • Direct-to-chip liquid cooling is the default above roughly 30-40 kW per rack
  • Requires a coolant distribution unit (CDU), redundant piping, and leak detection integrated into facility design
  • Retrofitting an air-cooled facility for liquid cooling is a construction project measured in months, plan it before hardware arrives
  • Immersion cooling remains niche; direct-to-chip is the practical enterprise default for B200/GB200 density

PUE Targets and What They Mean for Your Bill

Power Usage Effectiveness (PUE) measures total facility power divided by IT equipment power, and it is the single number that determines how much you pay beyond the GPUs themselves. Legacy enterprise data centers commonly run PUE 1.6 to 2.0, meaning 60 to 100 percent overhead on top of compute power for cooling and facility losses. Modern, well-designed air-cooled facilities achieve PUE 1.2 to 1.3. Liquid-cooled facilities, because they eliminate most of the fan and chiller overhead of air cooling, can reach PUE 1.1 to 1.15, which at GPU-cluster power scale translates into real annual savings: a 500 kW IT load at PUE 1.8 versus PUE 1.15 is the difference between roughly 900 kW and 575 kW of total facility draw, a meaningful line item at typical commercial electricity rates.

Power Distribution and Redundancy for GPU Clusters

GPU nodes draw power in sharp, correlated bursts during training steps, which stresses power distribution differently than steady-state enterprise IT load and needs to be sized with headroom rather than average draw. Plan for N+1 redundancy on power distribution units at minimum for production clusters, and budget for the reality that a single rack of B200-class hardware may require dedicated 3-phase feeds beyond what a standard data center rack circuit provides. Coordinate early with facility power capacity and your utility, since adding hundreds of kW of new GPU load to an existing building sometimes requires a utility service upgrade with its own multi-month lead time, independent of the hardware procurement timeline.

How Netray Plans Power and Cooling for On-Prem AI

Netray runs power and cooling planning as part of the same capacity assessment as GPU selection and cluster architecture, because a design that ignores facility constraints is not a real design. We size rack density against your actual site's cooling capability, flag when a deployment crosses the threshold requiring liquid cooling or facility upgrades, and coordinate with your facilities team or colocation provider before hardware is ordered so there is no gap between server delivery and the ability to safely power it on. For regulated manufacturers building out on-premises AI infrastructure inside existing facilities, this planning step is frequently the difference between a smooth deployment and a multi-month delay discovered after the GPUs arrive.

Frequently Asked Questions

At what rack density do I need liquid cooling instead of air cooling?

Air cooling remains practical up to roughly 15 to 20 kW per rack with proper hot aisle/cold aisle containment. Above roughly 30 to 40 kW per rack, direct-to-chip liquid cooling becomes the default design. Since a single 8x H100 or H200 node already draws 10-10.5 kW, most H100/H200 racks stay at 2-3 nodes to remain air-coolable, while B200 and GB200 rack-scale systems, which can exceed 120 kW per rack, require liquid cooling as mandatory infrastructure, not an option.

What is a good PUE target for a GPU data center?

Modern well-designed air-cooled facilities target PUE 1.2 to 1.3. Liquid-cooled facilities can reach PUE 1.1 to 1.15 because they eliminate most fan and chiller overhead. Legacy data centers commonly run PUE 1.6 to 2.0, meaning 60 to 100 percent power overhead beyond the GPUs themselves. At cluster scale, moving from PUE 1.8 to 1.15 can cut total facility power draw by roughly a third for the same IT load.

How much power does an 8x H100 GPU server need?

A fully loaded 8x H100 SXM node draws roughly 10 to 10.5 kW. An equivalent 8x B200 node draws roughly 12 to 15 kW due to the higher per-GPU TDP. At rack scale with networking and storage overhead included, a rack with 2 to 3 such nodes typically reaches 25 to 40 kW, which is at or beyond the practical limit for conventional air cooling.

Key Takeaways

  • 1Rack Density Thresholds by GPU Generation: Air cooling remains viable up to roughly 15 to 20 kW per rack with well-designed hot aisle/cold aisle containment and sufficient computer room air handler capacity, which covers a lightly populated H100 rack or most CPU-only infrastructure. A single fully populated 8x H100 or H200 node draws 10 to 10.5 kW, meaning even 2 nodes per rack (20-21 kW) sits at the edge of air cooling's practical limit, and most enterprise H100/H200 deployments end up at 2 to 3 nodes per rack specifically to stay air-coolable without a facility retrofit.
  • 2When Liquid Cooling Becomes Necessary, Not Optional: Direct-to-chip liquid cooling, where coolant loops attach directly to cold plates on the GPU and CPU packages, is now the default design for any rack exceeding roughly 30 to 40 kW, and it is required infrastructure for B200 and GB200 deployments at meaningful scale. Retrofitting an existing air-cooled facility for liquid cooling means adding a coolant distribution unit (CDU), redundant piping, leak detection, and often a dedicated facility water loop, which is a multi-month construction project, not a hardware order.
  • 3PUE Targets and What They Mean for Your Bill: Power Usage Effectiveness (PUE) measures total facility power divided by IT equipment power, and it is the single number that determines how much you pay beyond the GPUs themselves. Legacy enterprise data centers commonly run PUE 1.6 to 2.0, meaning 60 to 100 percent overhead on top of compute power for cooling and facility losses.

Planning GPU rack density for a new or existing facility? Netray will assess your power and cooling capacity before you commit to hardware that cannot be safely deployed on site.