AI Data Center Power and Cooling Calculator for On-Prem GPU Clusters
This free AI data center power and cooling calculator sizes the electrical and thermal footprint of an on-prem GPU cluster, and it is built for IT directors, facilities engineers, and plant leaders evaluating private AI. Enter GPU count and class, expected utilization, host overhead, facility PUE, and your electricity rate, and the tool returns GPU load, total IT load, facility load, required cooling tonnage, and annual energy cost. Most manufacturers discover that the constraint is not GPU budget but the 30 to 60 kilowatts of usable power and matching cooling their existing server room simply cannot deliver.
Your numbers
Total accelerators in the planned cluster. An 8-GPU server is one node, so 64 GPUs is eight nodes.
Board power at sustained load. Liquid-cooled flagship parts can exceed 1,000W each.
Training runs sit near 90-100%; mixed inference clusters usually average 50-75%.
CPU, memory, NVMe, NICs, and fans as a percentage of GPU power. Dense AI nodes commonly run 30-45%.
Power usage effectiveness: total facility power divided by IT power. Lower is better.
Include demand charges and delivery, not just the energy rate on the bill.
Your results
Planning estimates only. Actual load depends on workload mix, firmware power caps, and ambient conditions. Have a licensed electrical and mechanical engineer validate any design before procurement.
Get your full power and cooling sizing report
We will email you a personalized load, cooling, and energy cost breakdown with rack density guidance and redundancy options, and a Netray infrastructure specialist will follow up with a phased build plan.
No spam. Your results stay private. Unsubscribe anytime.
How the load calculation works
The model starts from board power. Sixty-four 700W accelerators at 70% sustained utilization draw about 31.4 kW. Host servers, NVMe, InfiniBand or Ethernet fabric, and fans add roughly 35% in dense AI nodes, taking total IT load to about 42.3 kW. Multiplying by a facility PUE of 1.3 gives about 55 kW at the building meter. Cooling is derived from the IT load only, since that is the heat actually rejected into the room: 42.3 kW converts to roughly 12 tons of refrigeration using the standard 3.412 BTU per watt-hour and 12,000 BTU per ton conversions. At $0.12 per kWh running continuously, that cluster costs roughly $58,000 per year to power.
Benchmarks and design rules of thumb
The defaults come from real enterprise GPU deployments rather than vendor spec sheets, where sustained draw is consistently higher than marketing literature suggests. The biggest planning error we see is assuming a rack that comfortably held traditional servers can host an AI node without electrical and mechanical work. A single dense node can consume more power than an entire legacy rack, and the branch circuits, power distribution units, and floor cooling behind it were never designed for that density. Treat the figures below as design constraints to check against your facility before hardware arrives, not as targets to optimize toward later.
- Traditional enterprise racks are provisioned for 5-10 kW; a single 8-GPU AI node draws 6-11 kW on its own.
- Air cooling becomes impractical above roughly 30-40 kW per rack, which pushes dense clusters toward rear-door heat exchangers or direct liquid cooling.
- A good enterprise data center runs PUE near 1.3; an unimproved on-site server room is often 1.6 or worse.
- Size electrical service to peak board power, not average utilization, because synchronized training steps produce sharp coincident spikes.
Reading your results before you buy hardware
Compare facility load against your available spare capacity at the panel, not the nameplate capacity of the building. If the calculator returns 55 kW and your server room has 20 kW of headroom, the real project is electrical distribution and cooling, and that work typically takes longer to permit and install than the GPUs take to arrive. Cooling tonnage tells you whether existing CRAC units can cope; add 25-30% for N+1 redundancy if the cluster will run production inference. Annual energy cost belongs in your total cost of ownership model next to hardware amortization, because over four years power and cooling commonly reach 20-30% of the capital cost.
How Netray helps you plan an on-prem AI build
Netray designs private AI infrastructure for aerospace, defense, and electronics manufacturers where ITAR and CMMC obligations rule out public cloud inference. We translate a workload description into a right-sized cluster: model footprint, concurrency targets, GPU class, power and cooling envelope, and a phased procurement plan that starts small enough to prove value. Because we also build the ERP and application layer on top - SyteLine, LN, M3, and ServiceMax integrations - we size for the workloads you will actually run rather than a generic benchmark. Most engagements begin with a facility and workload assessment that produces a costed reference architecture.
Frequently Asked Questions
Why does the calculator use IT load rather than facility load for cooling?
Cooling equipment has to remove the heat generated inside the room, and that heat equals the electrical energy consumed by the IT equipment itself. Facility load includes the energy the cooling plant consumes to do that work, so using it would double-count. The IT load figure of about 42 kW converts to roughly 12 tons. Add redundancy on top of that number, typically 25-30% for an N+1 design supporting production inference.
Do I need liquid cooling for a small AI cluster?
Usually not below about 30 kW per rack. Eight-GPU nodes with 700W accelerators draw roughly 8-11 kW each, so two to three nodes per rack stays inside what good air containment plus rear-door heat exchangers can handle. Beyond that, or with 1,000W-class parts, direct-to-chip liquid cooling becomes the practical option. The decision is driven by per-rack density and room airflow design, not by total cluster size.
How accurate is a PUE assumption for an on-site server room?
Less accurate than most planners expect, which is why the tool offers discrete choices instead of a free-form field. Purpose-built data centers reliably achieve 1.2-1.4. Converted office server rooms with mixed-age CRAC units, poor containment, and no economizer frequently measure 1.7-2.0 once you meter honestly. If you have never measured, assume 1.5 for planning and commission a proper energy audit before signing an electrical contract.
Get a right-sized on-prem AI reference architecture with power, cooling, and phased procurement mapped to your facility.
Related Tools
GPU Cluster Utilization Calculator
Turn GPU capital, amortization, and operating cost into an effective cost per productive GPU hour, and find the utilization threshold where owning beats renting.
On-Prem AIAI Inference Latency Calculator
Estimate decode throughput, time to first token, and end-to-end response time for a self-hosted model from GPU memory bandwidth, parameter count, and quantization.
On-Prem AIOn-Prem AI Security Hardening Checklist
A practical control checklist for securing self-hosted language models, covering model provenance, network isolation, data governance, host hardening, and audit readiness.
Go Deeper
Enterprise GPU Cluster Planning for AI Workloads
Plan an enterprise GPU cluster for AI workloads: H100 vs L40S sizing, networking, power, cooling, and cost models for on-prem LLM inference and training.
On-Prem LLM Deployment Architecture: Reference Guide
Reference architecture for on-prem LLM deployment: inference servers, GPU sizing, RAG pipelines, and security zones for regulated manufacturers.