Every watt a GPU uses turns into heat. How you take that heat away decides which chips you can run, how many fit in one rack, and which data centers can host you.

The short version

Smaller setups can be cooled with air, the way a fan cools a laptop. The newest and densest systems make too much heat for air, so they're cooled with liquid that runs through the equipment.

  • Air cooling works for servers that hold eight GPUs: H100, H200, B200 and B300.
  • Liquid cooling is required for the rack-sized systems, GB200 and GB300.
  • Everything in the new Rubin generation needs liquid, at any size.

What "rack density" means

A rack is the tall cabinet that holds the servers. Rack density is how much power, and so how much heat, is packed into one rack. It's measured in kilowatts (kW).

POWER DRAW, KILOWATTSAir: one 8-GPU H200 server10Air: one 8-GPU B300 server15Liquid: one DGX Rubin NVL8 server24Air: rack of four B300 servers60Liquid: GB200 NVL72 rack132Liquid: GB300 NVL72 rack142Liquid: Vera Rubin NVL72 rack190-230REPORTED, NOT PUBLISHED BY NVIDIA
Maximum or planning figures, in kilowatts. B300 and GB300 values are from NVIDIA. H200 and GB200 values are planning figures. The Rubin NVL8 value is NVIDIA's early figure for DGX Rubin NVL8. The Vera Rubin NVL72 range is industry reporting: NVIDIA has not published a rack power figure.

A rack of four air-cooled GPU servers draws about 40 to 60 kW. A liquid-cooled GB300 rack draws up to 142 kW. That's more than twice the heat in the same floor space.

Air cooling: easier to place, takes more room

Air cooling blows cold air through the servers. It's simple, and many existing data centers can do it.

The limit is how much heat air can carry away. One 8-GPU server draws 10 to 15 kW. Four of them in a rack is 40 to 60 kW, which is already more than many data halls were built for. So an air-cooled cluster gets spread across more racks and takes more floor space.

An example: 50 B300 servers, which is 400 GPUs, draw about 700 to 750 kW. At four servers to a rack, that's 13 racks.

Liquid cooling: more GPUs per rack, fewer places can do it

Liquid carries heat far better than air. In a liquid-cooled system, pipes bring coolant straight to the hottest chips.

That's what lets NVIDIA put 72 GPUs in a single rack. A GB200 rack draws about 132 kW and a GB300 rack up to 142 kW. No amount of air can cool that. The same 400 GPUs that took 13 racks on air fit in six.

The trade-off is the building. A liquid-cooled site needs water piped to the racks and pumping equipment to move it. Fewer data centers have that today, so you have fewer places to choose from.

One thing people miss: a liquid-cooled rack still needs some air cooling. Only the hottest parts are on liquid. NVIDIA's own reference design says "the most power-intensive components, such as GPUs and CPUs, are liquid cooled. Other components are air cooled."

Vera Rubin: NVL8 and NVL72

Rubin is NVIDIA's next generation after Blackwell. With Rubin, air cooling is off the table. NVIDIA has announced it in two sizes.

NVL8 is the smaller one: a server with 8 Rubin GPUs. NVIDIA's early figure for its DGX Rubin NVL8 is about 24 kW per server, and it calls the system "a liquid-cooled AI system." A B300 server with the same number of GPUs draws about 14 to 15 kW and runs on air. So a Rubin server needs roughly 60 to 70 percent more power, and it needs liquid.

NVL72 is the big one: a full rack with 72 Rubin GPUs. It's cooled with warm water, 45°C going in, which lets a building do without the chillers that colder systems need. NVIDIA hasn't published how much power the rack draws. Industry reporting puts it between 190 and 230 kW. Treat any single number with caution, and ask the operator for the rated figure in writing.

What this means in practice: a data center that hosts B300 servers on air today can't host Rubin without adding liquid cooling.

Here's the difference in numbers, before networking, storage and the building's own overhead:

  • 10 servers, 80 GPUs: about 240 kW on Rubin NVL8, against 140 to 150 kW on B300.
  • 50 servers, 400 GPUs: about 1.2 MW on Rubin NVL8, against 700 to 750 kW on B300.
  • 100 servers, 800 GPUs: about 2.4 MW on Rubin NVL8, against 1.4 to 1.5 MW on B300.

Which one do you need?

  • Running models for users, or fine-tuning them: air-cooled 8-GPU servers on H200, B200 or B300. They're the easiest to place.
  • Training very large models: liquid-cooled racks, GB200 or GB300.
  • Anything on Rubin: liquid.

Five questions to ask a data center

  • How many kW per rack can you support today, and with which kind of cooling?
  • Is water already piped to the racks?
  • Who supplies the pumping equipment for liquid cooling, you or us?
  • Is the site built to Tier III standards? NVIDIA's reference design for GB300 clusters calls for Tier 3, which means any part of the site can be serviced without shutting it down.
  • What's the minimum term? On the capacity we broker, minimum terms run 36 to 60 months.

Where XIRR fits

XIRR Advisors brokers two things: Tier III colocation space and power for hardware you own, and reserved GPU compute capacity (H100/H200/B200/B300/GB200/GB300/Vera Rubin) from neocloud operators. We do not sell GPUs or hardware.

When you send a colocation requirement, put the cooling method and the kW per rack next to the market, total power and timing. It's the first thing an operator asks.

For the chips themselves, see From Hopper to Rubin: A Plain-English Guide to NVIDIA Data Center GPUs.

Sources

NVIDIA product pages for DGX B300, GB300 NVL72, DGX Rubin NVL8 and Vera Rubin NVL72. NVIDIA Enterprise Reference Architecture for GB300 NVL72. NVIDIA DGX SuperPOD reference architecture for DGX GB300. H200, B200 and GB200 values are planning figures. The Vera Rubin NVL72 rack range is from industry reporting, not from NVIDIA.

Liquid coolingRack densityColocationVera RubinNVIDIA
Tell Us What You're Sourcing

Share your requirements. We'll canvas the market.

Tell us what you need (for colocation space: market, power in kW or MW, cooling and timing; for reserved GPU capacity: region, GPU type, cluster size and timing) and we'll canvas colocation and neocloud operators in parallel. Shortlist in 48 hours. We don't sell hardware. We broker colocation space and reserved GPU compute capacity.

Earlier conversations get better terms. When you engage early, we have time to negotiate with vendors before you need to commit. We're paid only when a deal closes, and the fee structure is agreed in writing before any introduction.