Alchemist Server is live. Browse our AI infrastructure line-up or talk to our team about a custom build.Talk to our team

Skip to main content

Technical Guides

Power, Not GPUs, Is the Constraint on Your Next Cluster

By Alchemist Server

Most cluster plans are sized on GPU count and then discover the room cannot power them. Working the other way round produces a design that actually runs.

A conversation about a new training cluster almost always starts with a GPU count. Sixty-four B300s, someone says, and the discussion moves on to fabric topology and storage throughput. The number that decides whether any of it is buildable, the kilowatts the room can actually deliver, comes up somewhere near the end, usually when a facilities manager is finally in the room.

The arithmetic that gets skipped

An eight-way B300 node draws roughly 14kW under sustained load. Eight of those is 112kW before a single switch, PDU or storage shelf is counted. A conventional enterprise rack is provisioned for somewhere between 7kW and 15kW. Even a well-specified colocation cabinet at 30kW takes two nodes and no more.

So sixty-four GPUs is eight nodes is roughly 120kW is four to sixteen racks depending on the facility, not the single rack the original conversation imagined. The GPU count was never the hard part.

Air stops working before the rack is full

Cooling fails earlier than power does. Above roughly 40kW in a rack, air cooling becomes impractical regardless of how much power is available: the volume of air required exceeds what can be moved through the cabinet, and the temperature rise across the rack pushes inlet temperatures on the rear equipment out of specification.

This is why liquid cooling appears in these designs at a specific density rather than as a premium option. It is not about efficiency at that point. It is about whether the hardware can run at its rated clocks at all.

Working backwards instead

Start with the facility. Ask what a rack can draw, what the cooling capacity is, and what the floor is rated to carry. Those three numbers produce a maximum node count per rack, and the node count produces the GPU count. The answer is often smaller than the one the project started with, which is uncomfortable and considerably cheaper than discovering it after the purchase order.

Where the workload genuinely justifies more density than the facility allows, liquid cooling changes the ceiling. A 300kW in-row CDU supports a row of eight-way nodes that air could not. That is a facility project with a lead time of its own, which is exactly why it belongs at the start of the conversation rather than the end.

What to measure before committing

Per-rack power available, in kilowatts, at the PDU rather than on the building schematic. Cooling capacity and whether facility water is available. Floor loading in kilograms per square metre. Three-phase availability and the breaker headroom above steady-state draw, because these systems have transient peaks well above their typical figure.

Every one of those is knowable before anything is ordered. None of them is knowable after.

Related reading

Planning an AI infrastructure build?

Tell us the workload and the power envelope you have to work within. We will come back with a configuration, a lead time and a written quote.