Alchemist Server is live. Browse our AI infrastructure line-up or talk to our team about a custom build.Talk to our team

Skip to main content

White Papers

Why We Size From the Facility, Not the Catalog

By Alchemist Server

A short statement of how we approach cluster design, and why the first conversation is usually about electrical distribution rather than GPUs.

Most infrastructure vendors start a conversation with a configuration. We start it with a question about the room, and this is a short explanation of why.

Configurations are easy, deployments are not

Producing a bill of materials that meets a GPU count is straightforward. Producing one that will run at rated performance in a specific building, on a specific electrical supply, with a specific cooling capacity, is the actual work. The second one is what the customer needs and the first one is what usually gets sold.

The failure mode we see repeatedly

A cluster is specified on capability, purchased, delivered, and then discovered to exceed what the facility can power or cool. The remedies at that stage are all expensive: spread across more racks than planned, run at reduced density, retrofit cooling under time pressure, or return hardware.

Every one of those outcomes was avoidable at design time with three numbers that take an afternoon to establish.

What this means in practice

Our first conversation is usually about available power per rack, cooling capacity and floor loading. It is a less exciting conversation than one about GPU counts, and it is the one that determines whether the project succeeds.

Occasionally it produces an answer the customer does not want: the facility supports less than the plan assumed. That is a considerably better conversation to have before a purchase order than after one.

Related reading

Planning an AI infrastructure build?

Tell us the workload and the power envelope you have to work within. We will come back with a configuration, a lead time and a written quote.