White Papers
Choosing a Fabric Topology for Your Node Count
When a single tier suffices, when a fat tree is required, and the cost of over-designing for growth that does not arrive.
Fabric topology should follow node count. Over-designing for growth that does not arrive is a common and expensive habit.
Single tier
Up to the port count of a single switch, a single tier is the right answer and there is no reason to build anything more elaborate. One switch, every node one hop from every other, no oversubscription to reason about and nothing to tune. A 64-port switch covers a starter cluster comfortably.
The temptation is to build a two-tier fabric anyway because the cluster might grow. That buys latency, cost and operational complexity today against expansion that in many cases never happens, and if it does happen the fabric can be rebuilt then with better information.
Two tier, fat tree
Beyond a single switch, a fat tree is the standard answer. The decision that matters is the oversubscription ratio between tiers. Non-blocking is correct for training, where collective traffic is synchronised and any oversubscription becomes a bottleneck precisely when every node is communicating at once. Inference fleets tolerate oversubscription well, because their traffic is neither synchronised nor collective-heavy.
Plan the expansion, do not build it
The useful middle path is to design so expansion is additive rather than a rebuild: leave the spine ports, plan the cable routes, and confirm the rack positions and power for the second tier. Then build the single tier. The expansion becomes a scheduled addition rather than a redesign, and nothing is paid for until it is needed.
Optics deserve specific attention in either case. Specify per link rather than per switch, measure the actual cable path including vertical runs, and order spares. A fabric is limited to the port count of its working optics.