3 ms·
Working from the 10K end-host number in the paper with a radix of 43 and 722 total router nodes, it’s the case that a given switch connects ~15 end-hosts. Fift
by rnxrx 8y ago
Working from the 10K end-host number in the paper with a radix of 43 and 722 total router nodes, it’s the case that a given switch connects ~15 end-hosts. Fifteen downlinks to 43 uplinks seems pretty wasteful in its own right.
The numbers you cite are closer to realistic: a modified Clos fabric consisting of a four-way spine of 8-slot chassis with 36 100G blades connecting a full complement of 288 leaves, each with 4x100G up and 40x10G down (no oversubscription, at least nominally) leaves us with 11.5K hosts connected by a total of 292 switches. Even if the modular switches are 10X the cost of the ToR this is still markedly cheaper than 700+ ToR’s (and this doesn’t include actual power/cooling costs and the opportunity cost of space lost).
This also doesn’t include cabling/transceiver cost differences: 722 * 43 * 2 (62K) vs 288 * 4 * 2 (2.3K). That’s literally an order of magnitude difference.
Even if the design were approached using single-speed connections (ex: 48x10G up divided over 16 spines, 48x10G down locally) there’s still a pretty compelling numerical advantage to Clos: ~210 96x10G ToR’s and 16 16-slot chassis for 10K hosts is 226 devices and ~20K transceivers. If we assume the spines cost 10X the leaves then 370X is still almost half of 722X.
- dgaudet 8y agomore clos advantages: - simplified operations: can service/remove/add ToRs without affecting global routing or the forwarding capacity of the fabric. - tor uplink flexibility: if a rack has higher bandwidth needs then double up the links (4 vs 8 in your example) - tor location mobility: it's a lot easier to manage 4 fiber runs per tor than it is to manage 43 different fiber runs... backhaul to 4 spine blocks vs. a complex web of interconnect spreading all over the datacenter floor. with the fly network it's unlikely you can move a rack once it has been placed, at least not until you're ready to tear down the whole fabric and build something new. so you better get your rack density and layout just right. with tor you're stuck with your spine locations, but everything else can be moved around. the fly advantages for homogeneous supercomputers built and decom'd N years later are clear... but for datacenters which grow and evolve with heterogeneous devices, fly doesn't seem to really hold up well compared to clos.