3 ms·
That’s well and good, but the limitation here is PCIe or other Bus bandwidth, and memory bandwidth. PCIe 5.0 gets us to 1Tbps per x16 slot. The top-of-the-line
by virtuallynathan 7y ago
That’s well and good, but the limitation here is PCIe or other Bus bandwidth, and memory bandwidth. PCIe 5.0 gets us to 1Tbps per x16 slot. The top-of-the-line CPUs Max our at 6.4Tbps of memory bandwidth (Power10 when it is released).
This would require a large ASIC or FPGA to be useful, loaded with high-flocked HBM2.
- AdrianB1 7y agoThey are looking at the entire system and they are willing to challenge (improve or replace) everything, including buses like PCIe. This is also an effort not meant for regular servers, but very high end. We already have 128 PCIe lanes or more in a single server, that means 8 Tbps; not in a single connection or slot, but aggregated. Also dual socket CPU's are very popular and more than that is still accessible; multiply the memory bandwidth by the number of sockets. Think about total throughput, not per bus, per socket, per NIC, etc.
- wmf 7y agoThey are looking at the entire system and they are willing to challenge (improve or replace) everything That's also what I thought; just put the NIC inside the processor and connect it to the internal fabric. (This still leaves plenty of software challenges.) But then DARPA says "The hardware solutions must attach to servers via one or more industry-standard interface points, such as I/O buses, multiprocessor interconnection networks, and memory slots, to support the rapid transition of FastNICs technology." Even if your "NIC" uses all 128 lanes of PCIe 5 that's only 4+4 Tbps. If you get rid of serdes and use something like IFOP that's ~600 Gbps per port then you'd still need something like 16 of those links.
- chongli 7y agoYou would also need to replace the operating system and all the software. The traditional operating system driver stack does way too much copying for any of this to work. Perhaps a better model would involve some way of time-sharing or multiplexing direct access to the hardware by user applications. I'm not an operating system expert so take my comments with a grain of salt.
- ComputerGuru 7y agoThe standards get ratified years (literally) before the first implementations ship out, but note that PCIe 6.0 is already slated to provide 4 TB/s in a x16 slot. TBH, I’m kind of surprised x32 didn’t happen in the time between PCIe 3.0 and 4.0 (or maybe it did and I just didn’t hear about it), as there are now “plenty” of enterprise-class chips that have sufficient lanes to make it feasible to saturate such a pipe, although I’m guessing custom silicon already makes sense at that level of specialization where you can do x32 if you want to without waiting on a formalized interface.
- AdrianB1 7y agoNo, it is just 128 GB/s in each direction for a x16 slot.
- ComputerGuru 7y agoWell, yes. I guess it depends on the context but I suppose you’re right since when you’re sending or receiving data that fast, you likely are going point-to-point and not switching/routing it, meaning you’ll end up mostly doing more of one than the other.
- AdrianB1 7y agoI was saying you miscalculated the PCIe bandwidth.
- ComputerGuru 7y agoI don’t know what I was thinking. I completely misread and didn’t sanity check the numbers.