6 ms·
Core to core latency data on large systems
- jauntywundrkind 3y agoIt'll be interesting to see how CXL shakes out. It might end up being not much more than cross socket access! 150ns to go between sockets is about what we see here & is in the realm of what CXL had been promising. Having a super short lightweight protocol like CXL.mem to talk over such fast fabric has so much killer potential. These graphs are always such a delight to see. It's a network map, of how well connected cores are, and they reveal so many particular advantages and diaadvantages of the greater systems architecture.
- loxias 3y agoI, too, am excited for CXL. Not enough people got to _feel_ the awesome of pmem. I think if more people had, pmem would be in all our laptops, desktops, servers.
- rsaxvc 3y agoBack in the days before Oracle, Sun would sell you a dual socket Opteron desktop and you could add your own FPGA right on the hypertransport in the second socket. Exciting to see that capability becoming more standardized with CXL. Edit: phrasing.
- formerly_proven 3y agoIt's almost poetic to have those mid-1990s Pentiums there, with about 2-3x the inter-socket latency of the current state-of-the-art, 30 years later.
- gpderetta 3y agoVery interesting. Now do bandwidth next!
- bee_rider 3y agoThe NUMA nature of recent* chips has made me wonder if there’s ever going to be a movement to start using message passing libraries (like MPI) on shared memory machines. * actually, not even that recent, Zen planted this hope in my brain.
- nvartolomei 3y agoThread-per-core software architectures are doing this https://penberg.org/papers/tpc-ancs19.pdf https://penberg.org/papers/tpc-ancs19.pdf Real world examples are scylladb and Redpanda, both built on the seastar framework (C++ https://seastar.io/message-passing/ https://seastar.io/message-passing/). And for rust there is glommio https://www.datadoghq.com/blog/engineering/introducing-glommio/ https://www.datadoghq.com/blog/engineering/introducing-glomm...
- RedlineTriad 3y agoThere is also another thread-per-core implementation by ByteDance (TikTok) for Rust called Monoio with benchmarks[0] comparing it to Tokio and Glommio. [0] https://github.com/bytedance/monoio/blob/master/docs/en/benchmark.md https://github.com/bytedance/monoio/blob/master/docs/en/benc...
- sapiogram 3y agoDoes thread per core necessarily imply message passing? I don't see why the two need to be related.
- yencabulator 3y agoThe thread-per-core manifesto has a goal of not sharing data between cores, and thus the communication inside the process becomes message passing, handing off ownership of a chunk of data to the recipient core. This lack of sharing is what enables the performance (no locks etc needed, outside of the message passing). This is a good watch (first half is pure background, second half talks about the motivation): https://www.youtube.com/watch?v=PbgTyCSDPrs https://www.youtube.com/watch?v=PbgTyCSDPrs
- hinkley 3y agoI was misreading these charts for too long. Maybe I still am. Am I seeing that none of these processors implement a toroidal communication path? I thought that was considered basic cluster topology these days so I’m surprised that multi core chips don’t implement it.
- twic 3y agoIf your chip is fabricated on a flat rectangular piece of silicon, that would involve links running from each edge, across the chip, to the other edge, in both orientations. I can imagine that would be very demanding of chip resources, slow, etc. If your chip is fabricated on the surface of a torus, or on a rectangle in highly curved space, then it would be a very natural architecture. But i am not aware of any chips that are.
- hinkley 3y agoI would presume the first couple of layers of silicon would be wires instead of gates. At least at the edges. Top left to top right, Bottom left to bottom right, top left to bottom left, top right to bottom right. The middle of the chip could contain logic.
- deleted 3y ago[deleted]
- nwmcsween 3y agoIf I'm reading this right socket-to-socket latency hasn't really improved much in a long time, why?
- undersuit 3y agoI like the end of the article. >If Pentium could run at 3 GHz and the FSB got a proportional clock speed increase, core to core latency would be just over 20 ns. Ran the test against my closest equivalent. CPU: Intel(R) Celeron(R) G5905T CPU @ 3.30GHz Num cores: 2 Num iterations per samples: 5000 Num samples: 300 1) CAS latency on a single shared cache line 0 1 0 1 25±0 Min latency: 25.3ns ±0.2 cores: (1,0) Max latency: 25.3ns ±0.2 cores: (1,0) Mean latency: 25.3ns Just wish I had a dual socket Pentium for the last 40 years.