3 ms·
Larry (SGI) had lived through IRIX fine grained locking and even SGI's NUMA hardware cache coherency based on Stanford research right? Was his take that the com
by gregw2 1y ago
Larry (SGI) had lived through IRIX fine grained locking and even SGI's NUMA hardware cache coherency based on Stanford research right? Was his take that the complexity wasn't worth it given his experiences at SGI, or that it was just too much for an open source community to tackle without owning the hardware layers?
(And did Maddog (DEC) with a different set of experiences agree?)
- vacuity 1y agoThe trend of multicore and NUMA means that hardware increasingly looks like a traditional network of many separate computers. The natural conclusions of single-core scaled up to, say, 4 cores, shift when there are 8+ cores. Locality becomes crucial; just as you wouldn't split up data-path dependencies across LANs, you shouldn't split them up across NUMA sockets either. Ignoring arguments about locking, message passing, cache management, and whatever, the most pressing argument for multikernels (or at least, far increased per-core state and reduced shared state) is that locality is essential for performance.
- layla5alive 1y agoYup, data movement and contention and coherencey are the things that will increasingly dominate power use as core scaling continues. Exploiting locality is a must for high performance systems. Linux would benefit from a scheduler per CCD (in AMD parlance) approach being a first-class option. CCD pinning is a mechanism to push in this direction today, but partitioning kernel scheduler(s) along hardware boundaries would reduce complexity and overhead for a lot of use cases..