4 ms·
There seems to be at least one incorrect statement in the article. "The point Intel is making in the chart above is that it has mastered NUMA scaling, which is
by ThenAsNow 10y ago
There seems to be at least one incorrect statement in the article.
"The point Intel is making in the chart above is that it has mastered NUMA scaling, which is no surprise since the company has been at it for decades."
SGI was one the pioneers of ccNUMA (Origin 2000?) in the late 90s. That would constitute decades.
Around the same time, Intel delivered the ASCI Red supercomputer using P6 processors with a snooping coherency scheme.
I don't think Intel had significant NUMA experience (at least in any commercial products) prior to AMD K8 (2003), and their first NUMA product would have been IA64 (E8870 chipset?) or the Nehalem microarchitecture chips. That is not quite a decade, let alone decades.
- creshal 10y agoI was under the impression that Itanium had it a while longer, but you're right. Itanium got it even later (2010) thanks to tape-out delays.
- ryao 10y agoThe SGI Altix 3000 had ccNUMA in 2003, but that was using SGI's ccNUMA technology to connect Intel's microprocessors rather than anything Intel developed. I believe that other UNIX vendors that adopted Intel's chips used their own ccNUMA technology to link them together too. Anyway, the article's author likely made the mistake of thinking that multisocket implies NUMA. It is an easy mistake to make for those who do not know the details of how systems are implemented.
- ThenAsNow 10y agoSGI did it with MIPS processors in the 90s before they used Intel chips. For example: http://dl.acm.org/citation.cfm?doid=264107.264206 http://dl.acm.org/citation.cfm?doid=264107.264206 This was the system in which the directory logic was formally verified: http://www.sgidepot.co.uk/origin/compcon97_dv.pdf http://www.sgidepot.co.uk/origin/compcon97_dv.pdf I agree the author probably made a simple mistake, but it is misleading to people at the "enthusiast" level who then parrot things like this as flamewar fodder. A lot of respect is owed to computer engineers at corporations that are defunct or a fraction of their former greatness (e.g., SGI, DEC, HP). As a further example, Bob Colwell's book, "The Pentium Chronicles", is a great read and teaches some powerful lessons about how to successfully do engineering in the kind of large organizations that can scale production of a design. But if you take it at face value, the book makes no meaningful attempt to inform the reader that other companies had done Out-of-Order well before P6. It leaves a naive reader with the idea Intel pioneered OOO as well, which is no way true. I guess I'm just (maybe overly) sensitive to sloppiness and omission leaving an inaccurate impression of the history, especially when it's so easy these days to fact-check.
- setpatchaddress 10y agoMy recollection is that "The Pentium Chronicles" said explicitly that OoO had been done, but not for x86 CPUs, and that there was a significant question at the time as to whether it was actually viable for x86.
- gpderetta 10y agoPossibly the article is worded imprecisely. NUMA by itself is not the desirable property. What you want is multisocket scaling, NUMA is just a way to get there. Certainly for 2 and 4 sockets Intel uncore (i.e. the connecting fabric) is state of the art and has been for a while (but not 2 decades).
- ThenAsNow 10y agoI agree it's a nit, and your interpretation is probably on the right track, but it's an inaccurate statement by the original author as-written. Also agree that the crux of the point is multisocket scaling. Though with enough sockets (if memory serves, in mid-2000s, the rough consensus was 4), I don't think computer architects think there is any way to scale well without a NUMA approach of some kind, so it's not just an academic point. With today's memory bandwidth and "memory transaction volume", perhaps even 2 sockets wouldn't scale well in a uniform memory access configuration.