7 ms·
A Deep Dive into AMD’s Rome Epyc Architecture
- mmrezaie 7y agoThere must be a simulation for this kind of architectures to see what is the best combination of size and components while making it practical! I wonder if anyone knows something like that? A tool to minmax these choices and estimate if this can be done with resources they have got.
- willis936 7y agoI’m pretty sure they roll their own. I would be surprised if simulation tools were not the most well guarded secrets of these companies.
- chrisseaton 7y agoThey have many levels of simulations for their processor designs yes but they’re proprietary.
- Symmetry 7y agoOh yes. That's been the dominant approach at least since Computer Architecture: A Quantitative Approach came out in '89. The search space is pretty complicated, though, given the physical interactions as well as the logical ones.
- sroussey 7y agoSounds like a lot of parameters for some ML setup
- repolfx 7y agoYeah but they're often trying to predict where the software industry will go years in advance. You can see that in the disagreements between AMD and Intel on AVX-512: there's a chicken and egg situation where it's not always clear what's right to optimise for, as it depends on changing workloads and software platforms. For instance the big chip companies were caught out by the need for low precision maths for AI inferencing. That all came out of Google.
- ajross 7y agoTools like that are a core part of the design process. You write that software along with the choice of parametrization of the design. It's not an off the shelf thing. But yes, that's how it works. It's also important to note that decisions like this are hugely workload-specific. There's no single best processor for all applications. In extreme examples: almost every transistor on a vector SIMD unit is wasted when trying to optimize for a client Javascript benchmark; streaming symmetric encryption gets no benefit from L3 cache (which is like half the chip these days!); etc...
- mmrezaie 7y agoDo people like Jim Keller are the experts in finding the right balance? Is that why he and his team are so important?
- zazagura 7y agoMaybe the time has come for applications specific CPU variations? One optimized for node.js tasks, one for databases, ...
- tempguy9999 7y ago> one for databases as a DB guy, there's no 'one task' for DBs . The only thing I can think that is nearly characteristic of DBs I've worked on is that they're IO bound. That's possibly true of most things except Floating point and graphics.
- snaky 7y agoWhile that's true, you can achieve impressive results offloading particular parts of particular DBMS code to the specialized CPUs, https://www.ibm.com/support/knowledgecenter/en/SSEPEK_11.0.0/perf/src/tpc/db2z_ibmziip.html https://www.ibm.com/support/knowledgecenter/en/SSEPEK_11.0.0...
- pjc50 7y agoARM got a Javascript oriented instruction: http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc.dui0801g/hko1477562192868.html http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc.... At least two different tasks, audio and graphics, have special application specific processors. Networking and crypto are also often offloaded.
- brosenlof 7y agohttp://gem5.org http://gem5.org
- positr0n 7y agohttp://gem5.org/Main_Page http://gem5.org/Main_Page Is an open source CPU simulator. I used it in undergrad to run benchmarks with different cache sizes and cache coherence strategies to see which were more effective. I'm sure Intel and AMD have much more advanced simulation tools though. Most likely multiple, or at least multiple levels of granularity (so you could do stuff like, simulate these potential branch predictor designs at a gate level, and then turn around and simulate the entire CPU at a higher level of abstraction.)
- deepnotderp 7y agoSo yes, plenty of simulators exist, many internal ones as well as gem5* but automated search space exploration, say like with mcmc, isn't in widespread use yet iirc, although plenty of academic papers have explored the topic. *which I swear every company has their own version of
- mjw1007 7y agoUp until around 2012, realworldtech.com and anandtech.com used to publish rather more detailed descriptions of the microarchitecture inside each core. Is anyone publishing things like that these days? I mean pages like these: https://www.realworldtech.com/haswell-cpu/4/ https://www.realworldtech.com/haswell-cpu/4/ https://www.anandtech.com/show/6355/intels-haswell-architecture/8 https://www.anandtech.com/show/6355/intels-haswell-architect... (I noticed that Agner Fog's chapter on Ryzen is conspicuously missing a "Literature" section.)
- throwaway2048 7y agoThe servethehome review of Rome is a pretty detailed look at the architecture. https://www.servethehome.com/amd-epyc-7002-series-rome-delivers-a-knockout/ https://www.servethehome.com/amd-epyc-7002-series-rome-deliv...
- rrss 7y agohttps://www.anandtech.com/show/14525/amd-zen-2-microarchitecture-analysis-ryzen-3000-and-epyc-rome/ https://www.anandtech.com/show/14525/amd-zen-2-microarchitec...
- mjw1007 7y agoThat, and the servethehome review, seems to be basically putting the presentation slides into words. A few years ago they seemed to have additional sources of information (they'd talk about things like instruction-to-port assignments and penalties for moving data between integer and FP domains).
- rrss 7y agoMaybe one of the write ups of AMD's presentation at hot chips next week will have what you want. To be honest, though, I don't see a substantial difference between the haswell article you linked and the Zen 2 article, provided you are willing to look past the AMD slides. The haswell article is also just "putting the presentation slides into words," just from IDF 2012 instead of AMD Tech Day 2019, and apparently the author felt a need to do the block diagrams themselves. (Also, FWIW, Agner's manual does have a literature section for Ryzen, it is just not numbered for some reason).
- ramshanker 7y agoMy gut feeling is that Intel also lays out / develops IO block and cores seperately. It's just that they are all put on single silicon.
- chx 7y agoBut separate silicon is what gives AMD an almost insurmountable cost advantage. They can bin each chiplet separately, their yields are much higher because each die is smaller and the cherry on top is the different, cheaper process for the I/O die.
- shaklee3 7y agoThis didn't really seem like a deep dive compared to the anandtech article. I was hoping for some memory bandwidth benchmarks, since this should be the first chip that has 8 channels without caveats (looking at you power 9). It's also not clear if it's 16 channels with 2S, but I suspect not. Edit: the picture from AMD in this review makes me think it can hit 16 memory channels with the two socket version. Does anyone know if this is true?
- wtallis 7y ago> the picture from AMD in this review makes me think it can hit 16 memory channels with the two socket version. Does anyone know if this is true? Yes, if the motherboard provides all the necessary slots. The inter-socket communication is achieved by re-purposing CPU pins used for PCIe, not pins used for DRAM. Each CPU has the full 8 DRAM channels of its own.
- LargoLasskhyfv 7y agoSomewhere at 50 to 60% down in the article: "There are a total of eight DDR4 memory controllers on this hub chip, the same number in total that were on the Naples complex; both support one DIMM per channel and have two channels per controller, but Rome memory runs slightly faster – 3.2 GHz versus 2.67 GHz – and therefore with all memory slots filled, yields a maximum of 410 GB/sec of peak memory bandwidth per socket. That’s 45 percent higher than the Cascade Lake Xeon SP processor, which has six memory controllers for a total of 282 GB/sec of memory bandwidth running at 2.93 GHz and 21 percent higher than the 340 GB/sec that Naples turns in running that 2.67 GHz DRAM. (Those are ratings for two-socket servers.)"
- thinkersilver 7y agoThe poster is holding a line of bash to the standard of code and is illustrating that readability should be the goal and a way of bringing bash commands to a standard of readability for something like a PR. Readability is really there to show _intent_ I would say though that if you are bringing this to the code standards of today then this should really be wrapped up in some kind of unit test (https://github.com/sstephenson/bats https://github.com/sstephenson/bats )for it to pass the PR. That would make the code a bit more maintainable and can be integrated as a stage in your CI/CD pipeline. If we do that then the intent would be clarified by the input and the expected output of the test. Then then the code would at least be maintainable and the readability problem becomes less of an issue when it comes to technical debt. I've done this plenty of times with my teams and its certainly helped.
- insulanus 7y agoAre you replying to this thread? https://news.ycombinator.com/item?id=20724679 https://news.ycombinator.com/item?id=20724679
- thinkersilver 7y agoYes I was. I've posted the comment to the correct story now. I don't how that happened.
- MayeulC 7y ago> “We like features that improve both power and performance,” Clark elaborated. “Being on the right path more often is important because the worst use of power is executing instructions that you are just going to throw away. We are not throwing work away after we figure out dynamically that we were wrong to do it. This definitely burns more power on the front end, but it pays dividends on the back end.” Every documentation I've seen is quite light on the branch prediction improvements. Going by the slides, they improved is accuracy by 1/3; I'd be curious to know how. Side note: if your superscalar is big enough (yeah, those registers use power), couldn't you just get rid of branch prediction at no performance cost (doing something else while waiting for the data)? My only grudge against Zen (as a consumer) is that the AM4 socket is intended for both APUs and CPUs. While this is a good thing, I have a couple utterly useless video outputs on my motherboard. I would have liked AMD to include some display driver circuitry on every chip. Maybe in the I/O die, if they use such a thing in all of their designs going forward? I mean, I would be quite content with using software rendering when I need to drive a screen, or even spare a bit of memory bandwidth and CPU cycles to drive an extra display from my desktop's graphics card.
- piadodjanho 7y ago> Every documentation I've seen is quite light on the branch prediction improvements. In one of the pictures in the article, it says the new architecture uses the TAGE Branche predictor. This is likely based on the work of Andre Seznec. There are many articles on the implementation (but they can be difficult to understand if you are not already familiar with his work). I've implemented the bare bone predictor on a computer architecture course, you can see an abridged version of my presentation slides here[1]. Note this only describes the bare bone predictor, in recent work Andre Senzec added a Loop predictor and a Statistical Correlation Unit to increase the accuracy. There are some work using TAGE with perceptions in the Statistical Correlation unit. [1] https://docs.google.com/presentation/d/1aUrwD-ENYPB7pMrCoYmEcLamuwyk_WfrggBT_3vOZBs/edit?usp=sharing https://docs.google.com/presentation/d/1aUrwD-ENYPB7pMrCoYmE...
- MayeulC 7y agoThank you, I hadn't realized those branch predictors were actually documented, and thought that they were referring to internal names. It is nice to see research being applied to new mainstream chips relatively quickly. In complement to your slides, there is a short overview here [1] (this is actually the first search result).