3 ms·
"Boomv2 achieves 3.92 CoreMark/MHz (on the taped out BOOM), vs 3.71 for the Cortex-A9." - it's a bit of a letdown. I was hoping it'd be closer to x64 performanc
by AllSeeingEye 9y ago
"Boomv2 achieves 3.92 CoreMark/MHz (on the taped out BOOM), vs 3.71 for the Cortex-A9."
- it's a bit of a letdown. I was hoping it'd be closer to x64 performance than to Cortex-A, but it's probably not achievable for RICS-V budget.
- phkahler 9y agoThere is still a lot of low hanging fruit in the boom design. Hopefully Chris will still be contributing some improvement to Boom.
- asb 9y agoI'm sure Chris will drop by to expand this, but it should be noted that figure was for a particular instantiation. There's no reason you can't explore different power/performance/area trade-offs, or invest further development effort to further improve performance.
- _chris_ 9y agoSorry to disappoint, AllSeeingEyes. What I taped out is a fairly modest instantiation of BOOM. I was trying to reduce risk and we had a very small area to play with, so I settled for ~4 CM/MHz. One potential win here would have been to use my TAGE-based predictor, which is an easy >20% IPC improvement on Coremark. Of course, a lot more reworking would be needed to achieve x86-64 clock frequencies.
- microcolonel 9y ago> which is an easy >20% IPC improvement on Coremark. Of course, a lot more reworking would be needed to achieve x86-64 clock frequencies. Maybe these should be expressed as Instructions Per Second (at peak and at the point of diminishing returns?) or something like that, rather than two independent numbers. Higher clock frequency actually seems like a bad thing, all else being equal. It seems to me that throughput ought to trend toward infinity, clock frequency toward zero. ;- )
- _chris_ 9y agoOf course it's important to always keep the "Iron Law" in mind, but it's far easier to compare ideas and talk about things in terms of IPC. For example, if a switch out one branch predictor for another, we're talking about an algorithmic change (implemented in hw) that will have an effect on IPC, and no effect on clock period (assuming we didn't screw something up). This is particularly useful when talking about a processor design, and not a specific processor in particular. As you said, there's a lot of good about slower clock frequencies, so you'll see the same ARM design being deployed at a variety of frequencies. Far easier to talk separately about a design's IPC from its achievable clock frequency (although both are important!).
- AllSeeingEye 9y agoNo worries, Chris, I didn't expect entirely new processor design to play in Intel league :) Really hoping to get my hands on RISC-V-based SBCs next year.
- _chris_ 9y agoMe too!
- mtgx 9y agoNobody is going to try to compete with Intel Core i7 or AMD Ryzen 7 from year one of a new ISA, for the same reason ARM chip makers didn't try to do that for years. It's just too risky and the investments need to be much bigger. Instead you "grow your way up" as ARM chips have done. Now, Cavium's Vulcan-based ThunderX2 seems to be beating Intel's Skylake server chips: https://www.nextplatform.com/2017/11/27/cavium-truly-contender-one-two-arm-server-punch/ https://www.nextplatform.com/2017/11/27/cavium-truly-contend...
- monocasa 9y agoCavium is still no where near intel on single core performance; they just threw way more simpler cores onto a chip.
- mtgx 9y agoI can still see it as a great option for say shared hosting providers. Multiple times as many customers could be supported for the price they pay for Intel chips.
- monocasa 9y agoAbsolutely. Just saying that in the context of performance/cycle/core, they're closer to a bunch of Atoms than a Xeon (which, again, is totally fine for some use cases).