3 ms·
It's a cool project but I do wish these open source processor initiatives targetted more realistic design points. In particular there's often a desire to push
by gchadwick 3y ago
It's a cool project but I do wish these open source processor initiatives targetted more realistic design points.
In particular there's often a desire to push out of order design into the micro-architecture where the resulting performance just doesn't justify it. In this they're achieving a CoreMark/MHz of 2.44 (from the paper here: https://upcommons.upc.edu/bitstream/handle/2117/384912/sargantana_preprint.pdf https://upcommons.upc.edu/bitstream/handle/2117/384912/sarga...). This is very low performance (on a par with the Arm M0+). Now CoreMark certainly isn't the be all and end all of Benchmarks. In particular it has very little relevance to high performance compute or application cores in general. However it's a useful performance smoke test. It is easy to perform well e.g getting close to 1.0 IPC for a single issue design such as Sargantana, CoreMark doesn't really stress the memory system so a major source of stalls that you need to hide latency for just isn't there. So if you're not hitting that you've definitely got work to do on the microarchitecture. They may well have been better off trying to build something simpler and putting more design time into improving the performance of the basic microarchitecture.
The other crucial aspect that's often overlooked is verification. This is a major part of producing a new production quality CPU design and it doesn't appear to be discussed in the paper at all. Maybe once they've released the RTL they'll also release the testbench so you can see what they have done.
- gchadwick 3y agoThough on the CoreMark benchmark they haven't published the IPC achieved. You get a large swing in results depending upon the compiler used and switches (For RV32 at least I've found GCC out-performs LLVM comfortably). They do have an IPC number for Dhrystone (another tiny benchmark that tells you little about real-world performance but you should be able to perform well on), that looks to be 0.7.
- phkahler 3y agoAny of these efforts not performing as well as BOOM may be suffering from "not invented here". Its already there and getting good IPC. Why not start from that.
- cf1241290841 3y agoI believe we might be at the point where supply chain security (and code base security) might warrant the question why you cant implement something on an M0+. If you really need higher speeds for reaction time, use an ASIC or FPGA. We already do this with USB3 or Ethernet controllers.