6 ms·
Absolutely! Chip designers have a several tools to do this. First, they create detailed software models (usually in C++) of their chips to estimate performance
by universal_sinc 3y ago
Absolutely! Chip designers have a several tools to do this.
First, they create detailed software models (usually in C++) of their chips to estimate performance as closely as they can before laying out a single transitory. These models can run code just like a real hardware device, albeit slowly.
Once the chip is designed, verilog simulators are programs used to generate the exact logical output of a circuit, which can be used to measure performance on a workload. However, this method is even slower than the first!
For larger workloads and higher speed, they use extraordinarily expensive FPGA-based platforms called Emulators. This allows circuits to be run at speeds in the MHz range before ever being sent to a fab. Booting an OS, running a complex multicore workload with shared memory, they can measure almost any workload. But this method is not available until late in the design phase and the boxes themselves are prohibitively expensive from being deployed very widely.
The software models are the most useful for estimating performance, as long as they are written early and well :)
- vintagedave 3y ago> they create detailed software models (usually in C++) of their chips to estimate performance as closely as they can How does this work? Do they model at the transistor level, or at the level of logical functions, or..? I'm particularly curious how this can estimate performance if it's anything higher-level than a direct transistor-for-transistor, layout-aware, emulation. I'd be really interested in learning more if there's anything you could share, please. I can find info about chip design software and languages like Verilog (as you mention) but not this sort of modeling.
- nudgeee 3y agoWhen I was in chip design about 15 years ago, we did transaction level modeling (TLM) using SystemC. Not sure if it’s still a thing these days. https://en.m.wikipedia.org/wiki/Transaction-level_modeling https://en.m.wikipedia.org/wiki/Transaction-level_modeling
- amelius 3y agoBasically you model all the elements of a chip (queues, memory, alus, etc), and how much time they take. You use a virtual clock so your simulation model can run at a different pace.
- Propelloni 3y agoThere are a few specialized languages for hardware description, Verilog is common, as is VHDL. A good point for starting is the Wikipedia page about hardware description languages [1]. This is a slow moving area, so even old resources should be useful. I only encountered HDLs during my university years and that's longer ago than I care to remember. I recall we did something with MIPS (back then we did everything hardware-near on MIPS) and used a book by O'Reilly, something something Systems Design or so. Couldn't find it, probably wrong name, but I found this [2], maybe useful? [1] https://en.wikipedia.org/wiki/Hardware_description_language https://en.wikipedia.org/wiki/Hardware_description_language [2] https://freecomputerbooks.com/langVHDLBooks.html https://freecomputerbooks.com/langVHDLBooks.html
- universal_sinc 3y agoThe idea is to write a C++ model that that produces cycle accurate outputs of the branch predictor, core pipeline, queues, memory latency, cache hierarchy, prefetch behaviour, etc. Transistor level accuracy isn't needed as long as the resulting cycle timings are identical or near identical. The improvement in workload runtime compared to a Verilog simulation is precisely because they aren't trying to model every transistor, but just the important parameters which effect performance. Let's take a simple example: Instead of modeling a 64-bit adder in all its gory transistor level detail, you can just have the model return the correct data after 1 "cycle" or whatever your ALU latency is. As long as that cycle latency is the same as the real hardware, you'll get an accurate performance number. What's particularly useful about these models is they enable much easier and faster state space exploration to see how a circuit would perform, well before going ahead with the Verilog implementation, which relatively speaking can take circuit designers ages. "How much faster would my CPU be if it had a 20% larger register file" can be answered in a day or two before getting a circuit designer to go try and implement such a thing. If you want an open source example, take a look at the gem5 project (https://www.gem5.org https://www.gem5.org). It's not quite as sophisticated as the proprietary models used in industry, but it's a used widely in academia and open source hardware design and is a great place to start.
- vintagedave 3y agoThis was really interesting to learn. Thankyou!
- KenArrari 3y agoA good example is one of the classic Computer Architecture class assignments which is to simulate a cache. So the way that looks is you have a stream of memory accesses and you "simulate it" by parsing that file and simulating the actions that would be taken. ie: "ok this block would be put in cache. This next access was a hit, this next access was a miss, etc". So then you just count those actions and estimate the performance by tallying all that up. That's the behavioral model part and IRL they do basically the same thing to decide what behavior they actually want the hardware to do. The next step is the circuit-level model done in verilog which actually simulates the logic-gates and does involve viewing a signal at every clock cycle.
- matt_d 3y agoThe following is a pretty good overview: "A Survey of Computer Architecture Simulation Techniques and Tools" - IEEE Access 2019 - Ayaz Akram, Lina Sawalha - https://ieeexplore.ieee.org/document/8718630 https://ieeexplore.ieee.org/document/8718630 For more see also: https://github.com/MattPD/cpplinks/blob/master/comparch.md#emulation--simulation https://github.com/MattPD/cpplinks/blob/master/comparch.md#e...
- FirmwareBurner 3y ago>they create detailed software models (usually in C++) of their chips to estimate performance as closely as they can before laying out a single transitory. Usually System Verilog instead of C++ but it has C++ interfaces https://en.wikipedia.org/wiki/SystemVerilog https://en.wikipedia.org/wiki/SystemVerilog