4 ms·
Why is it supposed to be so complicated to do "processor verification"? Why can't you simply upload the design to an FPGA, and then check that it can: 1. Boot
by devit 6y ago
Why is it supposed to be so complicated to do "processor verification"?
Why can't you simply upload the design to an FPGA, and then check that it can:
1. Boot all available operating systems (Linux, *BSD, Windows, etc.)
2. Successfully compile and run the testsuites for a bunch of open-source software (several languages like Rust have a standardized repository and method to build and run tests, so this is very easy)
3. Correctly run stress testing software (Prime95, etc.)
4. Correctly run several software unit tests that you write to exercise instructions that may not be produced by LLVM/GCC
5. Correctly run tests you write to exercise specific processor/cache states
6. Properly handling fuzzed code without freezing the whole CPU (using afl-fuzz)
Start with the simplest possible in-order core so that you get it working very easily, and then evolve to your desired end-state with a series of small commits, and if the verification fails use `git bisect` if needed to find the offending commit, insert any instrumentation you might need to detect the issue and fix it.
I don't see why you would need a specialized tool for that, or even what a specialized tool could possibly do.
- rrss 6y agoI think you are glossing over a lot of complexity in > 4. Correctly run several software unit tests that you write to exercise instructions that may not be produced by LLVM/GCC and > 5. Correctly run tests you write to exercise specific processor/cache states These two alone seem like they could be really quite complicated. Also, "run the world" is pretty slow when you have to do it in simulation (or emulation if you wait to have a netlist to find out how broken it is). I suspect the coverage from your list is substantially lower than you might expect. Would this have caught F00F? FDIV? AMD Phenom's TLB bug?
- joosters 6y agoHave a look at the extensive errata Intel publish for their CPUs. There are hundreds of mistakes in the chips’ behaviour, and yet each buggy CPU would pass your set of tests with flying colours. While you could never release a CPU that didn’t pass the tests you describe, they don’t even begin to exercise all the corner cases for a chip. Multiplying two specific numbers together, while the instruction crosses two memory pages, when an interrupt arrives? How do you even test for that kind of thing?
- wallacoloo 6y agoI feel the SW world is affected by some analogous bugs though. Any sort of race condition between two different threads accessing the same resource, maybe throw in some other piece like having the data always be valid unless a third thread happens to free some downstream resource at the same time... We know about techniques to reduce large classes of errors. Data races in particular can be prevented by some languages statically. Other types of “once in a blue-moon” errors that happen as a result of two coupled systems doing something in tandem can be reduced by introducing stronger boundaries between the systems, and then you can test each system independently and make sure it works regardless of what the other system does (I.e. dependency testing, or maybe even fuzzing). These approaches aren’t bulletproof, but I think they do illustrate a point: that there are techniques to reduce the likelihood of the errors you highlight. Whether they do it at a competitive cost to existing industry practices or not, I have no idea.
- TheCoelacanth 6y agoHardware is usually orders of magnitude more reliable than software, so what makes you think that they aren't already using those techniques or something better?
- wallacoloo 6y agoHmm? Maybe they do, I hope they do. I was replying more specifically to this part: > Multiplying two specific numbers together, while the instruction crosses two memory pages, when an interrupt arrives? How do you even test for that kind of thing? I.e. trying to despell the idea that large systems are intrinsically difficult to test.
- sweden 6y agoThis is a very well formulated question, well asked. People are already doing every step you mentioned but there are three problems: - Processors nowadays are so advanced and complex that you can't simply approach them as if it was one single block. You need to divide the processor into smaller blocks, develop those smaller blocks and put them together by the end of the development. Like in any engineering problem. The main problem is that it takes a lot of time for all sub blocks to be mature and stable enough for top level integration. - Once you have all your blocks ready, you can start integration and bring up the system with FPGAs. But now you face the problem that FPGAs are really slow and are not really usable as a normal system. You can run some preliminary tests, the short ones but it would take months to properly execute a normal benchmark. - You could create and tape out test chips but then you would also need to create all the infrastructure needed for the processor to work like the memory system, memory RAM, communication buses, firmware and etc.
- gchadwick 6y agoWell if you've just run your verification on FPGA like that chances are your silicon will fall over on first boot because you've totally missed a whole bunch of corners cases that only occur on the real memory system. Yes you can attempt to emulate these in FPGA but that's one of the reasons verification is not as easy as it seems. > Start with the simplest possible in-order core so that you get it working very easily, and then evolve to your desired end-state with a series of small commits You can't just trivially evolve a simple design into something more complex, much in the same way when Linux does a new major release they haven't started with some stripped down basic *nix and worked their way up from there.
- Koshkin 6y agoA tl;dr kind of answer is simply that an FPGA-based implementation is closer to a software emulator than to the real chip.
- not2b 6y agoFirst problem: "an FPGA". Unless it is far behind the state of the art, your processor design won't fit on a single FPGA chip. You need to partition the design, and run it on many FPGAs, and do the partition in a way that is correct and doesn't drop the performance to almost nothing. This is what the EDA emulation vendors will sell you: the systems with FPGAs organized into boards and racks of boards, and the software to target those systems, because to fit a cutting edge processor into the system you're going to need hundreds or thousands of FPGA chips. Once you do all that, you can try carrying out your program as described above. But it guaranteed that the first time you try it, it won't work, because your design will have bugs. Then what? You need the ability to debug. This means you need to have probes, you need to be able to extract the data, and you need very high bandwidth. You need to have testbenches that are partly in software and partly on the FPGA hardware. Again, that's what the EDA industry will sell you: the hardware and software to do it, as well as the expert consultants to walk you through the process. And your device needs to interact with the environment. Some of the hardest verification problems have to do with the timing of interrupts; if one comes when the processor is just at the right point, and that case isn't handled in the design, it could lock up. THose cases have to be covered. Now, for your example of a small core, perhaps it's small enough that you could get it to synthesize and fit into one large FPGA chip and avoid some of these issues. Good luck doing something that can boot Android in that size.
- monocasa 6y agoI thought the big cadence emulators weren't made of mainly FPGA chips, but instead arrays of absolutely tiny processor cores that only know logic and branch ops. Not that this makes a difference for the core of the point you're making, it's more an aside.
- _chris_ 6y agoOnly speaking for myself, I think of FPGA and hw emulation interchangeably. At the end of the day, I don't care (or know) how it's implemented. What I can say is a) they aren't super fast, only ~1Mhz b) they take forever to compile down to, c) they have some, but not great visibility to what went wrong, and d) they cost a stupid amount of money.
- imtringued 6y agoImagine if your test suite with 400 tests simply returned true or false instead of telling you which tests have failed. It's going to take forever to find the failing tests.
- _chris_ 6y agoGreat question. For starts though, you've just given me a list worth hundreds of trillions of instructions of tests. If you can actually test on an FPGA (which is usually not possible, and if it is, it's not representative of the actual silicon/analog system you're building anyways), you can get 1T instructions in roughly ~6 hours at 50 MHz. But most hardware emulations are ~1 MHz, so now you're looking at weeks to hit 1T instructions (SPECint alone is 20T). And what happens when you hit a bug, 2 weeks in? It may not be because you actually wrote new, buggy code, but because a new, higher performance branch predictor uncovered existing bugs. But you'll never know, because the FPGA historically gives you terrible visibility. But in simulation (where testing is actually done), you're looking at ~1 Hz for a cpu core. Ouch. Obviously a better approach is required (unit-tests against models, formal, etc.), since at the level of detail you can't test much of anything.
- _ph_ 6y agoJust a random FPGA won't cut it, if you want to simulate a reasonably large processor design. Also, you would only be able to simulate the basic logic for its validity. But that is only the very first and easiest step in the verification chain. Things start to get more interesting, when you look into the real-world analog properties of you chip. You want to check the correct timing so that the logic still works correctly at high clock speeds - depends very much on the acutal placing of the components, e.g. wire length. Then there is the question of thermal behavior. Long term stability (years of operation). And then we enter the space of manufacturing. You need to optimize your design so that it can be manufactured reliably. There are some systematic variations across the wafer and your design should account for that. There are also a lot of constraints given by the production process, how you have to distribute components on your chip. Getting from an initial logic design to a manufactured chip is a big adventure. This requires a lot of layers of software, a lot of it highly specified. And yes, you can buy very good simulators for chips, just search for Cadence Palladium for example. They are huge monsters.