4 ms·
[ time to re-share my Dr. Nicely story ] Dr. Nicely caused quite a bit of excitement at Intel. I was on the p6 architecture team when he discovered the FDIV bu
by wscott 3y ago
[ time to re-share my Dr. Nicely story ]
Dr. Nicely caused quite a bit of excitement at Intel. I was on the p6 architecture team when he discovered the FDIV bug. Our FPU was formally verified and didn't have the same bug.
To be nice to Dr Nicely we sent him a pre-release p6 development system to test with his program to demonstrate that his bug was fixed. He was working on a prime number sieve program and came back reporting that the p6 ran at 1/2 the speed of a Pentium for his code. Wow, another blackeye/firestorm caused by Dr. Nicely. He had too much of an audience for him to report to the world this new processor was slower.
So I got to spend a lot of time learning how to sieve works and what is happening. For the most part, it allocates a huge array in memory with each byte representing a number. You walk the array with a stride of known primes setting bytes and whatever is left must be prime. ie. every 3 is not prime, every 5 is not prime, every 7...
So in the steady state, you are writing a single byte to a cache line without reading anything. And every write hits a different cache line.
Now p6 had a write-allocate cache, but the Pentium would only allocate on read, so on the Pentium a write that misses the cache would become a write to memory. On the p6 that write would need to load the cache line from memory into the cache and then the line in the cache was modified. And since every line in the cache was also modified we had to flush some other cache line first to make room. So every 1-byte write would become a 32-byte write to memory followed by a 32-byte read from memory.
Normally write-allocate is a good thing, but in this case, it was a killer. We were stumped.
Then the magic observation: 99% of these writes were marking a space that was already marked. When you get up to walking by large strides most of those were already covered by one of the smaller factors.
So if you change the code from:
array[N] = 1
to:
if (!array[N]) array[N] = 1
Now suddenly we are doing a read first, and after that read we skip the write so the data in the cache doesn't become modified and can be discarded in the future.
Also, the p6 was a super-scalar machine that ran multiple iterations of this loop in parallel and could have multiple reads going to memory at the same time. With that small tweak, the program got 4X faster and we went from being 1/2X the speed of a Pentium to being twice the speed. And this was at the same clock frequency. The test hardware ran 100Mhz, we released at 200Mhz and went up from there.
- nextaccountic 3y agoCould this transformation be still effective in modern processors?
- retrac 3y agoIt wouldn't be needed today on x86_64. The underlying idea - priming the processor with hints about what and when - is now standard practice. There are explicit instructions for manipulating the cache and marking data as needed-soon. E.g. PREFETCH, available as __builtin_prefetch in GCC. Added to the Pentium III, I think.
- ghaff 3y agoThere are a lot of cache-related performance behaviors that can be triggered by the "wrong" code. I was the product for a minicomputer system once and an outside performance consultant stumbled on one of those which caused a bit of a kerfuffle at the time. As I recall, it wasn't really an issue for real-world workloads but if you did things just wrong, it could be a significant performance hit.
- sponaugle 3y agoThe Pentium Pro is by far the CPU I remember being the most surprised about, and I was an engineer at Intel when it was released. Such an amazing design and manufacturing technique, and 200Mhz with 256 or 512k L2 cache! That seemed like such a huge cache at the time. Performance on video codecs with some tweaks was a 2x jump as well. Funny enough I still have a working Pentium Pro 200 system. good times!
- inversetelecine 3y agoI regret getting rid of mine. Harder to find now, or 'pricey' on eBay compared to 10 years ago. I'll never forget the size of the cpus compared to others at the time. Running them in dual-cpu setups was fun also.
- sponaugle 3y ago