3 ms·
The article was about testing a specific implementation of floor, ceil, and round using CPU instructions. You can run floor, ceil, and round on the GPU, but tha
by daveFNbuck 6y ago
The article was about testing a specific implementation of floor, ceil, and round using CPU instructions. You can run floor, ceil, and round on the GPU, but that won't tell you whether your new implementation using CPU instructions is correct.
- dragontamer 6y agoI thought the post was about how the 32-bit space could be exhausted in minutes with common hardware in the year 2014. Which means that any problem I have which is in the 32-bit space is a candidate for brute force. Given the advances in computers over the past few years: exhausting the 64-bit (or 2x 32-bit) space is "reasonable" for the year 2020. At least, reasonable for people who are willing to wait a month (on one GPU), or willing to buy a cluster of GPUs.
- mlyle 6y agoNo, the article is that if you have some code that operates on a 32 bit value, it's reasonable on most modern hardware to try all 2^32 values to ensure quality. There may be some GPU code that is reasonable to test on all 2^64 values, but it doesn't help us verify the implementation of SSE (Intel architecture) ceil/round/etc against reference implementations at all, because trying all double precision values would not complete in a practical amount of time and offloading it to a GPU means we are not testing the SSE implementation that we're trying to verify anymore.
- a1369209993 6y ago> trying all double precision values would not complete in a practical amount of time Actually we're getting close-ish: AVX512(=8x64) at 4GHz (or two dependency chains at 2GHz) times 64 CPU cores (which TFA's author has mentioned having) is 2 teraops. Double-precision space of 2^64 / 2Tops is 3 months/instruction on a single computer. So you could brute-force verify a simple function in a year or two, and at least the SIMD width and number of cores is likely to continue increasing.
- mlyle 6y agoKeep in mind you need to run the non-SIMD baseline function, too, to perform the actual comparison/verification... So, not yet. Also, AVX512 throttles down, so you're not going to running it on all cores full tilt at 4GHz today. I think it's much more reasonable for 64 bit space to have a reasonable set of human-chosen test cases + a large randomized test sequence.
- daveFNbuck 6y agoThe post was about exhausting the space to test code. The example they used of code that should have been tested this way was CPU code. If you're testing CPU code, exhausting the space quickly on a GPU isn't helpful.
- masklinn 6y ago> I thought the post was about how the 32-bit space could be exhausted in minutes with common hardware in the year 2014. Which means that any problem I have which is in the 32-bit space is a candidate for brute force. No, the post is about exhaustive testing, and that it's completely reasonable on 32 bit values because it only takes a few minutes. > Given the advances in computers over the past few years: exhausting the 64-bit (or 2x 32-bit) space is "reasonable" for the year 2020. That's not actually reasonable, it's feasible for very specific contexts. > At least, reasonable for people who are willing to wait a month (on one GPU), or willing to buy a cluster of GPUs. And whose code specifically can run on GPU. TFA is about CPU-optimised code, the entire point was to exhaustively test an SSE3 implementations. How do you do that on a GPU?
- dragontamer 6y ago> That's not actually reasonable, it's feasible for very specific contexts. I was inventing a new hash function for myself, just for giggles a few weeks ago. I needed a constant: I started by choosing an arbitrary constant (I started with 0x31415926...), but then I realized that the 32-bit space, and even 50+ bit spaces were feasible on my GPU. I ended up rewriting the random-number generator on the GPU, and searched for the number which caused the largest number of "avalanche" bit-flips across the entire 32-bit seed space of the RNG. Where an "avalanche" bit flip is 16 bits flipped, and 16-bits remaining the same, after the RNG was applied to the seed. For example: seed = 1 * K, where K == 0xAAAAAAAA will cause 16-bit flips, and 16-bits to remain the same. Repeat for all 32-bit values of seed, searching for the optimal K. 0xAAAAAAAA wasn't the best number (aka: 1010101010101010 binary), but you can see where my thought process was. My RNG was more complicated than just a single multiply, but I think you can see where the general thought process was with my above example. -------- There was a time when constants were arbitrarily chosen for these kinds of functions. But today, we can literally test ALL numbers and just pick the best. If K is constrained by some other requirements (in my case: it must be odd, and a few other tidbits to work with my RNG), then you shrink the search space from 64-bits down to something that can be accomplished in just a few days of search. ------------- As far as I'm concerned, the blog post is talking about how modern computers have conquered the 32-bit space. My current post is how the 64-bit space could be feasibly explored. It wasn't even that long ago when 64-bit space was considered cryptographic-secure (see DES crypto algorithm or WEP).