4 ms·
Wow, totally unaware of the performance speed-up 5x over sha1 and blake2b. Even more for others! Are there cases where the benchmark doesn't hold true? Differ
by 2bluesc 5y ago
Wow, totally unaware of the performance speed-up 5x over sha1 and blake2b. Even more for others!
Are there cases where the benchmark doesn't hold true? Different architectures?
I'm inclined to use this on some less security sensitive things I need a hash for.
- staticassertion 5y agoThere shouldn't be any meaningful reduction in security afaik. Most of the performance comes from reducing the number of rounds used because it was shown that it was unnecessary iirc.
- searealist 5y agoBLAKE3 only reduces the rounds to 7 from 10 in BLAKE2. That only accounts for a 1.42x speedup, not the > 5x speedup seen in BLAKE3 for a single thread. Most of the speedup comes from the hashing mode which breaks the input into a tree of independent chunks and then using SIMD to hash several of them in parallel. Additionally, the Rust implementation even allows even more parallelism using threads.
- staticassertion 5y agoGotcha, I wasn't sure if it was the parallelism or the reduced rounds that made the bigger difference.
- anfilt 5y agoStill BLAKE3 still is not a ton faster for short data inputs especially if your implementation has a lot overhead for parallelism like threading like you mentioned in a rust implementation. Good for files and similar if your data stream is long enough to effectively use the tree structure.
- rowanG077 5y agocheck out https://github.com/Cyan4973/xxHash https://github.com/Cyan4973/xxHash for I use it for performance sensitive non-crypto stuff.
- oconnor663 5y ago> Are there cases where the benchmark doesn't hold true? Different architectures? Yes, that 5x figure represents a big speedup that comes from using SIMD (specifically AVX-512) to compress multiple blocks in parallel. To take full advantage of AVX-512, BLAKE3 needs to have at least 16 KiB of input. So for any input shorter than that, the speedup compared to BLAKE2b/s will be smaller. For inputs less than 2 KiB, the only speedup is the ~1.4x that you get from the round reduction. Architectures other than x86 tend to have less in the way of SIMD. ARM has NEON, but that's currently only 128 bits wide. So the big single-threaded speedups are only on x86 today. (For multithreading speedups, the architecture doesn't matter as much. Just the number of cores you have.) > I'm inclined to use this on some less security sensitive things I need a hash for. This makes sense insofar as BLAKE3 is a new design, and it pays for cryptographic applications to be conservative. But to be clear, if BLAKE3 turns out not to uphold the same security properties as SHA-2 or BLAKE2s, that would represent a catastrophic failure of the design. (It's possible that some flaw could be discovered that affects BLAKE2 and BLAKE3 equally, but that the lower round count of BLAKE3 would make the flaw more severe. In that case, I'd expect everyone would also want to migrate off BLAKE2 pretty quickly anyway.)
- bradknowles 5y agoInteresting then that the latest Intel chips are dropping AVX-512. Makes me wonder if BLAKE3 will be changed to suit, or if Intel chips are just going to suck when it come to performance with the standard?
- oconnor663 5y agoI'm not sure any changes are necessary. The implementations will use AVX-512 where it's available, and they'll fall back to something else (often AVX-2) if it's not.