18 ms·
Reasons to Prefer Blake3 over Sha256
- ndsipa_pomu 3y agoAt this rate, it's going to take over 700 years before we get Blake's 7
- benj111 3y agoI had to scroll disappointingly far down to get to the Blake's 7 reference. Thank you for not disappointing though. The down side of that algorithm though is that everything dies at the end.
- dragontamer 3y ago> BLAKE3 is much more efficient (in time and energy) than SHA256, like 14 times as efficient in typical use cases on typical platforms. [snip] > AVX in Intel/AMD, Neon and Scalable Vector Extensions in Arm, and RISC-V Vector computing in RISC-V. BLAKE3 can take advantage of all of it. Uh huh... AVX/x86 and NEON/ARM you say? https://www.felixcloutier.com/x86/sha256rnds2 https://www.felixcloutier.com/x86/sha256rnds2 https://developer.arm.com/documentation/ddi0596/2021-12/SIMD-FP-Instructions/SHA256H2--SHA256-hash-update--part-2-- https://developer.arm.com/documentation/ddi0596/2021-12/SIMD... If we're talking about vectorized instruction sets like AVX (Intel/AMD) or NEON (aka: ARM), the advantage is clearly with SHA256. I don't think Blake3 has any hardware implementation at all yet. Your typical cell phone running ARMv8 / NEON will be more efficient with the SHA256 instructions than whatever software routine you need to run Blake3. Dedicated hardware inside the cores is very difficult to beat on execution speed or efficiency. I admit that I haven't run any benchmarks on my own. But I'd be very surprised if any software routine were comparable to the dedicated SHA256 instructions found on modern cores.
- eatonphil 3y agoFrom another thread: > On my machine with sha extensions, blake3 is about 15% faster (single threaded in both cases) than sha256. https://news.ycombinator.com/item?id=22237387 https://news.ycombinator.com/item?id=22237387
- vluft 3y agofollowup to this now with further blake3 improvements, on a faster machine now but with sha extensions vs single-threaded blake3; blake3 is about 2.5x faster than sha256 now. (b3sum 1.5.0 vs openssl 3.0.11). b3sum is about 9x faster than sha256sum from coreutils (GNU, 9.3) which does not use the sha extensions. Benchmark 1: openssl sha256 /tmp/rand_1G Time (mean ± σ): 576.8 ms ± 3.5 ms [User: 415.0 ms, System: 161.8 ms] Range (min … max): 569.7 ms … 580.3 ms 10 runs Benchmark 2: b3sum --num-threads 1 /tmp/rand_1G Time (mean ± σ): 228.7 ms ± 3.7 ms [User: 168.7 ms, System: 59.5 ms] Range (min … max): 223.5 ms … 234.9 ms 13 runs Benchmark 3: sha256sum /tmp/rand_1G Time (mean ± σ): 2.062 s ± 0.025 s [User: 1.923 s, System: 0.138 s] Range (min … max): 2.046 s … 2.130 s 10 runs Summary b3sum --num-threads 1 /tmp/rand_1G ran 2.52 ± 0.04 times faster than openssl sha256 /tmp/rand_1G 9.02 ± 0.18 times faster than sha256sum /tmp/rand_1G
- 0xdeafbeef 3y agoInterestingly, sha256sum and openssl don't use sha_ni. iced-cpuid $(which b3sum) AVX AVX2 AVX512F AVX512VL BMI1 CET_IBT CMOV SSE SSE2 SSE4_1 SSSE3 SYSCALL XSAVE iced-cpuid $(which openssl ) CET_IBT CMOV SSE SSE2 iced-cpuid $(which sha256sum) CET_IBT CMOV SSE SSE2 Also, sha256sum in my case is a bit faster Benchmark 1: openssl sha256 /tmp/rand_1G Time (mean ± σ): 540.0 ms ± 1.1 ms [User: 406.2 ms, System: 132.0 ms] Range (min … max): 538.5 ms … 542.3 ms 10 runs Benchmark 2: b3sum --num-threads 1 /tmp/rand_1G Time (mean ± σ): 279.6 ms ± 0.8 ms [User: 213.9 ms, System: 64.4 ms] Range (min … max): 278.6 ms … 281.1 ms 10 runs Benchmark 3: sha256sum /tmp/rand_1G Time (mean ± σ): 509.0 ms ± 6.3 ms [User: 386.4 ms, System: 120.5 ms] Range (min … max): 504.6 ms … 524.2 ms 10 runs
- vluft 3y agonot sure that tool is correct; on my openssl it shows same output as you have there, but not aes-ni which is definitely enabled and functional. ETA: ahh you want to do that on libcrypto: iced-cpuid <...>/libcrypto.so.3: ADX AES AVX AVX2 AVX512BW AVX512DQ AVX512F AVX512VL AVX512_IFMA BMI1 BMI2 CET_IBT CLFSH CMOV D3NOW MMX MOVBE MSR PCLMULQDQ PREFETCHW RDRAND RDSEED RTM SHA SMM SSE SSE2 SSE3 SSE4_1 SSSE3 SYSCALL TSC VMX XOP XSAVE
- insanitybit 3y ago> I don't think Blake3 has any hardware implementation at all yet. > https://github.com/BLAKE3-team/BLAKE3 https://github.com/BLAKE3-team/BLAKE3 > The blake3 Rust crate, which includes optimized implementations for SSE2, SSE4.1, AVX2, AVX-512, and NEON, with automatic runtime CPU feature detection on x86. The rayon feature provides multithreading. There aren't blake3 instructions, like some hardware has for SHA1, but it does use hardware acceleration. edit: Re-reading, I think you're saying "If we're going to talk about hardware acceleration, SHA1 still has the advantage because of specific instructions" - that is true.
- jonhohle 3y agoI just tested the C implementation on a utility I wrote[0] and at least on macOS where SHA256 is hardware accelerated beyond just NEON, BLAKE3 ends up being slower than SHA256 from CommonCrypto (the Apple provided crypto library). BLAKE3 ends up being 5-10% slower for the same input set. As far as I'm aware, Apple does not expose any of the hardware crypto functions, so unless what exists supports BLAKE3 and they add support in CommonCrypto, there's no advantage to using it from a performance perspective. The rust implementation is multithreaded and ends up beating SHA256 handily, but again, for my use case the C impl is only single threaded, and the utility assumes a single threaded hasher with one running on each core. Hashing is the bottleneck for `dedup`, so finding a faster hasher would have a lot of benefits. 0 - https://github.com/ttkb-oss/dedup https://github.com/ttkb-oss/dedup
- RaisingSpear 3y agoKeep in mind that many CPUs out there don't support those instructions (notably Intel's Skylake and ARM's Cortex A72). BLAKE3 will be significantly faster than SHA2 on many platforms out there.
- tromp 3y agoFor short inputs, Blake3 behaves very similar to Blake2, on which it is based. From Blake's wikipedia page [1]: BLAKE3 is a single algorithm with many desirable features (parallelism, XOF, KDF, PRF and MAC), in contrast to BLAKE and BLAKE2, which are algorithm families with multiple variants. BLAKE3 has a binary tree structure, so it supports a practically unlimited degree of parallelism (both SIMD and multithreading) given long enough input. [1] https://en.wikipedia.org/wiki/BLAKE_(hash_function) https://en.wikipedia.org/wiki/BLAKE_(hash_function)
- cesarb 3y agoWhile I really like Blake3, for all reasons mentioned in this article, I have to say it does have one tiny disadvantage over older hashes like SHA-256: its internal state is slightly bigger (due to the tree structure which allows it to be highly parallelizable). This can matter when running on tiny microcontrollers with only a few kilobytes of memory.
- londons_explore 3y agoThe internal state is no bigger when hashing small things though right? I assume most microcontrollers are unlikely to be hashing things much bigger than RAM.
- oconnor663 3y agoIt's hard to give a short answer to that question :) - Yes, if you know your input is short, you can use a smaller state. The limit is roughly a BLAKE2s state plus (32 bytes times the log_2 of the number of KiB you need to hash). Section 5.4 of https://github.com/BLAKE3-team/BLAKE3-specs/blob/master/blake3.pdf https://github.com/BLAKE3-team/BLAKE3-specs/blob/master/blak... goes into this. - But it's hard to take advantage of this space optimization, because no libraries implement it in practice. - But the reason libraries don't implement it is that almost no one needs it. The max state size is just under 2 KiB, which is small enough even for https://github.com/oconnor663/blake3-6502 https://github.com/oconnor663/blake3-6502. - But it would be super easy to implement if we just put the "CV stack" on the heap instead of allocating the whole thing as an array up front. - But the platforms that care about this don't have a heap. @caesarb mentioned really tiny microcontrollers, even tinier than the 6502 maybe. The other place I'd expect to see this optimization is in a full hardware implementation, but those are rare. Most hardware accelerators for hash functions provide the block operation, and they leave it to software to deal with this sort of bookkeeping.
- gavinhoward 3y agoGood, terse article that basically reinforces everything I've seen in my research about cryptographic hashing. Context: I'm building a VCS meant for any size of file, including massive ones. It needs a cryptographic hash for the Merkle Tree. I've chosen BLAKE3, and I'm going to use the original implementation because of its speed. However, I'm going to make it easy to change hash algorithms per commit, just so I don't run into the case that Git had trying to get rid of SHA1.
- AdamN 3y agoSmart idea doing the hash choice per-commit. Just make sure that somebody putting in an obscure hash doesn't mess up everybody's usage of the repo if they don't have a library to evaluate that hash installed.
- gavinhoward 3y agoI agree. There will be a set of presets of hash function and settings; if BLAKE3 fails, then I'll actually have to add SHA3 or something, with a set of settings, as presets. The per-commit storage will then be an enum identifying the hash and its settings. This will let me do other things, like letting companies use a 512-bit hash if they expect the repo to be large.
- agodfrey 3y agoMaybe you’re already aware, but you glossed over something: Since you’re using the hash to locate/identify the contect (you mentioned Merkle and git), if you support multiple hash functions you need some assurance that the chance of collisions is low across all supported hash functions. For example two identical functions that differ only in the value of their padding bytes (when the input size doesn’t match the block size) can’t coexist.
- gavinhoward 3y agoYou are absolutely right. And yes, I am aware. Location will actually be done by prefixing the hash with the value of the enum for the hash function/settings pair that made the hash.
- EdSchouten 3y agoWhat I dislike about BLAKE3 is that they added explicit logic to ensure that identical chunks stored at different offsets result in different Merkle tree nodes (a.k.a. the ‘chunk counter’). Though this feature is well intended, it makes this hash function hard to use for a storage system where you try to do aggressive data deduplication. Furthermore, on platforms that provide native instructions for SHA hashing, BLAKE3 isn’t necessarily faster, and certainly more power hungry.
- lazide 3y agoHuh? The storage system doing this wouldn’t use that part of the hash, it would do it itself so no issues? (Hash chunks, instead of feeding everything in linearly) Otherwise the hash isn’t going to be even remotely safe for most inputs?
- persnickety 3y agoCould you point to how this is implemented and how it can be used? From the sound of it, you're trying to do something like rsync's running-window comparison?
- EdSchouten 3y agoImagine the case where you're trying to create a storage system for a large number of virtual machine images (e.g., you're trying to build your own equivalent of AWS Machine Images). There is obviously a lot of duplication between parts of images. And not necessarily at the same offset, but also at different offsets that are n*2^k bytes apart, where 2^k represents the block/sector size. You could consider building this storage system on top of BLAKE3's tree model. Namely you store blocks as small Merkle tree. And an image is basically a collection of blocks that has a different 'hat' on top of it. Unfortunately, BLAKE3 makes this hard, because the same block will end up having a different Merkle tree node depending on the offset at which it's stored.
- luoc 3y agoYou mean something like a CDC algorithm? I know that some Backup tools like Restic use this. https://en.m.wikipedia.org/wiki/Rolling_hash https://en.m.wikipedia.org/wiki/Rolling_hash
- stylepoints 3y agoUntil it starts coming installed by default on Linux and other mojor OS's, it won't be mainstream.
- theamk 3y agoPython 3.11 will have it https://bugs.python.org/issue39298 https://bugs.python.org/issue39298
- latexr 3y agoThat says “Resolution: rejected” and Python is currently at 3.12.0. Did the feature land?
- theamk 3y agooops I misread it.. seems it was rejected because it was not standard enough... https://github.com/python/cpython/issues/83479#issuecomment-1093850702 https://github.com/python/cpython/issues/83479#issuecomment-...
- syonfox 3y agomurmur3
- Snawoot 3y agomurmur3 is not a cryptographic hash, so it's not even in the same field.
- latexr 3y agoIt bears mentioning `shasum` is better supported in that it ships with operating systems (macOS, I guess Linux depends on the distro, don’t know about Windows) and is available directly from programming languages (Ruby, Swift, Python, …) without the need for libraries. Even if BLAKE3 is massively faster, it’s not like I ever noticed SHA256’s supposed slowness. But I do notice its ubiquity. Based on the article, I would consider switching to BLAKE3 immediately where I use checksums. But until it gets wider support (might be easier with a public domain license instead of the current ones) I can’t really do it because I need to do things with minimal dependencies. Best of luck to the BLAKE3 team on making their tool more widely available.
- upget_tiding 3y ago> might be easier with a public domain license instead of the current ones There reference implementation is public domain (CC0) or at your choice Apache 2.0 https://github.com/BLAKE3-team/BLAKE3/blob/master/LICENSE https://github.com/BLAKE3-team/BLAKE3/blob/master/LICENSE
- latexr 3y agoYou’re right, I misread the CC0 part.
- zahllos 3y agoI agree that if you can, BLAKE3 (or even BLAKE2) are nicer choices than SHA2. However I would like to add the following comments: * SHA-2 fixes the problems with SHA-1. SHA-1 was a step up over SHA-0 that did not completely resolve flaws in SHA-0's design (SHA-0 was broken very quickly). * JP Aumasson (one of the B3 authors) has said publicly a few times SHA-2 will never be broken: https://news.ycombinator.com/item?id=13733069 https://news.ycombinator.com/item?id=13733069 is an indirect source, can't seem to locate a direct one from Xitter (thanks Elon). Thus it does not necessarily follow that SHA-2 is a bad choice because SHA-1 is broken.
- gavinhoward 3y agoAll that may be true. However, I don't think we can say for sure if SHA2 will be broken. Cryptography is hard like that. In addition, SHA2 is still vulnerable to length extension attacks, so in a sense, SHA2 is broken, at least when length extension attacks are part of the threat model.
- zahllos 3y agoIf you want to be pedantic we can say there is definitely a collision in SHA-2. Assume we have 2^256 unique inputs. Hash them all and assume no collisions. Now, if we have one more unique input (so 2^256 + 1 inputs) we have a collision. The same logic applies to BLAKE3. However we do actually know quite a bit on how to design hash functions to make this hard to do in practice. The latest cryptanalysis (to actually find a collision) either requires a vastly reduced number of rounds or is is computationally infeasible. There's no clear flaw like there was with SHA1, where the path to finding a collision has been known since ~2004. Length extension "attacks" sure, that's an unfortunate design choice. But it doesn't impact at all on collision resistance, which is what is implied by suggesting SHA1 is vulnerable then SHA2 is. In the end, if you can use BLAKE3 or BLAKE2, great, I probably would as well. There isn't always a choice (e.g. there's no blake3 support in most crypto hardware) and if there isn't, sha3 or sha2 are fine choices.
- Godel_unicode 3y agoI don’t understand why people use sha256 when sha512 is often significantly faster: https://crypto.stackexchange.com/questions/26336/sha-512-faster-than-sha-256 https://crypto.stackexchange.com/questions/26336/sha-512-fas...
- oconnor663 3y agoA couple reasons just on the performance side: - SHA-256 has hardware acceleration on many platforms, but SHA-512 mostly doesn't. - Setting aside hardware acceleration, SHA-256 is faster on 32-bit platforms, like a lot of embedded devices. If you have to choose between "fast on a desktop" vs "fast in embedded", it can make sense to assume that desktops are always fast enough and that your bottlenecks will be in embedded.
- adrian_b 3y agoOn older 64-bit CPUs without hardware SHA-256 (i.e. up to the Skylake derivatives), SHA-512 is faster. Many recent Arm CPUs have hardware SHA-512 (and SHA-3). Intel will add hardware SHA-512 starting with Arrow Lake S, to be launched at the end of 2024 (the successor in desktops of the current Raptor Lake Refresh). Most 64-bit CPUs that have been sold during the last 4 years and many of those older than that have hardware SHA-256.
- garblegarble 3y agoThis may only be applicable to certain CPUs - e.g. sha512 is a lot slower on M1 $ openssl speed sha256 sha512 type 16 bytes 64 bytes 256 bytes 1024 bytes 8192 bytes 16384 bytes sha256 146206.63k 529723.90k 1347842.65k 2051092.82k 2409324.54k 2446518.95k sha512 85705.68k 331953.22k 707320.92k 1149420.20k 1406851.34k 1427259.39k
- nayuki 3y agoIt's an interesting set of reasons, but I prefer Keccak/SHA-3 over SHA-256, SHA-512, and BLAKE. I trust the standards body and public competition and auditing that took place - more so than a single author trumpeting the virtues of BLAKE.
- jasonwatkinspdx 3y agoIronic, because the final NIST report explaining their choice mentions that BLAKE has more open examination of cryptanalysis than Keccak as a point in favor of BLAKE.
- tptacek 3y agoI'd probably use a Blake too. But: SHA256 was based on SHA1 (which is weak). BLAKE was based on ChaCha20, which was based on Salsa20 (which are both strong). NIST/NSA have repeatedly signaled lack of confidence in SHA256: first by hastily organising the SHA3 contest in the aftermath of Wang's break of SHA1 No: SHA2 lacks the structure the SHA1 attack relies on it (SHA1 has a linear message schedule, which made it possible to work out a differential cryptanalysis attack on it). Blake's own authors keep saying SHA2 is secure (modulo length extension), but people keep writing stuff like this. Blake3 is a good and interesting choice on the real merits! It doesn't need the elbow throw.
- ianopolous 3y agoWould be interesting to hear Zooko's response to this. (Peergos lead here)
- pclmulqdq 3y agoMost people who publicly opine on the Blake vs. SHA2 debate seem to be relatively uninformed on the realities of each one. SHA2 and the Blakes are both usually considered to be secure. The performance arguments most people make are also outdated or specious: the original comparisons of Blake vs SHA2 performance on CPUs were largely done before Intel and AMD had special SHA2 instructions.
- ianopolous 3y agoThe author is one of the creators of blake3, Zooko.
- tptacek 3y agoSorry, I should have been more precise. JP Aumasson is specifically who I'm thinking of; he's made the semi-infamous claim that SHA2 won't be broken in his lifetime. The subtext I gather is that there's just nothing on the horizon that's going to get it. SHA1 we saw coming a ways away!
- 3y ago
- RcouF1uZ4gsC 3y agoSHA-256 has the advantage that it is used for BitCoin. It is the biggest bug bounty of all time. There see literally billions riding on the security of SHA-256.
- aburan28 3y agoThere has been a mountain of cryptanalysis done on SHA256 with no major breaks compared to a much smaller amount analysis on blake3.
- jrockway 3y agoFast hash functions are really important, and SHA256 is really slow. Switching the hash function where you can is enough to result in user-visible speedups for common hashing use cases; verifying build artifacts, seeing if on-disk files changed, etc. I was writing something to produce OCI container images a few months ago, and the 3x SHA256 required by the spec for layers actually takes on the order of seconds. (.5s to sha256 a 50MB file, on my 2019-era Threadripper!) I was shocked to discover this. (gzip is also very slow, like shockingly slow, but fortunately the OCI spec lets you use Zstd, which is significantly faster.)
- coppsilgold 3y agosha256 is not slow on modern hardware. openssl doesn't have blake3, but here is blake2: type 16 bytes 64 bytes 256 bytes 1024 bytes 8192 bytes 16384 bytes BLAKE2s256 75697.37k 308777.40k 479373.40k 567875.81k 592687.09k 591254.18k BLAKE2b512 63478.11k 243125.73k 671822.08k 922093.51k 1047833.51k 1048959.57k sha256 129376.82k 416316.32k 1041909.33k 1664480.49k 2018678.67k 2043838.46k This is with the x86 sha256 instructions: sha256msg1, sha256msg2, sha256rnds2
- dralley 3y ago"modern hardware" deserves some caveats. AMD has supported those extensions since the original Zen, but Intel CPUs generally lacked them until only about 2 years ago.
- adrian_b 3y agoFor many years, starting in 2016, Intel has supported SHA-256 only in their Atom CPUs. The reason seems to be that the Atom CPUs were compared in Geekbench with ARM CPUs, and without hardware SHA the Intel CPUs would have obtained worst benchmark scores. In their big cores, SHA has been added in 2019, in Ice Lake (while Comet Lake still lacked it, being a Skylake derivative), and since then all newer Intel CPUs have it. So except for the Intel Core CPUs, the x86 and ARM CPUs have had hardware SHA for at least 7 years, while the Intel Core CPUs have had it for the last 4 years.
- ur-whale 3y agoOne metric that is seldom mentioned for crypto algos is code complexity. I really wish researchers would at least pay lip service to it. TEA (an unfortunately somewhat weak symmetric cipher) was a very nice push in that direction. TweetNaCl was another very nice push in that direction by djb Why care about that metric you ask? Well here are a couple of reasons: - algo fits in head - algo is short -> cryptanalysis likely easier - algo is short -> less likely to have buggy implementation - algo is short -> side-channel attacks likely easier to analyse - algo fits in a 100 line c++ header -> can be incorporated into anything - algo can be printed on a t-shirt, thereby skirting export control restrictions - algo can easily be implemented on tiny micro-controllers etc ...
- oconnor663 3y agoWe put a lot of effort into section 5.1.2 of https://github.com/BLAKE3-team/BLAKE3-specs/blob/master/blake3.pdf https://github.com/BLAKE3-team/BLAKE3-specs/blob/master/blak..., and the complicated part of BLAKE3 (incrementally building the Merkle tree) ends up being ~4 lines of code. Let me know what you think.
- rstuart4133 3y ago> One metric that is seldom mentioned for crypto algos is code complexity. ... TEA (an unfortunately somewhat weak symmetric cipher) was a very nice push in that direction. Spec is also a push in that direction [0]. It's code looks to be as complex as TEA's (1/2 a page of C), blindingly fast, yet as far I know has no known attacks despite being subject to a fair bit of scrutiny. About the only reason I can see for it not being largely ignored is it was designed by NSA. SHA3 is also a simple algorithm. Downright pretty, in fact. It's a pity it's so slow. [0] https://en.wikipedia.org/wiki/Speck_(cipher) https://en.wikipedia.org/wiki/Speck_(cipher)
- colmmacc 3y agoIt's very hard to see Blake3 getting included in FIPS. Meanwhile, SHA256 is. That's probably the biggest deciding factor on whether you want to use it or not.
- vluft 3y agoI dunno, if your crypto choices were just "the best thing that won't be included in FIPS" you would do pretty well; blake3, chacha20, 25519 sigs & dh...
- Retr0id 3y agoBlake3 is a clear winner for large inputs. However, for smaller inputs (~1024 bytes and down), the performance gap between it and everything else (blake2, sha256) gets much narrower, because you don't get to benefit from the structural parallelization. If you're mostly dealing with small inputs, raw hash throughput is probably not high on your list of concerns - In the context of a protocol or application, other costs like IO latency probably completely dwarf the actual CPU time spent hashing. If raw performance is no longer high on your list of priorities, you care more about the other things - ubiquitous and battle-tested library support (blake3 is still pretty bleeding-edge, in the grand scheme of things), FIPS compliance (sha256), greater on-paper security margin (blake2). Which is all to say, while blake3 is great, there are still plenty of reasons not to prefer it for a particular use-case.
- LegibleCrimson 3y agoHow does the extended output work, and what's the point of extended output? From what I can see, BLAKE3 has 256 bits of security, and extended output doesn't provide any extra security. In this case, what's the point of extended output over doing something like padding with 0-bits or extending by re-hashing the previous output and appending it to the previous output (eg, for 1024 bits, doing h(m) . h(h(m)) . h(h(h(m))) . h(h(h(h(m))))). Either way, you get 256 bits of security. Is it just because the design of the hash makes it simple to do, so it's just offered as a consistent option for arbitrary output sizes where needed, or is there some greater purpose that I'm missing?
- oconnor663 3y ago> From what I can see, BLAKE3 has 256 bits of security, and extended output doesn't provide any extra security. 128 bits of collision resistance but otherwise correct. As a result of that we usually just call it 128 bits across the board, but yes in an HMAC-like use case you would generally expect 256 bits of security from the 256 bit output. Extended outputs don't change that, because the internal chaining values are 256 bits even when the output is larger. > extending by re-hashing the previous output and appending it to the previous output It's not quite that simple, because you don't want later parts of your output to be predictable from earlier parts (which might be published, depending on the use case). You also want it to be parallelizable. You could compute H(m) as a "pre-hash" and then make an extended output something like H(H(m)|1)|H(H(m)|2)|... That's basically what BLAKE3 is doing in the inside. The advantage of having the algorithm do it for you is that 1) it's an "off the shelf" feature that doesn't require users to roll their own crypto and 2) it's slightly faster when the input is short, because you don't have to spend an extra block operation computing the pre-hash. > what's the point of extended output? It's kind of niche, but for example Ed25519 needs a 512 bit hash output internally to "stretch" its secret seed into two 256-bit keys. You could also use a BLAKE3 output reader as a stream cipher or a random byte generator. (These sorts of use cases are why it's nice not to make the caller tell you the output length in advance.)
- LegibleCrimson 3y agoThat makes sense. I hadn't thought about using that as a PRNG, but the idea is interesting to me. I might play around with it and profile it to see how these use cases play out. Implementing a BLAKE3-backed Rust rand::RngCore sounds like a fun little exercise, and would make it easy to profile compared to other PRNGs. Actually, looking at that trait right now, I see that there are already ChaCha implementations, so the concept is already being exercised in the same family. Thanks for the explanation. I'm far from a security expert, so more off-the-shelf bits at my disposal means fewer opportunities for me to accidentally implement security vulnerabilities by trying to do it myself.
- 15155 3y agoKeccak is my preference. Keccak is substantially easier to implement in hardware: fewer operations and no carry propagation delay issue because there's no addition.
- sylvain_kerkour 3y agoAt the end of the day, what really matters for most people is 1) Certifications (FIPS...) 2) Speed. SHA-256 is fast enough for maybe 99,9% of use cases as you will saturate your I/O way before SHA-256 becomes your bottleneck[0][1]. Also, from my experience with the different available implementations, SHA-256 is up to 1.8 times faster than Blake3 on arm64. [0] https://github.com/skerkour/go-benchmarks/blob/main/results/scaleway_amp2_c8.txt https://github.com/skerkour/go-benchmarks/blob/main/results/... [1] https://kerkour.com/fast-hashing-algorithms https://kerkour.com/fast-hashing-algorithms
- oconnor663 3y agoI mostly agree with you, but there are a couple other bullet points I like to throw in the mix: - Length extension attacks. I think all of the SHA-3 candidates did the right thing here, and we would never accept a new cryptographic hash function that didn't do the right thing here, but SHA-2 gets a pass for legacy reasons. That's understandable, but we need to replace it eventually. - Kind of niche, but BLAKE3 supports incremental verification, i.e. checking the hash of a file while you stream it rather learning whether it was valid at the end of the stream. https://github.com/oconnor663/bao https://github.com/oconnor663/bao. That's useful if you know the hash of a file but you don't necessarily trust the service that's storing it.
- jandrewrogers 3y agoI think SHA-256 is still marginal for speed in modern environments unless your I/O is unusually limited relative to CPU. Current servers can support 10s of GB/s combined throughput for network and storage, which is achievable in practice for quite a few workloads. Consequently, you have to plan for the CPU overhead of the crypto at the same GB/s throughput since it is usually applied at the I/O boundaries. The fact that SHA256 requires burning the equivalent of several more cores relative to Blake3 has been a driver in Blake3 anecdotally creeping into a lot of data infrastructure code lately. At these data rates, the differences in performance of the hash functions is not a trivial cost in the cases where you would use a hash function (instead of e.g. authenticated encryption). The arm64 server case is less of a concern for other reasons. Those cores are significantly weaker than amd64 cores, and therefore tend to not be used for data-intensive processing regardless. This allows you to overfit for AVX-512 or possibly use SHA256 on arm64 builds depending on the app. There is a strong appetite for as much hashing performance per core as possible for data-intensive processing because it consumes a significant percentage of the total CPU time in many cases. Due to the rapid growing scale, non-cryptographic hash functions are no longer fit for purpose much of the time.
- digger495 3y agoI'm holding out for BLAKE7, personally
- deleted 3y ago[deleted]
- aborsy 3y agoWhen it’s said SHA2 will remain secure in foreseeable future, are there estimates on the number of decades? The quantum computers apparently don’t help much with hash attacks, and SHA2 has received a lot of cryptanalysis.
- footlose_3815 3y agoI replaced sha with blake in my deduplication needs, and it sped up the comparisons by a factor of 4 at least. For my use case, it is great