4 ms·
> Imagine how fast a ChaCha ASIC could run Not as fast. Chacha20 uses 32-bit additions which are fast in software but expensive and slow in hardware. In additi
by dependenttypes 7y ago
> Imagine how fast a ChaCha ASIC could run
Not as fast. Chacha20 uses 32-bit additions which are fast in software but expensive and slow in hardware. In addition protecting Chacha20 from power analysis attacks is more difficult compared to AES.
> just generic SIMD
Constant-time AES with SSE2 is actually faster than the naive variable-time AES. See https://www.bearssl.org/constanttime.html#aes https://www.bearssl.org/constanttime.html#aes
In addition Chacha20 is not nearly as fast as AES when using the AVX-512 Vector AES instructions.
> but also Adiantum
Which uses AES (once per sector).
- karim 7y agoGenuinely curious, would you mind explaining why some operation can be fast in software but slow in hardware?
- kop316 7y agoI think the parent comment is saying it is fast in software on a modern CPU, but making tha into an ASIC would either be a) slow or b) expensive due to the 32-bit additions. IIRC (I can't find it right now), when NIST had the contest for AES, AES hhad to run on low power hardware in the late 90s/early 2000s. This required things like everything to be fast on an 8-bit microcontroller.
- dependenttypes 7y agoTo implement 32-bit + in hardware you need 31 full adders and one half adder, each of which uses multiple gates and depends on the result of the previous adder. Meanwhile + and bitwise and tend to take the same amount of cycles to be processed, and each cycle takes the same amount of time, see https://gmplib.org/~tege/x86-timing.pdf https://gmplib.org/~tege/x86-timing.pdf Chacha20 in hardware would not be any slower than chacha20 in software, but it would be slower than other algorithms which do not use 32-bit +.
- JoshTriplett 7y ago> To implement 32-bit + in hardware you need 31 full adders and one half adder, each of which uses multiple gates and depends on the result of the previous adder. This is not how CPUs typically implement addition, or other ALU operations. Carry-lookahead adders have existed since the 1950s: https://en.wikipedia.org/wiki/Carry-lookahead_adder https://en.wikipedia.org/wiki/Carry-lookahead_adder
- sls 7y agoThank you, I love this citation so much. > Charles Babbage recognized the performance penalty imposed by ripple-carry and developed mechanisms for anticipating carriage in his computing engines.
- Twirrim 7y ago> In addition Chacha20 is not nearly as fast as AES when using the AVX-512 Vector AES instructions. Note that Cloudflare opted for Xeon Silver chips that aren't good at AVX-512, unless doing pure AVX-512 operations.
- ZeroCool2u 7y agoAnd their 10th gen prod servers switched to AMD which, as far as I know, have SIMD support, but not AVX-512 support specifically.
- dr_zoidberg 7y agoThat is correct, Zen 2 doesn't support AVX512 (no AMD chip does).