4 ms·
I'm interested in if there's a simple design for a ChaCha/Poly1305 accelerating ISA extension for RISC-V (outside the general crypto extension, not sure even as
by microcolonel 6y ago
I'm interested in if there's a simple design for a ChaCha/Poly1305 accelerating ISA extension for RISC-V (outside the general crypto extension, not sure even as a member of that where it is going).
I feel like it is a lot simpler, if for no other reason than there being effectively a single mode for the ChaCha primitive rather than three or more. The whole operation is composed of bitwise rotations and adders that can be arranged in a static network, and (as far as I know) is more or less not a source of timing sidechannel information one way or another.
Maybe I could try banging my head against the XCrypto repository this weekend.
- brandmeyer 6y agoARX ciphers suffer somewhat on vanilla RISC-V due to the lack of native rotations. You have to synthesize rotations from shifting and logical ops. Prime field algorithms suffer somewhat on vanilla RISC-V due to the lack of good bignum support. For example, there is no add-with-carry, nor a convenient way to get the carry-out at a low level. You have to use the set-if-less-than instruction after an addition to separately compute the carry bit.
- a1369209993 6y ago> For example, there is no add-with-carry, nor a convenient way to get the carry-out at a low level. IIUC, you're supposed to break the bignum up into (say) 56-bit chunks, and use the upper bits of the register as the carry. I'm sceptical that that works as well as it should for practical bignum applications, though. (I haven't had occasion to try it out.)
- microcolonel 6y agoI think this becomes less terrible the wider your bignum is.
- dependenttypes 6y agomicrocolonel was interested for an ISA extension rather than software implementation.
- brandmeyer 6y agoISA extensions cover quite a lot of ground. One design point is to implement a very narrowly scoped accelerator that only addresses the highest latency (or lowest work-per-instruction) parts of a software implementation. Both Chacha and Poly1305 were specifically designed for fast execution in software implementations, so IMO this approach is well-suited to those algorithms. You could easily argue that the ARMv7-M instruction "unsigned multiply accumulate accumulate long" exists solely as a bignum accelerator, for example. Perhaps fusing addition with xor is similarly worthwhile for ChaCha.