5 ms·
> I realised when I implemented EdDSA for Monocypher that optimisations compound. I feel you're missing the whole point. It's immaterial whether anyone can ge
by simplotek 4y ago
> I realised when I implemented EdDSA for Monocypher that optimisations compound.
I feel you're missing the whole point.
It's immaterial whether anyone can get to optimizations that compound multiplicatively. The whole point is that halving something that costs nothing earns you nothing. That's the whole point. Go ahead and shave off that millisecond. Will anyone actually notice whether you add or remove that penalty? Odds are, not at all.
- loup-vaillant 4y agoPeople did notice. Quite a few happy users are glad signature verification took less than a second instead of more than 3. Or 30, if you compare to some of the alternatives. Others love the fact it uses 2KB of stack space instead of 5. Monocypher's speed was actually an important component in its success in the embedded market, even though I didn't explicitly target it initially (I was lucky my portability driven decisions made it a good fit there).
- tharkun__ 4y agoNot your parent but I think that is exactly the point lots of people here are making. There definitely are niches where there are quite a few performance optimization opportunities that users do care about. In your example making something a user is actively waiting for go from 3 seconds to less than one is a great optimization target. What is not a great optimization target is making something the user is actively waiting for and that takes 30ms take 25ms instead. That's wasted money on developer time. If your "user" is a developer of embedded software with memory constraints and using your library leaves them more room that's awesome. If your user was someone using the library on a general purpose computing device with loads of memory then the 2 vs. 5 does nothing.
- rerdavies 4y agoYou need to do some research on how to get modern C/C++ compilers to vectorize.;-) No assembly required, and not that hard to restructure code. (But MUCH easier in C++).
- anonymoushn 4y agoSo far it seems like if you have some parsing or formatting task that can be trivially vectorized, the compiler will never do that, and you absolutely must use intrinsics.
- loup-vaillant 4y agoI tried auto-vectorisation, and it worked pretty well. But the code became just as big as using intrinsics would have (that with explicitly unrolling loops and rearranging things in memory), and intrinsics generated code that was easily 35% faster. I decided not put it off for later, and keep things simple for now.