3 ms·
> this might save megawatts of power world wide That'd sure be neat! I really have no idea, though. > A small nitpick: not sure whether those macros bought an
by zwegner 7y ago
> this might save megawatts of power world wide
That'd sure be neat! I really have no idea, though.
> A small nitpick: not sure whether those macros bought anything, I guess the optimizer could inline function calls to intrinsic wrappers just fine as well?
Yeah, that's a bit weird. It's my hacky generics-in-C method, that allows both AVX2 and SSE4 implementations to be compiled in one codebase, while keeping the base algorithm clean. The idea is that you can compile both versions by including multiple times, like so:
#define AVX2
#include "z_validate.c"
#undef AVX2
#define SSE4
#include "z_validate.c"
#undef SSE4
...which would allow a later runtime dispatch based on CPUID, etc.
- thechao 7y agoHow much faster than a "naive" scalar implementation is this?
- zwegner 7y agoIt depends on the implementation. A good starting point would be looking at the various implementations benchmarked here: https://github.com/lemire/fastvalidate-utf-8 https://github.com/lemire/fastvalidate-utf-8 It looks like the main naive validator there (validate_utf8) clocks in at 1.25 cycles/byte for ASCII and 11 cycles/byte for UTF-8, which is respectively a 15.8x and 41.5x slowdown from my code.