4 ms·
Hmm. I'm not convinced any of your routines are more "high-performance" than the standard-library, have you benchmarked them? For example, I don't think your g
by fredericoqq 9y ago
Hmm. I'm not convinced any of your routines are more "high-performance" than the standard-library, have you benchmarked them?
For example, I don't think your getorder() could possibly be faster than __builtin_ctz(), or even a wrapper around ffs(). You also hardcoded the size of size_t to 32 bits, so it's not more portable than those options.
Edit: I checked, ffs is a tiny branch-free routine in glibc:
(gdb) disassemble ffs
Dump of assembler code for function ffsl:
0x0007e9d0 <+0>: mov $0xffffffff,%edx
0x0007e9d5 <+5>: bsf 0x4(%esp),%eax
0x0007e9da <+10>: cmove %edx,%eax
0x0007e9dd <+13>: add $0x1,%eax
0x0007e9e0 <+16>: ret
End of assembler dump.
- apjana 9y ago> __builtin_ctz avoided it to stay out of compiler-specific stuff. we do have plans to replace getorder() with ffsl(). Perhaps I would be more accurate if I say nnn is faster by design. Movement of data around memory is minimal. No redundant bytes are allocated. We use quicksort and optimize further by pushing non-matches down right away so they never appear in a filter comparison again. And of course, using non-lib custom functions enable using static linkage, having a controlled binary size and removing redundant checks/processing because the limits and borderline cases are known.
- deleted 9y ago[deleted]