4 ms·
You want a benchmark that shows that one >>, one &, two [] and a + are going to be fast? The test loop that wraps the lookup code will likely take as much CPU
by apankrat 6y ago
You want a benchmark that shows that one >>, one &, two [] and a + are going to be fast?
The test loop that wraps the lookup code will likely take as much CPU as what's being benchmarked. Even if you are to use RDPMC and even if you are to use it very sparingly.
- yorwba 6y agoThe naive solution uses just one [], so it's reasonable to ask whether the compressed solution is faster. (Especially for the higher compression levels with even more memory accesses.) And the relevant workload to benchmark wouldn't be just case-converting a single character, but a whole text, as would be done e.g. for constructing a case-insensitive search index. Because most texts contain only a limited range of characters, it's entirely possible that a large lookup table works just fine because only a tiny portion of it needs to fit into the cache.
- Asooka 6y agoI'm fairly confident that since the lookup table fits comfortably in L1 cache, both algorithms will be about equally fast. You may see a difference if you have to case-fold the entire Library of Congress several times per user operation. The other case where there may be meaningful difference in performance would be embedded devices with small caches and slow memory.
- vardump 6y agoIndirect memory lookups are slow, even if the data is in the L1 cache.