3 ms·
(I posted this comment on the blog as well) The poor performance of Turbo Boyer-Moore might be due to cache locality issues. If you jump from checking the end
by twp 16y ago
(I posted this comment on the blog as well)
The poor performance of Turbo Boyer-Moore might be due to cache locality issues. If you jump from checking the end of the needle to the start of the needle, then there's a good chance that you'll cause a cache miss - which has the same cost as several hundred instructions. A modern Turbo Boyer-Moore would probably work backwards within the same cache line before jumping to the start of the needle. A couple of carefully chosen cache-preload instructions could make Turbo Boyer-Moore. For a good overview of cache effects, google "gallery of processor cache effects".
- onan_barbarian 16y agoThis is a plausible and knowledgeable-sounding argument; its only flaw is that it not true. The only relevant 'several hundred instructions' cache miss on a Core 2 Duo is a L2 miss ("Last Level Cache" miss). The hardware pre-fetcher will pick up accesses to the data stream after the first 1-2 misses. No amount of minor local jumping around in Turbo-BM is going to affect this.
- twp 16y agoThanks!