7 ms·
Thanks, and I must say, very old school web page:)
by plafl 5y ago
Thanks, and I must say, very old school web page:)
- errantmind 5y agoAgner Fog's writings are all well worth the read if you care about optimization. Some of the techniques he uses in his optimized versions of common C functions are really interesting. For example, he uses SIMD in several of these, but one of the performance problems associated with SIMD is the last 'n' bytes. If you processing, say, 16 or 32 bytes at a time, what do you do when you get towards the end of your buffer and that process would over-read past the end of the buffer? Most code I've seen just uses a simple linear process for the last few bytes but this is often very slow. Agner's solution was to just read past the end of the buffer, but use some assembly tricks to ensure there is always valid memory available after the end of the buffer, that way reading past the end of the buffer would (almost) never cause a problem.
- jandrewrogers 5y agoAnother useful trick for reading the last few bytes, especially if suitable buffer padding is not possible and the read would cross a memory page boundary, is to use PSHUFB aligned against the memory page boundary to pull the last few bytes into a register. Safe and fast.
- burntsushi 5y agoFor the specific case of substring/byte searching, another useful trick here is to do an unaligned load that lines up with the end of the haystack. Some portion of that load will have already been searched, but you already know nothing is in it.
- jhgb 5y ago> but use some assembly tricks to ensure there is always valid memory available after the end of the buffer Isn't rounding up your buffer's size an allocation technique rather than "as assembly trick"?