4 ms·
I like how getting 46% gain over already optimized C code is considered "insignificant". Moore's law is over, folks, 46% is a pretty decent gain in my book, esp
by melted 11y ago
I like how getting 46% gain over already optimized C code is considered "insignificant". Moore's law is over, folks, 46% is a pretty decent gain in my book, especially in GCC where facilities are available to you to keep your code portable. Just use function multi versioning and define a canonical fallback.
- nly 11y ago46% gain over C is irrelevant if your program only spends 1% of its time decoding base64
- melted 11y agoSure, if it's just 1% it probably wasn't worth optimizing even that plain C code. But you don't know that. And you especially don't know that if it's a part of a library or the like.
- jcranmer 11y agoI thought about trying to measure the input on the real-world code that I have to use it in (specifically, decoding email messages). The challenge of trying to do base64 decoding fast there is that you generally have to ignore whitespace characters (and possibly other junk as well). The code implementation here clearly requires that all the extraneous stuff get kicked out and the resulting string be compacted, an operation whose overhead by itself could easily negate merely a 46% gain; it's hard to see how it could be modified to make the stateful decisions while still keeping the tight gain. I also have a sneaking suspicion that the improved scalar version might turn out to be worse in real-world systems because the 4KiB table would cause L1 cache thrashing if you're doing lots of small base64 decodings followed by later processing (e.g., string conversion). So the real problem is that the measurement is done in a way which is different enough from real-world scenarios that it's questionable if any speedup is actually achievable in those scenarios, let alone one as significant as a 46% gain. It's easy to be fast if you ignore having to do it correctly.
- jcranmer 11y agoTo follow up: I did do a quick modification of the test harness using a gigantic mbox file I had and some really crappy MIME parsing code. $ ./email input size: 86023056 (that's only base64-encoded text, this is a ~562MB file). improved scalar... 1.528 scalar... 1.564 (speed up: 0.98) SSE... 1.533 (speed up: 1.00) Null... 1.447 (speed up: 1.06) So about 5% of the test was spent doing base64 decoding, and the improved scalar had a 48% over regular scalar and SSE had a speedup of 40% over regular scalar on just the base64 section. So the short conclusion is that SSE doesn't really get any speedup on short strings. The additional test code, for full disclosure (yes, it's not correct, but this is easier to code and get some numbers rather than nothing. A more correct solution would really hurt the SSE bit more than other ones.): template <typename T> void run_email(T callback) { static std::string cte{"content-transfer-encoding"}; std::ifstream mbox("/tmp/test.mbox"); std::string line; bool inHeader = true; bool base64Body = false; while (std::getline(mbox, line)) { if (line[line.length() - 1] == '\r') line.pop_back(); if (inHeader) { std::transform(line.begin(), line.end(), line.begin(), ::tolower); if (line.find(cte) == 0 && line.find("base64") != std::string::npos) base64Body = true; if (line.empty()) inHeader = false; } else { if (line.substr(0, 5) == "From " || line.substr(0, 2) == "--") { inHeader = true; base64Body = false; continue; } if (base64Body) { while (line[line.size() - 1] == '=') line.pop_back(); while (line.size() % 16 != 0) line.push_back('A'); uint8_t *text = (uint8_t*)&line[0]; callback(text, line.size(), text); } } } }
- bmm6o 11y agoAnd on the other hand there are plenty of applications where you can assume that the encoded data doesn't need any pre-processing, either because your program was the one that did the encoding or because validation occurs on a higher level (e.g. a cryptographic MAC). The cache issue is a good point, I wonder how dedicated your process has to be to processing base64 text for that to not be a factor.
- wmu 11y ago"I like how getting 46% gain over already optimized C code is considered "insignificant". Author here (again!) I used SIMD in different algorithms and often speedup was greater than 2, 3, or more. So, from my skewed point of view speedup less than 2 isn't very impressive. :) But I buy your opinion, next time I'll try to be more enthusiastic.
- melted 11y agoSure, in lucky situations I've been able to speed things up by an order of magnitude by just changing memory layout and using SSE. But 46% in a tight spot is worthwhile change. Basically, when dealing with a highly performance intensive code (storage, if you must know), my rule of thumb was, 2% of more of improvement in overall end-to-end throughput on a long running benchmark is worthwhile if change is not too complicated. This could mean 10x improvement in one of the parts of the pipeline, or 2% improvement across the board, or anything in between: the goal was overall throughput. Likewise, nothing that regressed the performance could ever go in. You'd be surprised what you can squeeze out of your code if you establish these ground rules and run with them for a year.