11 ms·
Validating UTF-8 bytes using only 0.45 cycles per byte (AVX edition)
- the_clarence 8y agoI see a lot of applications trying to take advantage of SIMD, but what when you try to run them on systems that don't support these instructions? My guess is that you need to write multiple files taking advantage of different sets of instructions and then dynamically figure out which to use at runtime with cpuid, but isn't that cumbersome and a way to inflate a codebase dramatically?
- en4bz 8y agohttps://gcc.gnu.org/wiki/FunctionMultiVersioning https://gcc.gnu.org/wiki/FunctionMultiVersioning
- deleted 8y ago[deleted]
- jmgrosen 8y agoGenerally speaking, I think if you care enough about performance to write manual SIMD code, being a little more cumbersome is a tradeoff you’re willing to make.
- why_only_15 8y agoIn my understanding when you use intrinsics and build for a processor without support for the intrinsics then GCC for example will replace it with equivalent code.
- eesmith 8y agoThat is true. Here's a couple of negatives. First, you still need to build once for each architecture, either as different executables, or as different object files, and provide some dispatch mechanism to use the right one based on what hardware is available. Second, if the intrinsics aren't built-in then there may be faster alternatives than using the GCC emulated version.
- mcbain 8y agoUnfortunately, no. That is the case with GCCs __builtin functions. With a few exceptions, intrinsics are basically macros for inline asm that the compiler can reason about. If on x86-64 you use a _mm256* intrinsic and compile without AVX support you just get a compile error, not a pair of equivalent SSE instructions.
- rurban 8y agoEven worse. You mostly get run-time errors when the built machine supported that feature, your machine doesn't, and the features aren't separated into multiversioning or loading different shared libs.
- oconnor663 8y ago> inflate a codebase dramatically This is usually only done for very specific algorithms. Unicode validation, hash functions, things like that. Unless you have an absolutely tiny application (which you might, if you're some kind of microcontroller), it's going to be a small percentage of your overall code size.
- londons_explore 8y agoIn a microcontroller, I don't think you'll be needing AVX2...
- Rebelgecko 8y agoI'm not sure where exactly the line is drawn between a microcontroller and a CPU, but even some of the lower end ARMs support SIMD instructions.
- wmu 8y agoSpeaking of the Intel world it's not that bad. There are three major version right now: SSE4.1, AVX and AVX2 (AVX512 is not popular yet). In the past (roughly 10 years ego) it was a problem, as there were: MMX, SSE, SSE2, SSE3, SSSE3, SSE4.1, SSE4.2, XOP, 3DNow and perhaps a few more extensions. it's not a typo, there are three 'S' :)
- wmu 8y agoSorry, I forgot that in HN comments the asterisk char is an italics indicator. There should be a mark after SSSE3.
- saagarjha 8y agoDarwin platforms ship binaries with different slices for different versions of Intel processors. You have the generic x86_64 and the newer x86_64h which supports more features.
- akarambir 8y agoWhat does linux utilities like sed, awk use for text manipulation because they were very slow when I was changing a few table names in a sql file.
- zorked 8y agoI don't think they use anything in common. Try to set your locale to "C" as otherwise string comparisons will do extra work handling your locale's notions of equivalent characters.
- masklinn 8y agoNote that this and that are not necessarily related: you're talking about performing unicode-aware text matching and manipulation, TFA is solely about validating a buffer's content as UTF-8.
- akx 8y agoHow slow? On my 2013 MBP, `gsed` (sed from coreutils) can do a replacement like that at about 350 MiB/s (of which most seems to be spent writing to disk, since writing to /dev/null hikes it up to 800 MiB/s).
- akarambir 8y agoIt was sed substitute command on a ~800Mb file on Thinkpad T470 with SSD. It was taking around 40-50 sec for each substitution. Though as others have pointed, it may not be directly related to article in discussion.
- coldtea 8y ago>It was taking around 40-50 sec for each substitution. Substitution should not be really a relevant metric as it wouldn't influence the result much. Sed/Awk will still have to go through the whole file to find all occurrences they should substitute (and when they do find an occurrence, the substitution would take nanoseconds). The size of the file is a better metric (e.g. how many seconds for that 800mb in total). Also, whether you used regex in your awk/sed, and what kind. A badly written regex can slow down search very much.
- kissiel 8y agoI wonder about the Joules per byte. AFAIK AVX units are quite expensive energy-wise.
- masklinn 8y agoDon't they also tend to work at a lower clock due to their higher energy requirements? edit: though this is AVX2 ("AVX-256") rather than AVX-512, and Lemire has covered AVX and the possibility of throttling (with or without AVX) in the past so they're probably aware of the potential issue and consider that they either won't get triggered or the gain is good enough to compensate the lower frequency.
- kissiel 8y agoNice. So I understand that AVX2 is not bringing the CPU's clock down. Got any sources for power consumption figures/comparisons of those AVX units?
- lorenzhs 8y agoHeavy use of complex AVX2 operations causes downclocking, too, but typically less so than AVX-512. More details are documented in https://en.wikichip.org/wiki/intel/frequency_behavior https://en.wikichip.org/wiki/intel/frequency_behavior -- also see e.g. https://en.wikichip.org/wiki/intel/xeon_gold/6138#Frequencies https://en.wikichip.org/wiki/intel/xeon_gold/6138#Frequencie... for an example how the frequencies differ depending on the number of active cores. I think the reason for reducing clock speed when vector units are in heavy use is to keep power usage in check. You might also find https://blog.cloudflare.com/on-the-dangers-of-intels-frequency-scaling/ https://blog.cloudflare.com/on-the-dangers-of-intels-frequen... helpful, which goes into detail about a specific case where dynamic frequency scaling resulted in AVX-512 code running slower than AVX2 code.
- masklinn 8y agoAnd here are some of Lemire's own posts on the subject: * https://lemire.me/blog/2018/04/19/by-how-much-does-avx-512-slow-down-your-cpu-a-first-experiment/ https://lemire.me/blog/2018/04/19/by-how-much-does-avx-512-s... * https://lemire.me/blog/2018/08/13/the-dangers-of-avx-512-throttling-myth-or-reality/ https://lemire.me/blog/2018/08/13/the-dangers-of-avx-512-thr... * https://lemire.me/blog/2018/08/15/the-dangers-of-avx-512-throttling-a-3-impact/ https://lemire.me/blog/2018/08/15/the-dangers-of-avx-512-thr... * https://lemire.me/blog/2018/08/24/trying-harder-to-make-avx-512-look-bad-my-quantified-and-reproducible-results/ https://lemire.me/blog/2018/08/24/trying-harder-to-make-avx-... * https://lemire.me/blog/2018/08/25/avx-512-throttling-heavy-instructions-are-maybe-not-so-dangerous/ https://lemire.me/blog/2018/08/25/avx-512-throttling-heavy-i... * https://lemire.me/blog/2018/09/04/per-core-frequency-scaling-and-avx-512-an-experiment/ https://lemire.me/blog/2018/09/04/per-core-frequency-scaling... * https://lemire.me/blog/2018/09/07/avx-512-when-and-how-to-use-these-new-instructions/ https://lemire.me/blog/2018/09/07/avx-512-when-and-how-to-us...
- bradleyjg 8y agoUnder the new string model in java > 8 a fairly frequent workflow is: 1) get external string 2) figure out if it is UTF-8, UTF-16, or some other recognizable encoding 3) validate the byte stream 4) figure out if the code points in the incoming string can be represented in Latin-1 5) instantiate a java string using either the Latin-1 encoder or the UTF-16 encoder I know some or all of these steps are done using hotspot intrinsics, and then the JIT/VM does inlining, folding and so on, but I wonder how fast a custom assembly function to do all these steps at once could be.
- Twirrim 8y agoYou might be interested in his blog on the same subject a few days ago: https://lemire.me/blog/2018/10/16/validating-utf-8-bytes-java-edition/ https://lemire.me/blog/2018/10/16/validating-utf-8-bytes-jav...
- adamretter 8y agoIf you are given the external string as bytes, which is all you can have if you don't know the encoding. Then steps 2,3,4 can all be done as one step I would have thought. Something like - https://github.com/adamretter/utf8-validator/blob/optimize-utf8-validation/src/main/java/uk/gov/nationalarchives/utf8/validator/Utf8Validator.java#L196 https://github.com/adamretter/utf8-validator/blob/optimize-u...
- jwilk 8y agoPrevious blog post on HN: https://news.ycombinator.com/item?id=17081571 https://news.ycombinator.com/item?id=17081571