5 ms·
Transcoding Unicode strings at crazy speeds with AVX-512
- mcraiha 3y ago[flagged]
- 867-5309 3y ago[flagged]
- shanielh 3y agoFunny thing I just published a PR to RoaringBitmap and got answered by him. Small world.
- pvg 3y agodiscussed recently https://news.ycombinator.com/item?id=37197921 https://news.ycombinator.com/item?id=37197921
- deleted 3y ago[deleted]
- clausecker 3y agoThat article was about a Latin 1 to UTF-8 routine, while this one is about UTF-8 from/to UTF-16.
- clausecker 3y agoAuthor here. Please let me know if you have any questions.
- k0k0r0 3y agoGreat to see you here!
- amluto 3y agoI wish these libraries (simdutf, simdjson, etc) supported translation of possibly broken utf8 to valid utf8 with replacement of errors. IMO this ought to be available out of the box, just like Python’s decode(errors='replace'). This is extremely common — utf8 is the dominant codec, and plenty of applications receive supposedly-utf8 data from untrusted sources and need to display it or further process it without just giving up if it’s bad. Yes, I realize that this can be done by hand using simdutf’s API. No, users should not be expected to roll it themselves.
- clausecker 3y agoThe algorithms described in the paper all perform input validation and transcode until the first encoding error or until end of input, whichever comes first. They also tell you how much of the input they transcoded, so it's indeed quite easy to call the transcode function and then to manually fix up errors if it chokes. We had considered adding a variant that could automatically fix up malformed input, but as there are multiple mutually incompatible conventions for dealing with encoding errors, we decided to ignore this case for now. Perhaps a future version of simdutf may provide support for best-effort transcoding of errorneous input.
- amluto 3y agoI can indeed "quite easily" hack this up. But this is C++, and C++ string manipulation is infamously unsafe and error-prone, and I really have very little desire to hack this up, write tests, etc, when I'd rather not think about UTF-8 at all. For an example of how unsafe and error-prone C++ string manipulation is, from simdutf's very own documentation: > simdutf_warn_unused size_t convert_latin1_to_utf8(const char * input, size_t length, char* utf8_output) noexcept; Eww, this literally cannot be safe in the sense of "free from memory-safety errors when called carelessly" (like Rust or almost any GC language would expect). Look, takes a single pointer to an output array with no length or other limit. It overruns the buffer unless you get it exactly right. (ThePhD has a very on-point rant or two [0] about this, and ztd.text IMO (and in his opinion) gets this right.) Also, for all that simdutf is really quite fast, it's really not obvious how to transcode latin-1 to utf-8 in a single pass, because you apparently need to compute the length first. I think one of the great strengths of safer languages is that they fight back when you try to design an API like this. [0] https://thephd.dev/output-ranges https://thephd.dev/output-ranges