3 ms·
We went through a lot of pain to get this right in Tamgu (https://github.com/naver/tamgu https://github.com/naver/tamgu). In particular, emojis can be encoded a
by clauderoux 7y ago
We went through a lot of pain to get this right in Tamgu (https://github.com/naver/tamgu https://github.com/naver/tamgu). In particular, emojis can be encoded across 5 or 6 Unicode characters. A "black thumb up" is encoded with 2 Unicode characters: the thumb glyph and its color.
This comes at a cost. Every time you extract a sub-string from a string, you have to scan it first for its codepoints, then convert character positions into byte positions. One way to speed up stuff a bit, is to check if the string is in ASCII (see https://lemire.me/blog/2018/05/16/validating-utf-8-strings-using-as-little-as-0-7-cycles-per-byte/ https://lemire.me/blog/2018/05/16/validating-utf-8-strings-u...) and apply regular operator then.
We implemented many techniques based on "intrinsics" instructions to speed up conversions and search in order to avoid scanning for codepoints.
See https://github.com/naver/tamgu/blob/master/src/conversion.cxx https://github.com/naver/tamgu/blob/master/src/conversion.cx... for more information.
- arcticbull 7y agoIt's not sufficient to use code points right? Some characters, for instance your example of the black thumbs up emoji, are grapheme clusters [1] composed of multiple code points. I think you have to iterate in grapheme clusters and convert that back to an offset in the original underlying encoding. If you just rely on code points you risk splitting up a grapheme cluster into (in your example) two graphemes, one in each sub-string, the left representing "black" and the right representing "thumbs up." Further, you actually need to utilize one of the unicode normalization forms to perform meaningful operations like comparison or sorting. This is one thing Rust's string API gets right, allowing you to iterate over a string as UTF-8 bytes in constant time -- and, by walking, your choice of codepoints and (currently unstable, or in the unicode-segmentation crate) grapheme clusters. Even that though is a partial solution. [2] Definitely a tough problem! [1] https://mathias.gaunard.com/unicode/doc/html/unicode/introduction_to_unicode.html https://mathias.gaunard.com/unicode/doc/html/unicode/introdu... [2] https://internals.rust-lang.org/t/support-for-grapheme-clusters-in-std/7339/4 https://internals.rust-lang.org/t/support-for-grapheme-clust...
- clauderoux 7y agoExactly my point. Most modern emojis cannot rely on pure codepoints to be extracted.
- arcticbull 7y agoMakes sense! Did you end up implementing normalization for your sub-string find, or did you work around it some other way? I couldn't seem to see it when skimming.
- clauderoux 7y agoYou can have a look on: s_is_emoji...
- clauderoux 7y agoIn https://github.com/naver/tamgu/blob/master/include/conversion.h https://github.com/naver/tamgu/blob/master/include/conversio..., I have implemented a class: agnostring which derives from "std::string". There are some methods to traverse a UTF8 string: begin(): to initialize the traversal end() : is true when the string is fully traversed next(): which goes to the next character and returns the current character. s.begin(); while (!s.end()) { u = s.next(); }
- arcticbull 7y agoVery cool, thanks!