3 ms·
You can. The last byte of an emoji is never the same as 'a'. UTF-8 is self-synchronizing, a trailing byte can never be misinterpreted as the start of a new cod
by ynik 3y ago
You can.
The last byte of an emoji is never the same as 'a'. UTF-8 is self-synchronizing, a trailing byte can never be misinterpreted as the start of a new codepoint.
This makes `memmem()` a valid substring search on UTF-8! With most legacy multi-byte encodings this would fail, but with UTF-8 it works!
- pezezin 3y agoAssuming that your strings are normalized, otherwise precomposed characters will not match decomposed characters.