3 ms·
> As pointed out elsewhere, you are wrong because UTF-16 'characters' can still be comprised of compounded elements. (Base character plus additional composition
by udp 8y ago
> As pointed out elsewhere, you are wrong because UTF-16 'characters' can still be comprised of compounded elements. (Base character plus additional composition elements to create a final character.) In simple terms, you can't just treat it as an array to index to any given character.
I am aware of that having implemented UTF-8 myself, and I thought I was fairly clear in my comment about characters vs. bytes.
> The assumption about finding specific characters (for example, command flags) in a string or array of strings is also... complicated.
As an example: https://en.wikipedia.org/wiki/Halfwidth_and_fullwidth_forms https://en.wikipedia.org/wiki/Halfwidth_and_fullwidth_forms
Also aware of that. But it only matters if you're actually searching for those characters. Parsing something like XML or JSON, therefore, is as easy with UTF-8 as it would be with ASCII, because you're only looking for characters like < or { which are the same byte values in UTF-8. You don't need to worry about accidentally finding a continuation byte because UTF-8 sets high bits on continuation bytes for exactly this reason.
- mjevans 8y agoIs this UTF-16(BE|LE) and UTF-8 mixed? I'm unclear about the context you're proposing. If you're trying to match for a Unicode character in either UTF-16 bytestream dialect then determining the matching format first is important. The mixed format just seems too crazy to touch.