4 ms·
Correct me if I'm wrong: UTF-16 can be faster if you're iterating through characters, but for parsing iterating through bytes is usually sufficient. It's fine
by udp 8y ago
Correct me if I'm wrong:
UTF-16 can be faster if you're iterating through characters, but for parsing iterating through bytes is usually sufficient. It's fine to use something like plain old strchr to search for a { or a < in a UTF-8 string. Characters only get really important when you want to print the string.
- Const-me 8y ago> It's fine to use something like plain old strchr to search for a { or a < in a UTF-8 string. Correct. > only get really important when you want to print the string. Print, split into fixed-length pieces, render, layout, typeset — GUI apps do that a lot with their strings. Even web browser do, probably that’s why JavaScript strings are UTF-16 despite web in general is mostly UTF-8, see section 6.1.4 “The String Type” on page 67: http://www.ecma-international.org/publications/files/ECMA-ST/Ecma-262.pdf http://www.ecma-international.org/publications/files/ECMA-ST...
- udp 8y agoVery true, but once you're at the level of complexity where you're rendering a GUI are you really concerned about cache invalidation due to incorrect branch prediction iterating strings? GUI rendering is very expensive for lots of reasons, but I don't believe that's one of them - even in the UTF-8 case. You're ultimately either drawing thousands of pixels or communicating over a (relatively compared to the CPU cache) slow bus with the GPU.
- Const-me 8y ago> You're ultimately either drawing thousands of pixels or communicating over a (relatively compared to the CPU cache) slow bus with the GPU. In modern software, GPU draws pixels. Before it does that, CPU lays out these glyphs. Because GPUs are ridiculously fast these days, the layout step is typically slower than painting. I use MS edge browser. The built-in profiler said this comments page took 14ms to layout and only 6ms to paint. This page contains just a few tiny images, the majority of the content is text. I think the layout step spent most of the time iterating over the characters, looking up glyphs and measuring various blocks of text on this page.
- mjevans 8y agoAs pointed out elsewhere, you are wrong because UTF-16 'characters' can still be comprised of compounded elements. (Base character plus additional composition elements to create a final character.) In simple terms, you can't just treat it as an array to index to any given character. There's also all the detriments that still apply. (Including UTF-16LE / UTF-16BE, BOM, not being able to concatenate two valid string sequences (blind) and always have a valid result, etc.) You're also incorrect, or at least not correctly describing the search operation. The assumption about finding specific characters (for example, command flags) in a string or array of strings is also... complicated. As an example: https://en.wikipedia.org/wiki/Halfwidth_and_fullwidth_forms https://en.wikipedia.org/wiki/Halfwidth_and_fullwidth_forms For /command/ flags the half and full width equivalent characters might want to be 'folded' back over the traditional ASCII namespace. Sometimes a specific loss of precision can be desirable.
- udp 8y ago> As pointed out elsewhere, you are wrong because UTF-16 'characters' can still be comprised of compounded elements. (Base character plus additional composition elements to create a final character.) In simple terms, you can't just treat it as an array to index to any given character. I am aware of that having implemented UTF-8 myself, and I thought I was fairly clear in my comment about characters vs. bytes. > The assumption about finding specific characters (for example, command flags) in a string or array of strings is also... complicated. As an example: https://en.wikipedia.org/wiki/Halfwidth_and_fullwidth_forms https://en.wikipedia.org/wiki/Halfwidth_and_fullwidth_forms Also aware of that. But it only matters if you're actually searching for those characters. Parsing something like XML or JSON, therefore, is as easy with UTF-8 as it would be with ASCII, because you're only looking for characters like < or { which are the same byte values in UTF-8. You don't need to worry about accidentally finding a continuation byte because UTF-8 sets high bits on continuation bytes for exactly this reason.
- mjevans 8y agoIs this UTF-16(BE|LE) and UTF-8 mixed? I'm unclear about the context you're proposing. If you're trying to match for a Unicode character in either UTF-16 bytestream dialect then determining the matching format first is important. The mixed format just seems too crazy to touch.