3 ms·
My first thought on seeing this was Bjoern Hoehrmann's UTF-8 decoder, which encodes a finite state machine into a byte array: http://bjoern.hoehrmann.de/utf-8/
by SloopJon 8y ago
My first thought on seeing this was Bjoern Hoehrmann's UTF-8 decoder, which encodes a finite state machine into a byte array:
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ http://bjoern.hoehrmann.de/utf-8/decoder/dfa/
I happened to come across this decoder again recently in Niels Lohmann's JSON library for C++:
https://github.com/nlohmann/json https://github.com/nlohmann/json
I see that this is mentioned in a previous post:
https://lemire.me/blog/2018/05/09/how-quickly-can-you-check-that-a-string-is-valid-unicode-utf-8/ https://lemire.me/blog/2018/05/09/how-quickly-can-you-check-...
One thing I'd like to check in the new code is whether it's as picky about things like overlong sequences as Bjoern's code is.
- loeg 8y ago> One thing I'd like to check in the new code is whether it's as picky about things like overlong sequences as Bjoern's code is. I believe that's what checkContinuation() is doing, based on its use of the "counts" parameters. I don't understand how it works, but I don't see any other reason for count_nibbles() to compute the '->count' member.
- kwillets 8y agoThe counts are actually to find the length of the following continuation bytes, and mark them to look for overlaps or underlaps between the end of one code point and the next.
- loeg 8y agoUnderlaps and overlaps would occur if sequences are invalid, i.e., overly long (or short) compared to the initial byte's prefix.
- kwillets 8y agoIt checks overlongs. It's not too hard -- a few bits in the first byte, and a few in the second for some lengths.