3 ms·
C0 and C1 can do so though according to the author. > because all 4 of 2-byte UTF-8's mandatory content bits lie in the first-and-final start byte, we can expl
by gzitscrux 15d ago
C0 and C1 can do so though according to the author.
> because all 4 of 2-byte UTF-8's mandatory content bits lie in the first-and-final start byte, we can explicitly rule out 11000000 (0xC0) and 11000001 (0xC1) as permanently invalid bytes. They will never ever appear anywhere in a valid UTF-8000 code unit!