3 ms·
> "ȃ" uses two bytes in UTF-8 "â" uses two bytes when encoded in UTF-8, while "ȃ" (which was what bzxcvbn supplied as an example and you pasted in the quoted
by tzot 4y ago
> "ȃ" uses two bytes in UTF-8
"â" uses two bytes when encoded in UTF-8, while "ȃ" (which was what bzxcvbn supplied as an example and you pasted in the quoted section) uses three bytes when encoded in UTF-8.
>>> s="ȃ"
>>> len(s)
2
>>> s.encode('utf8')
b'a\xcc\x91'
>>> import unicodedata as ud
>>> [ud.name(c) for c in s]
['LATIN SMALL LETTER A', 'COMBINING INVERTED BREVE']