3 ms·
In CPython 2, a `str` was bytes period; no encoding was enforced, these bytes were not necessarily valid UTF-8, and a `unicode` had either UTF-16 or UTF-32 in-m
by tzot 5y ago
In CPython 2, a `str` was bytes period; no encoding was enforced, these bytes were not necessarily valid UTF-8, and a `unicode` had either UTF-16 or UTF-32 in-memory representation (decided at compilation time). In CPython 3, a `bytes` is bytes and a `str` has ASCII/UTF-16/UTF-32 in-memory representation (decided at runtime).
So CPython 2 strings were not “bytes of UTF-8”, and that is why strings `'\xfc\xf7\xe9'` worked fine, they even printed fine in the proper environments.
- mistrial9 5y ago> In CPython 2, a `str` was bytes -- like old C language, a str is really bytes, no encoding is enforced.