3 ms·
> for instance, Python now uses a weird mixture of ASCII and UCS2 internally. Really? Last I heard (PEP 393), the rule was: "8 bits if all codepoints are less
by rav 5y ago
> for instance, Python now uses a weird mixture of ASCII and UCS2 internally.
Really? Last I heard (PEP 393), the rule was: "8 bits if all codepoints are less than 256 (i.e. Latin-1); 16 bits if all codepoints are less than 2^16 (i.e. BMP); otherwise 32 bits". This means that text with all Latin-1 characters (which are approximately the first 256 codepoints of Unicode) will be stored internally as, well, Latin-1. This implies that ASCII strings are stored as ASCII.
- qalmakka 5y agoYep, that's what I meant. They either use ASCII, UCS2 or UCS4 depending on the type of string. Doesn't make a lot of sense to me, but I guess they couldn't just throw 16 bit chars away.
- chronial 5y agoI'm not 100% sure on this, but I don't backwards-compatibility mattered in this decision. They wanted memory-efficient and O(1)-indexable strings.