3 ms·
Wise move, in my opinion. I hope Python will do that too, eventually.
by faragon 8y ago
Wise move, in my opinion. I hope Python will do that too, eventually.
- kevin_thibedeau 8y agoPython implemented multi-representation strings in 3.3. It picks the best encoding from ASCII, UTF8, UCS2, or UCS4. UTF8 is the preferred encoding for strings exposed to the C API. https://www.python.org/dev/peps/pep-0393/ https://www.python.org/dev/peps/pep-0393/
- masklinn 8y agoCPython provides but does not use UTF8 internally, it's a cache for PyUnicode_AsUTF8 (formerly _PyUnicode_AsString), which exists to convert Python strings to char* in order to interoperate with C libraries. It is not the canonical representation of the string. pypy, on the other hand, made this change in the latest release (7.1.0): https://twitter.com/pypyproject/status/1095971192513708032 https://twitter.com/pypyproject/status/1095971192513708032
- ken 8y agoThe "kind" field is one of [uninitialized, Latin-1, UCS-2, UCS-4]. I'm not sure how it would pick UTF-8. (In memory usage, the Python way is strictly worse: adding one 4-byte character to an ASCII string forces the entire string to be upgraded to UCS-4. The advantage of the Python way is O(1) access to codepoints, but that's an operation that rarely comes up in practice, when dealing with Unicode strings.) I'm not sure how that would be compatible with Python, either. Either you're giving up string[index], or you're causing this one subscript operation to take O(n) time, or you need a new index type just for strings (which is what Swift does). None of these seem very 'Pythonic'. EDIT: Apparently recent versions of pypy solve this by computing their own index into your string, which seems like it would be terrible for memory usage, but I look forward to their blog post about it.
- sambe 8y agoWhat do you hope Python will do?