4 ms·
C and C++ are so widely used that transitions like this are made not at the language level but at the level of platforms or other communities. Some parts of the
by bdarnell 11y ago
C and C++ are so widely used that transitions like this are made not at the language level but at the level of platforms or other communities. Some parts of the C/C++ world made this transition relatively seamlessly, while others got caught in the same traps as Python.
The key is UTF-8: UTF-8 is a superset of 7-bit ASCII, so as long as you only convert to/from other encodings at the boundaries of your system, unicode can be introduced to the internal components in a gradual and mostly-compatible way. You only get in trouble when you decide that you need separate "byte string" and "character string" data types (which is generally a mistake: due to the existince of combining characters, graphemes are variable-width even if you're using strings composed of unicode code points, so you don't gain much by using UCS-4 character strings instead of UTF-8 byte strings).
My theory is that the python 3 transition would have gone much smoother and still accomplished its goals if they had left the implicit str/bytes conversion in place but just made it use UTF-8 instead of ASCII (although in environments like Windows where UTF-16 is important this may not have worked out as well).
- dietrichepp 11y agoYou are correct that there's no real benefit in UTF-32 over UTF-8, which is why Go and Rust (and others) have worked with UTF-8 in memory just fine. However, the actual encoding of a str object is irrelevant, and it's not the point. The whole point of the str/bytes difference in Python is that you make Python keep track of whether you've done the conversion or not. In Python 2, you can be sloppy, and the programs are buggy as a result! You're putting the cart before the horse here in terms of the Python 3 transition. The str/unicode fix was one of the driving factors for Python 3 to exist in the first place, and if you removed it, then what's the point? Again, look at Go or Rust. Both of them have separate types for strings and bytes, even though they have the same representation in memory, and as a result we don't have that kind of bug in our program.
- bdarnell 11y ago> The str/unicode fix was one of the driving factors for Python 3 to exist in the first place, and if you removed it, then what's the point? The reason python 3 exists is that most python 2 code had latent unicode-related bugs that would only manifest when they were exposed to non-ascii data. The backwards-incompatible barrier between str and bytes was the solution the python 3 team chose for this problem; adopting utf-8 as the standard encoding would have been another solution which I claim would have been more backwards-compatible (essentially moving to the go/rust model, which prove that you don't necessarily need separate byte and character string types for correct unicode handling).
- hetman 11y agoBut not every 8-bit byte string is valid UTF-8 so that could still cause a world of pain.
- e12e 11y agoBut apart from being "Internet compatible", it makes little sense to move from ascii only to utf-8. It only makes sense if you're already trying to deal with stuff that doesn't fit in ascii. Don't get me wrong, I think ascii was a terrible hack (I seem to recall one of the designers called it his biggest mistake/regret - and that they should've gone with some kind of prefix-encoding). But the thing with unicode-everywhere is that it's more beginner friendly: you can have unicode in your variable names, and strings without giving it another thought. $ cat u.py from __future__ import print_function å="æ" print(å) $ python2 u.py File "u.py", line 2 SyntaxError: Non-ASCII character '\xc3' in file u.py on line 2, but no encoding declared; see http://python.org/dev/peps/pep-0263/ for details $ python3 u.py æ Sure, the explanation of what's wrong is right there in the error-message, but it seems a bit of an unnecessary hurdle for people to get around to write important programs, that ask you to type in your name, and then prints out the name ten times ;-)
- marshray 11y agoThis is all true, but (speaking as one) I think the deeper reason is that C and C++ developers are just already used to great pain associated with manipulating character data, whereas scripting language developers expect these things for free.