3 ms·
FWIW, in Linux, this problem does not exist. Everything is UTF-8 and Python 2 would work just fine (and always did). In order to support Windows better, Python
by dannymi 4y ago
FWIW, in Linux, this problem does not exist. Everything is UTF-8 and Python 2 would work just fine (and always did).
In order to support Windows better, Python 3 introduced support for UCS-4 (or worse, UTF-16) strings (depending on a compilation setting when Python was compiled) and they had to introduce extra string types to distinguish readable strings from binary strings ("bytes").
These extra types made Python 3 a lot harder to teach (I teach 30 person classes every year).
So it's not all roses now.
In the end, I got used to it, BUT I just gave up asking encode()/decode() questions at the exams. Very few people understand it, or care enough (and I understand why--it's a ridiculous thing to have). You only need it if your OS somehow slept through the introduction of UTF-8, which is backward compatible with ASCII, resilient even if there are transfer errors and can encode all unicode characters.
Encoding problem used to be really common in UNIX (and before that, in mainframes), but with the introduction of UTF-8, all encoding problems I had vanished and never appeared again.
Even Windows 10 has an UTF-8 mode now and the Windows API functions that end in "A" can be made to use UTF-8.
Now, in a sense, Python 3 has this entire complication for no reason.
That said, Python 3 is ok to use now--and, conceptually, distinguishing byte strings from unicode strings is better (for example so that you don't accidentially print the former to the terminal). It just uses up brain cycles that you could be using for solving your actual problems.
- thrdbndndn 4y ago> I just gave up asking encode()/decode() questions at the exams. Very few people understand it, or care enough (and I understand why--it's a ridiculous thing to have). I get it from the "pass the exam" perspective, since that's one more thing to worry about. But from my experience in teaching others, doing the conversion between bytes and string implicitly (à la Python 2's way) hinders actual understanding of this very important concept, and it's quite harmful in further study. Bytes should be considered as a separate, more low-evel thing, away from int/float/strings; at the very least, it should be considered as bits/hex numbers. If you want strings, you explicitly encode/decode them in a way, even if everything is UTF-8. On top of that, "byte string" is just a confusing concept. It might works for English speaker (by "it's a ridiculous thing to have" I assume you mean that, "'english'.encode() is just b'english', why bother?"), not at all for Chinese speakers, even in UTF-8. There is no b'中文' -- only b'\xe4\xb8\xad\xe6\x96\x87' which has zero meanings in their own. And even from an easy-to-use perspective: most people don't even work on bytes often nowadays. A more abstract "string" type is all they need, without worrying about how it works under the hood (and if they do, they need to understand how encode/decode works properly anyway).
- dannymi 4y ago>doing the conversion between bytes and string implicitly There was no conversion. `bytes` and `str` were the same type. http://docs.python.org/whatsnew/2.6.html#pep-3112-byte-literals http://docs.python.org/whatsnew/2.6.html#pep-3112-byte-liter... says: > Python 2.6 adds bytes as a synonym for the str type, and it also supports the b'' notation. I just checked in Python 2.7: >>> bytes is str True >>> print("Hänsel") Hänsel >>> "Hänsel" 'H\xc3\xa4nsel' I'm working with Germans, Japanese and Polish that use a lot of special characters, including Kanji, umlauts, extra quote characters etc. I need the non-ASCII parts and had no problem with them in Python 2 on Linux (now, C++ libraries that reinvented their own string classes: many problems; C libraries: no problems). The point is when bytes is str, everything works just fine in Python 2 Linux with UTF-8 locale (which are used in all modern Linux distributions). No need to have a distinction between bytes and str. That how the rest of the OS works, too. Even a lot of Gtk, Glib and so on (for example the GNOME desktop environment) assume that you are in an UTF-8 locale for file names, for example. > A more abstract "string" type is all they need, without worrying about how it works under the hood (and if they do, they need to understand how encode/decode works properly anyway). Ehh, we had students write drivers for measurement apparatuses and they all used Python 2 str (without being prompted to do so). No encode or decode anywhere. Of the students, almost no one who tried Python 3 for that stayed with it (instead they were using Python 2). There was just no upside for this use case. I agree that, long term, having a distinction str vs bytes makes sense. But then you ARE juggling things that the OS doesn't need--it's basically busywork in Linux. I'm not trying to minimize your experience--but I don't think it would happen if you tried python2 on Linux today. Not sure it was worth it breaking compat for that.
- orf 4y ago> FWIW, in Linux, this problem does not exist. Everything is UTF-8 and Python 2 would work just fine (and always did). That's not true at all. I remember all kinds of encoding errors when dealing with the FS, the network or any user input when using Linux. Unless you're talking specifically about IDLE?