6 ms·
> the text model of Python 2 is a giant mess and makes it very hard to correctly deal with non-ascii text for any non-trivial software There are counter-argume
by sirclueless 10y ago
> the text model of Python 2 is a giant mess and makes it very hard to correctly deal with non-ascii text for any non-trivial software
There are counter-arguments to this. Armin Ronacher, author of (among other software) the excellent Flask web framework, thinks that Python 2's system of codecs and byte streams is better in practice [1][2]. Reasons include: You can do byte -> byte conversions with codecs that are no longer possible. You can better handle text encodings besides UTF-8 (and here he describes several embarassing failures of Python 3 to handle OS paths correctly). You can write single APIs that handle byte streams like gzip and text encodings like UTF-8.
[1]: http://lucumr.pocoo.org/2011/12/7/thoughts-on-python3/ http://lucumr.pocoo.org/2011/12/7/thoughts-on-python3/
[2]: http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/ http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/
- dom0 10y agoI don't think it's honest to post this without context, and without mentioning that in the five years that passed most of these things were remedied, and indeed, some things were already remedied at the time of his writing. Some points Armin makes are valid and remain valid for Linux-ish systems, but have been shown and refuted countless times for other operating systems; Python is not a Linux-only show. I won't re-iterate all that here.
- masklinn 10y ago> Armin Ronacher, author of (among other software) the excellent Flask web framework, thinks that Python 2's system of codecs and byte streams is better in practice. Armin Ronacher works in a very specific context of having to deal with byte/text interfaces in pretty much all his projects, and while I can see where he comes from I work at a different level and at the level at which I work the P2 model is a giant pain in the ass. > [1] http://lucumr.pocoo.org/2011/12/7/thoughts-on-python3/ http://lucumr.pocoo.org/2011/12/7/thoughts-on-python3/ http://lucumr.pocoo.org/2016/11/5/be-careful-about-what-you-dislike/ http://lucumr.pocoo.org/2016/11/5/be-careful-about-what-you-... Armin is no foe of Python 3. And as noted in the essaye Python 3 has undergone several improvements or features reintroductions e.g. PEP 461 reintroduced C-style formatting to bytestrings, making generating binary data (especially ascii-based formats) significantly more convenient than it is between 3.0 and 3.4. Also note that Armin has repeatedly praised Rust's text model, which is much more similar to P3's than P2's (except with static types and no messy legacy). > and here he describes several embarassing failures of Python 3 to handle OS paths correctly And (fucking surprise) the issue with that is the text model of FS paths is an embarrassing pile of garbage, Python 2 is convenient because it doesn't try to touch that mess at all and just hands the flaming bag of shit to whoever comes next.
- the_mitsuhiko 10y ago> Also note that Armin has repeatedly praised Rust's text model, which is much more similar to P3's than P2's (except with static types and no messy legacy). That is incorrect. Rust's text model has (almost) free (and copyless) transmutes from bytes to strings. Python does not. The text model of rust is much closer to Python 2 than 3 in many ways.
- masklinn 10y ago> That is incorrect. […] The text model of rust is much closer to Python 2 than 3 in many ways. Rust's text model strictly separates proper strings and bytestrings, defaults to proper strings and requires that strings be properly formed (so much so that it has additional completely separated platform-dependent types for dealing with OS-originated "stuff"). The one "difference" (which is more in the realm of implementation detail than language text model) is that Rust leverages its ownership system to make UTF8 "encoding" and "decoding" free (literally for the former, essentially for the former). The encoding and decoding are still there and explicit operations though. > Rust's text model has (almost) free (and copyless) transmutes from bytes to strings. Only for the specific case of input bytes already in the language's internal encoding (which granted will be common as most inputs would be ascii or utf-8) and with the same ownership constraints as the input, and that's mostly enabled by Rust's ownership model. > Python does not. Python doesn't generally do no-alloc/0-copy operations so that's not overly surprising.
- dom0 10y ago> Only for the specific case of input bytes already in the language's internal encoding (which granted will be common as most inputs would be ascii or utf-8) and with the same ownership constraints as the input, and that's mostly enabled by Rust's ownership model. Except of course on operating systems where text I/O is done entirely in UTF-16. Say, Windows. Since Python strings have no fixed encoding, but choose "the most efficient one" (heuristically) when decoding, they can cope better than a fixed UTF-8 encoding in these cases. >> Python does not. > Python doesn't generally do no-alloc/0-copy operations so that's not overly surprising. Indeed. Even when the encoding is not changed, the string will be always copied. One could think of an API that does that, though, to optimize all those cases were memory is already owned by a shim in the runtime.
- ubernostrum 10y agoYou should be aware Armin now has a more-or-less followup post telling people not to do what you just did (i.e., reference his 2011 post as an authoritative "Python 3 is bad" explanation, because both Python 3 and his own opinions have evolved since he wrote that post).
- ak217 10y agoArmin has backed off of this stance since then. And for good reason. As someone who works with Python text processing extensively, I can tell you that the Python 2.7 text model is broken and dangerous, due to the silent bytes-unicode coercion and misguided use of ascii instead of UTF-8 as the default text encoding. Many people don't realize this and will argue that it's not broken, because they have never fed non-ascii text through their app to watch it blow up! And once they realize that they have a problem, they then have to deal with a rat's nest of silent bytes-unicode coercions happening implicitly all over their app, sometimes impossible to deal with due to library code outside their control. There is a good discussion to be had on whether a language should prioritize bytes or unicode strings as the main data type, but there is no excuse for the "ticking timebomb" string data type design that pre-3 Python has with strings and the default encoding. For this reason alone I'm very happy that 2.7 is starting to lose its grip. Its continued support is a problem, and I have no love for people who are trying to hold on to it. There are many other features in 3 that I can no longer live without - most of them now available through backports modules - but types and asyncio can't be easily backported either, and people are starting to use them extensively.
- ianamartin 10y agoYeah, but if you are dealing only with a subset of the English Language in the U.S., and your API endpoint that you are scraping wants to serve to all peoples in all locales in all situations, you are fucked if you want to use Python3 and its csv module. You genuinely are better off using Python 2.7.x and its naive approach to text.
- rspeer 10y agoI don't understand what you mean by "your API endpoint that you are scraping wants to serve to all peoples in all locales in all situations". That would mean to me that the API endpoint could be sending me Unicode, in which case Python 3's Unicode-aware CSV is going to work great, and Python 2's csv is fucked. The limitations of Python 2's csv module was one of the key points that moved my company to Python 3. On Python 3, if you want to be naive about text (not sure why you're celebrating only working in a subset of English, but you have this option), you could open the file as Latin-1 and get the same results as Python 2. Many CSVs are made with Excel. Excel's only form of Unicode CSV is tab-separated UTF-16. Python 2's csv can't parse those at all, can it?