6 ms·
The Fundamental Problem in Python 3
- kresten 7y agoIt takes a specific type of personality to remain angry about Python 3 in Dec 2019. Didn’t that story finish?
- mehrdadn 7y agoWhy would the stories finish when the pains they refer to are still there like before...?
- jasondclinton 7y agoI know if at least a few big tech companies that haven't migrated from 2 yet. So, a lot of folks are discovering Python 3 for the first time.
- philipov 7y agoIn my experience, companies haven't yet migrated from Python 2 for the same reason they haven't yet migrated from COBOL. It has nothing to do with Python 3's benefits or flaws.
- KaiserPro 7y agoits a specific bug in python3. not just rant, its a specific, well reasoned bug report.
- jnwatson 7y agoIt is horrible compromise of a bad situation. 99% of the time file name bytes are UTF-8 decodable. What do you do in the other case? You could return a different type, but then that makes dealing with it harder. You could make a new filesystem name type, but that would make simple APIs complicated. Using a new control character is probably the best one could do.
- oefrha 7y agoJust a rehash of https://changelog.complete.org/archives/10053-the-incredible-disaster-of-python-3 https://changelog.complete.org/archives/10053-the-incredible... https://news.ycombinator.com/item?id=21606416 https://news.ycombinator.com/item?id=21606416 Arguing that python3’s str model doesn’t work well with POSIX’s “any bag of bytes can be a filename” model. Plus new rants about surrogateescape which the author has learned since publishing the last article. The author sure has a penchant for flamebait titles.
- mehrdadn 7y agoNo comment on the titles, but the issues are real (the linked page explains some of it decently [1]). I know I've found it excruciatingly difficult to write Python code that handles non-ASCII stdio correctly, especially one that might display on a terminal, especially a portable manner. Some of the compatibility issues are inherently hard problems in any language, but others are Python-related, and it didn't get better in Python 3. [1] http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/ http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/
- oefrha 7y ago> I've found it excruciatingly difficult to write Python code that handles non-ASCII stdio correctly, especially one that might display on a terminal, especially a portable manner, and it didn't get better in Python 3. I've been doing just that in Python 3 for half a decade at least -- got a fairly popular open source CLI application that hasn't seen a Unicode complaint for years. It doesn't come for free on all possible configurations (yeah, every developer with a wide enough userbase has seen the dreaded "'ascii' codec can't encode characters in position ...: ordinal not in range" at some point) but it's definitely not "excruciatingly difficult". Bad things only happen sys.stdout.encoding isn't utf-8, which is rare on * nix systems -- fixed by setting PYTHONIOENCODING or setting locale to * .UTF-8. Encoding on windows is always a headache (not specific to Python at all) but somehow we seemed to have managed to steer clear too. Anyway, just check sys.stdout.encoding. Meanwhile, this article is about problems that arise when you have garbage filenames that are neither Unicode nor whatever Microsoft's encoding; I wouldn't expect the average user to deal with files like that in day-to-day usage.
- adrian17 7y ago"into what is fundamentally a weakly-typed, dynamic language." Nit: isn't Python generally considered to be strongly, dynamically typed?
- wirrbel 7y agoYes
- njharman 7y agoInaccuracy and hyperbole are the tools of click baiters like OA.
- nhumrich 7y agoCorrect
- shean_massey 7y agoExactly. I stopped looking for insights as soon as I read that one.
- coldtea 7y agoYes, because any post containing an error can never also include insights...
- xapata 7y agoUse pathlib?
- mikl 7y ago"Python does not cater to my favourite edge case" != fundamental problem In the days before UTF-8-everywhere, file names with anything besides alphanumerics and safe symbols like dashes or underscores were always a problem. If you had special characters in your filenames, you were almost certain to run in to problems, since their meaning varied greatly depending on how the system was set up - codepages and whatnot. But this is only a problem if you have such files, and unless you’ve kept files around for decades, you don’t. So young programmers can grow up never having problems with this. Everything will be UTF-8 and it’ll just work. And as for broken old file names, who cares? Fix your file names and move on. There’s no reason that Python 3 should have workarounds for problems that were solved over a decade ago.
- KaiserPro 7y ago> And as for broken old file names, who cares? Fix your file names and move on. what if you are trying to process "foreign"(as in, not created by you) filenames, trying to validate/conform with python? I mean its a great DoS vector, which is difficult to protect against with python. it'll crash. Which the point the article is trying to make. > unless you’ve kept files around for decades, you don’t. Thats both untrue and not very helpful. Its perfectly possible to bump into files like this.
- viraptor 7y ago> trying to validate/conform with python Then you can use the 'bytes' variant of calls. If you process arbitrary data, there's lots of things you need to think about. This is just one of them. And if you'd crash otherwise, you should have exception handling in place. I understand it's a problem and it's not trivial to support this, but it's not difficult either.
- oefrha 7y ago> what if you are trying to process "foreign"(as in, not created by you) filenames, trying to validate/conform with python? I mean its a great DoS vector, which is difficult to protect against with python. 1. Have catch-all exception handling. Exception may not be your fault, exception-caused "crashing" whatever that means is entirely your fault. 2. Use os.fsdecode. https://docs.python.org/3/library/os.html#os.fsdecode https://docs.python.org/3/library/os.html#os.fsdecode 3. Don't process random untrusted filenames. Sanitize if it's some sort of HTTP upload, for instance.
- goatinaboat 7y agoThe truth is that 99% of programmers can do everything they need to in ASCII and the other 1% are working on tools to handle Unicode itself. It’s a mistake and as soon as it goes the way of the <blink> tag the better. At least that tag was amusing for a short while...
- bildung 7y agoThis is only true for the native English speakers. 96% of the world population aren't.
- doteka 7y agoAlways felt like a strange argument to me. I grew up bilingual, neither of these languages was English. Never in my life has it occurred to me to name a file I created with something else than ascii characters.
- bildung 7y agoMe neither, but I don't write the applications only for myself :) Most normal people are not aware of these problems, e.g. because they only have used Windows for their whole life and never crossed OS boundaries. Then they never experience the encoding problems that teached us the hard way how to properly name files.
- goatinaboat 7y agoneither of these languages was English. Never in my life has it occurred to me to name a file I created with something else than ascii characters. This is the totally normal experience of every programmer from every country. The only people pushing Unicode, ironically, are white native English speakers who think they’re saving the world.
- bildung 7y agoThe people pushing Uniode are those who have to support their own applications in the field.
- smitty1e 7y agoThe lack of any Pathlib discussion seemed curious.
- kthejoker2 7y ago> This was the most common objection raised to my prior post. “Get over it, the world’s moved on.” Gee I wonder why ...