6 ms·
PyPy - Python 3k Status #5 update
- bcambel 14y agoWhen do you think Py3K will be main stream ? I'm afraid this might be a failing effort. Is there any company out there using Py3K in production ?
- kisielk 14y agoDo we really need to rehash this tired argument over every single post that references Python 3? You can look at any other Python 3 related post in the last couple of years on HN and Reddit to find endless arguments about this topic. How about a comment thread that actually talks about the post contents for once?
- deleted 14y ago[deleted]
- anacrolix 14y agoIt's a legitimate concern, Py3k has been a huge bungle for users.
- saurik 14y agoPaths and file names on Unix fundamentally do not have encoding: the filesystem represents them as sequences of bytes, and it is entirely possible to have a directory full of folders where there is no codec that is capable of faithfully or even reasonably decoding all of the names contained (or even that Unicode itself is capable of representing the semantically correct decoding of the filenames if you knew what theoretical codec represented them in the first place). It is therefore a fundamental mistake and a misinterpretation of the semantics of filesystems by Python 3 to insist that file names are represented by Unicode strings with a locale-sensitive encoding; in fact, I question whether there are interesting security ramifications inherent in this mistake (such as allowing me to change the locale in which a process is running and thereby remap its import path; this coming up, of course, as this article is largely about sys.path and Unicode).
- thristian 14y agoWell, it's a misinterpretation of the semantics of traditional Unix filesystems. Other filesystems used by other operating systems (such as Windows' NTFS and OS X's HFS+) genuinely do store filenames as Unicode strings, so on those platforms (or on Unix, if you've mounted one of those filesystems) Python 3.x's approach is exactly correct. That said, recent versions of Python 3 include a workaround for exactly the problem you describe: when the interpreter gets a byte-sequence value from the OS, filenames that decode to Unicode cleanly will be represented as proper Unicode strings, while filenames that can't be decoded will have the raw bytes represented as code-points in Unicode's Private Use Area. That way, even if Python can't decode the contents of a string, you can still, say, get a parameter from the command-line, pass it to open(), and be confident that you'll actually get the file the user intended.
- saurik 14y agoWhile I reference the issue of a mixed-codec directory in order to make clear the flaw in the operating assumption, the actual problem I am interesting in here, and which I conclude my statements with a reference to, is handling things like sys.path (hence this being a reply to the article). To respond to your comments, however: actually, the behavior on, let's say a Linux box, if you mount one of those aforementioned filesystems, will not be "exactly correct". Python 3 will attempt to encode the Unicode string using the current locale, pass it to the underlying Unix open() function, which will then have no clue what to do as it hits the filesystem. In fact, rather than just idly claiming this, I went ahead and set up exactly this test setup on one of my servers. I created a python3 script in a known specific source encoding (UTF-8) and asked it, in each of two different locales, to make a file that included an accented character, while mounted on an HFS+ disk image. hfs+# mount | grep hfs /.../hfs+.img on /.../hfs+ type hfsplus (rw,force) hfs+# cat test.py #!/usr/bin/python3 # -*- coding: utf-8 -*- open("helloä", "w") hfs+# LANG=en_US.UTF-8 ./test.py hfs+# LANG=fr_FR.ISO-8859-1 ./test.py hfs+# LANG=en_US.UTF-8 ls -la -rw-r--r-- 1 root root 0 2012-07-11 00:57 hello? -rw-r--r-- 1 root root 0 2012-07-11 00:57 helloä -rwxr-xr-x 1 root root 64 2012-07-11 00:56 test.py* hfs+# As you can see, the behavior here is really poor for something that claims to support Unicode. What we would expect to have happen is that, as I opened the file with a Unicode name on a Unicode filesystem that I would actually get the specific Unicode string that I had wanted. Instead, because Python 3 is only pretending to understand the semantics of filesystems with regards to character sets, and in fact has no way of taking advantage of the Unicode support in HFS+, we ended up with the encoding of the user's locale environment breaking our filenames. This would be akin to me doing JSON.encode() in a browser, and having a Unicode JavaScript string get converted into different JSON (which is also represented as a Unicode JavaScript string) depending on what language the user's browser is configured to use: that is a miserable Unicode failure, and not a success story. FWIW, the exact same code does work on Mac OS X (using the correct locale name of fr_FR.ISO8859-1): both versions of the script get the exact same filename. To be honest, I was somewhat surprised that they got that right ;P. I sadly do not have a Windows computer handy, as I'd love to see how it handles NTFS (which is technically UCS-2 and doesn't have the weird canonicalization behavior that HFS+ sometimes does: I can easily imagine broken corner cases with invalid UTF-16 surrogate pairs). Regardless, I really do not believe that you need to break the behavior on Linux in order to make Mac OS X work correctly. Even if such a tradeoff were required, I question whether it should be resolved with Linux on the losing end. I am not going to say that the correct solution doesn't even use Unicode strings: it just needs to more intelligently handle who is in charge of the encodings and what they semantically mean than Python 3 is prepared to do. As far as possible solutions go, I certainly do not believe that the private Unicode codepoint solution is either sufficient (as it doesn't solve the de novo filename creation problem) nor even remotely reasonable (as if I attempt to then communicate these filenames with other systems, or heaven-forbid wanted to store a filename with a private Unicode codepoint in it, I'm now screwed). (edit: Wow, BTW. They added this "solution" after I had already given up on Python 3k, so I hadn't followed up to see what they did. It seems like they ended up not using "private codepoints" as they were considering in 2008, but are instead using surrogate-halves: they are encoding "malformed UTF-8 sequences as malformed UTF-16 sequences"[1]. Given that you are allowed to store malformed UTF-16 on NTFS, I wonder how they expect that to work. Still doesn't solve either of my complaints, though.) 1: http://hyperreal.org/~est/utf-8b/releases/utf-8b-20060413043934/kuhn-utf-8b.html http://hyperreal.org/~est/utf-8b/releases/utf-8b-20060413043...