16 ms·
The Python Unicode Mess
- repolfx 8y agoThat sounds more like a mess handling things that are not Unicode.
- masklinn 8y agoYes, but the issue here would be that Python forced these "things which are not unicode" into unicode.
- softblush 8y agohttps://web.archive.org/web/20181006121702/http://changelog.complete.org/archives/9938-the-python-unicode-mess https://web.archive.org/web/20181006121702/http://changelog....
- kabacha 8y ago> Python's unicode is a "mess" because of this single edge case I've encountered FTFY
- est 8y agomore like > Unicode is a "mess" because it can not unquote arbitrary backslash strings.
- burntsushi 8y agoThis article is kind of hard to evaluate, because the OP doesn't provide an example program with an example input that fails. So it's hard to judge whether the solution presented here is actually ideal. Instead, we're forced to just take the OP's word for it, which is kind of uncomfortable. I do somewhat agree with the general sentiment, although I find it difficult to distinguish between the problems specifically related to its handling of Unicode and the fact that the language is unityped, which makes a lot of really subtle things very implicit.
- yorwba 8y agoThe OP links to StackOverflow, where failing inputs are mentioned in the comments on the accepted answer. And the second-most upvoted answer explains that .decode('unicode_escape') only works for Latin-1 encoded text: https://stackoverflow.com/a/24519338 https://stackoverflow.com/a/24519338
- lozenge 8y agoThe question being how to parse character escapes (backslash sequences) in Python. To be honest, you could write a custom character-by-character parser easily or even use the regex module.
- burntsushi 8y agoBut the OP has their own problem with which they posted a solution, but didn't include the actual problematic program.
- upofadown 8y agoThe article is about a specific instance (filenames). In general, handling Unicode as a bunch of indexable code points as per Py3 turned out to be not that great. I guess the idea came from the era where people still thought that strings could be in some sense fixed length. These days we better understand that strings are inherently variable length. So there is no longer any reason to not just leave everything encoded as UTF-8 and convert to other forms as and if required. Strings are just a bunch of bytes again.
- ubernostrum 8y agoThe most correct way to expose Unicode to a programmer in a high-level language is to make grapheme clusters the fundamental unit, as they correspond to what people think of as "characters". Failing that, exposing strings as sequences of code points is a second-best choice. UTF-8 is a non-starter because it encourages people to go back to pretending "byte == character" and writing code that will fall apart the instant someone uses any code point > 007F. Or they'll pat themselves on the back for being clever and knowing that "really" UTF-8 means "rune == code point == character", and also write code that blows up, just in a different set of cases. And yes, high-level languages should have string types rather than "here's some bytes, you deal with it". Far too many real-world uses for textual data require the ability to do things like length checks, indexing and so on, and it doesn't matter how many times you insist that this is completely wrong and should be forbidden to everyone everywhere; the use cases will still be there.
- masklinn 8y ago> The most correct way to expose Unicode to a programmer in a high-level language is to make grapheme clusters the fundamental unit, as they correspond to what people think of as "characters". > UTF-8 is a non-starter because it encourages people to go back to pretending "byte == character" and writing code that will fall apart the instant someone uses any code point > 007F. These are somewhat different concerns, you can provide cluster-based manipulation as the default and still advertise that the underlying encoding is UTF-8 and guarantees 0-cost encoding (and only validation-cost decoding) between proper strings and bytes (and thus "free" bytewise iteration, even if that's as a specific independent view). > Far too many real-world uses for textual data require the ability to do things like length checks, indexing and so on This is not a trivial concern e.g. "real-world uses for length checks" might be a check on the encoded length, the codepoint length or the grapheme cluster length. Having a "proper" string type doesn't exactly free you from this issue, and far too many languages fail at at least one and possibly all of these use cases, just for length queries.
- lincolnq 8y agoIndeed py3 decided to make unicode strings the default. This fixes all sorts of thorny issues across many use cases. But it does indeed break filenames. I haven't dealt with this issue myself, but the way python was supposed (?) to have "solved" this is with surrogate escapes. There's a neat piece on the tradeoffs of the approach here: https://thoughtstreams.io/ncoghlan_dev/missing-pieces-in-python-3-unicode/ https://thoughtstreams.io/ncoghlan_dev/missing-pieces-in-pyt... Maybe handling the surrogates better would allow you to use 'str' everywhere instead of bytes?
- tyingq 8y agoHis examples are all stuff that isn't Unicode. The filename thing would probably work using a latin1 encoding, since that leaves 8 bit bytes undisturbed.
- nicolaslem 8y agoFor anyone interested in learning why Python 3 works this way I highly recommend the blog of Victor Stinner[0]. As for the article, this is nothing new. The problem is similar to the issues raised by Armin Ronacher[1]. These problems are well known and Python developers address them one at a time. Issues around these egde cases have improved since the initial release of Python 3.0. [0] http://vstinner.github.io http://vstinner.github.io [1] http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/ http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/
- zorkw4rg 8y agoI'm not so sure other languages do that any better (nodejs doesn't even support non-unicode filenames at all for instance). Modern python does a pretty good job at supporting unicode, very far away from being a "Mess" that's just very much not true at all. People always like to hate on python but then other languages supposedly designed by actually capable people do mess up other stuff all the time. Look at how the great Haskell represents strings for instance and what a clusterfuck[1] that is. [1] https://mmhaskell.com/blog/2017/5/15/untangling-haskells-strings https://mmhaskell.com/blog/2017/5/15/untangling-haskells-str...
- pvarangot 8y agoWhat's the deal with Haskell strings? It's not a mess, it's basically enforcing the same "unicode sandwich" approach Python recommends by using the type checker. Of course to do that you need one type for when the string is in the different possible different layers of the sandwich. There's added types for lazy vs. non-lazy but that's for performance optimization, and don't get me started on how Python "get messy" when you want to do performance optimization because it usually kicks you out of the language.
- bunderbunder 8y ago> What's the deal with Haskell strings? I think the linked article laid the case fairly well. Basically, Haskell has a bunch of string types you need to understand and fret about, and the one named "String" is the one you almost never want, but it's also the only one with decent ergonomics unless you know to install a compiler extension. I think it's a fair criticism. The "Lots of different string types" thing isn't (IMO) such a big deal coming from a language of Haskell's vintage. Given what Python's "decade-plus spent with a giant breaking change right in the middle of the platform hanging over our heads" wild ride has been like, I can't blame anyone for not wanting to replicate the adventure. But, for newcomers, the whole thing where you need to know to {-# LANGUAGE Support, Twenty, First, Century #-} is a pretty big stumbling block.
- pvarangot 8y ago
- prevedmedved 8y agoLooks like we need a Python 4. (/s)
- kbumsik 8y agoWell, Python already has a plan for Python 4. The Python 4 will be released after Python 3.8. There are already discussion on Python 4.0 in the dev group. It is just a new number after 3.8 so there won't be breaking issues like 2=>3.
- 1wd 8y agoI think that was just some idea that was discarded. "Seems that we've reached the consensus: we release Python 3.10 after Python 3.9. We maybe release Python 4.0 at some point if there's a significant backwards incompatible change." https://mail.python.org/pipermail/python-committers/2018-September/006159.html https://mail.python.org/pipermail/python-committers/2018-Sep...
- benatkin 8y agoPython 3 to be retired in 2050 =)
- gnud 8y ago> For a Python program to properly support all valid Unix filenames, it must use “bytes” instead of strings, which has all sorts of annoying implications. While in python 2, you had to use unicode strings for all sorts of actual text, which caused its own problems. > What’s the chances that all Python programs do this correctly? Yeah. Not high, I bet. Exactly.
- acdha 8y agoAs far as I can tell this is a long-form “I used to be able to ignore encoding issues and now it’s a ‘mess’ because the language is forcing me to be correct”. Each of the examples cited is something which was a source of latent bugs which he thought was working because they were ignored. Only his third bit of advice isn’t wrong and treating it as something unusual shows the problem: the only safe way to handle text has always been to decode bytes as soon as you get them, work with Unicode, and then encode it when you send them out. Anything else is extremely hard to get right, even if many English-native programmers were used to being able to delay learning why for long periods of time.
- masklinn 8y ago> As far as I can tell this is a long-form “I used to be able to ignore encoding issues and now it’s a ‘mess’ because the language is forcing me to be correct”. The problem with that view is that there are things for which you can not be correct, and there are no encoding issues because there is no encoding (or if there is one it does not map to proper unicode): * UNIX files and paths have no encoding, they're just bags of bytes, with specific bytes (not codepoints, not characters, bytes) having specific meaning * Windows file and path names are sequences of UTF-16 code units but not actually UTF-16 (they can and do contain unpaired surrogates), as above with specific code units (again not codepoints or characters) having specific meaning These are issues you will encounter on user systems, there is no "forcing you to be correct". A non-unicode path is not incorrect. On many systems it just is. OSX is one of the few systems where a non-unicode path is actually incorrect, and that means you will not encounter one as input so you have no reason to handle this issue at all. > Only his third bit of advice isn’t wrong and treating it as something unusual shows the problem: the only safe way to handle text That's where you fail, and to an extent so does Python: some text-like things are not actually text. Path names famously is one of this case. You're trying to hammer the square peg of path names in the round hole of unicode.
- acdha 8y agoThat’s just restating my point: Unix filenames are bytes (on most filesystems, anyway). The fact that many people were able to conflate them with text strings was a convenient fiction. Python no longer allows you to maintain that pretense but it’s easy to deal with by treating them as opaque blobs, attempt to decode and handle errors, or perform manipulations as bytes.
- IshKebab 8y agoThis is just the cost of using a dynamic language with implicit error handling (exceptions).
- ubernostrum 8y agoYou can garble filenames just as easily in statically-typed languages. Consider, for example, Windows' infamous 16-bit units that aren't actually well-formed UTF-16. I'm not aware of any widely-used programming language whose type system will save you from that sort of thing ("here's some bytes, figure out if they're a string and if so what encoding" is a historically very difficult problem).
- masklinn 8y ago> You can garble filenames just as easily in statically-typed languages. If the language assumes filenames are regular language strings, which not all do. > Consider, for example, Windows' infamous 16-bit units that aren't actually well-formed UTF-16. unix filenames are literally just bags of bytes with no known or specified encoding.
- ubernostrum 8y agoUnix filenames don't pretend to be something they aren't. Windows filenames like to present a convincing façade of being UTF-16 right up until they aren't.
- masklinn 8y agoWindows filenames "like to present a convincing façade of being UTF-16" in the exact same way unix filenames "like to present a convincing façade of being UTF-8". Both are common assumptions neither is actually true, and all of that is well-documented.
- ubernostrum 8y agoI would express it more as "programmers in Unix environments like to act as if everything still uses the C locale everywhere, all the time".
- aeturnum 8y agoMy main criticism of Python 3's changes to strings is that it has become much more specific about strings. In Python 2, if you have a series of bytes -or- a "string", the language has no opinion about the encoding. It just passes around the bytes. If that set of bytes enters and exits Python without being changed, its format is of no concern. Interactions do not force you to define an encoding. This is not correct, but it is often functional. Python 3, on the other hand, if you ever treat bytes as a string, forces you to have an opinion about the encoding. Same goes for if you convert back to bytes. For uncommon or unexpected encodings, the chance of this going wrong in a casual, accidental way is much higher. Of course, the approach is more correct, but it doesn't feel more correct to the programmer.
- ubernostrum 8y agoIn Python 2, if you have a series of bytes -or- a "string", the language has no opinion about the encoding This is incorrect. Python 2 lets you get away with a lot of things, but some string-y operations on Python 2 bytestrings will still trip the "need to know the encoding" issue. And absent you telling it the encoding, Python 2 will assume ASCII and begin throwing exceptions as soon as it sees a byte outside the ASCII range.
- ak217 8y ago> it doesn't feel more correct to the programmer. I agree with the details of what you said, but the insidious thing about how Python 2 organized strings and encodings is that most programmers were free to ignore it and produce buggy software. Then, later, people who had to use that software on non-ascii data would try to use it and it would blow up. This would lead to a very painful cycle of shaking out bugs that the original author may not even be motivated to fix. The decision to force encodings to be explicit and strings/bytes to be separate was a great design change. It literally made all our code more valuable by removing hidden bugs from it.
- josteink 8y ago> Python 3, on the other hand, if you ever treat bytes as a string, forces you to have an opinion about the encoding. Of course. How else can you translate the bytes into a meaningful textual representation without an encoding? Python 3 requires you to consider the real world, and that’s a good thing. (Apart from when you want to port sloppily written Python on 2.x code)
- apk-d 8y agoLet's be honest, the real mess is with UNIX filenames. I dare you to come up with a legitimate use case for allowing newlines and other control characters in a file name.
- gpderetta 8y agoBackward compatibility.
- Volt 8y agoWith what?
- gpderetta 8y agowith almost 50 years of unix history.
- Volt 8y agoI think the point was that UNIX got it wrong, and we've been dealing with the consequences ever since. It's of course too late to change it, so yeah.
- gpderetta 8y agoMaybe. But 50 years ago utf-8 didn't exist, unicode didn't exist, possibly not even latin-1 did exist. If unix had enforced a specific encoding (which implies constrains which byte values can appear in a path byte string), transition to newer encodings would have been significantly harder.
- dcbadacd 8y agoIt's like a built-in unit test - devs have to not mangle and assume anything about filenames they get from the system - though they still do, I've seen multiple times how my nice umlauts get mangled or my spaces cause scripts to fail.
- 0x006A 8y agoi love unicode handling in python3, it's so much better to work with. python2 was a mess, migrating old code requires looking at old code, the result is only better code, never a mess.
- garethrees 8y agoThere is a particular use case which leads to frustration with Python 3, if you don't know the latin1 trick. The use case is when you have to deal with files that are encoded in some unknown ASCII-compatible encoding. That is, you know that bytes with values 0–127 are compatible with ASCII, but you know nothing whatsoever about bytes with values 128–255. The use case arises when you have files produced by legacy software where you don't know what the encoding is, but you want to process embedded ASCII-compatible parts of the file as if they were text, but pass the other parts (which you don't understand) through unchanged (for example, the files are documents in some markup language, and you want to make automatic edits to the markup but leave the rest of the text unchanged). Processing as text requires you to decode it, but you can't decode as 'ascii' because there are high-bit-set characters too. The trick is to decode as latin1 on input, process the ASCII-compatible text, and encode as latin1 on output. The latin1 character set has a code point for every byte value, and bytes with the high bit set will pass through unchanged. So even if the file was actually utf-8 (say), it still works to decode and encode it as latin1, and multi-byte characters will survive this process. The latin1 trick deserves to be better known, perhaps even a mention in the porting guide.
- codedokode 8y agoWould not a better solution be to process the file as a byte string?
- zbentley 8y agoI don't think so. If you want to detect and operate on only the data that could represent ASCII characters, you could, certainly process it as a byte string if you wanted, but you'd have to track the presence of non-ASCII-range character codes yourself, and keep state around to represent whether you were in the middle of a multibyte character as you read through the bytes. If done right, it would be a (probably much slower) re-implementation of what happens when you use the latin1 trick mentioned. You have to get it right, though (sneaky edge cases abound--what if the file starts in the middle of an incomplete multibyte character?). TL;DR this could technically work but is a poor idea.
- minitech 8y ago> And the environment? [it’s not even clear.] https://stackoverflow.com/questions/44479826/how-do-you-set-a-string-of-bytes-from-an-environment-variable-in-python https://stackoverflow.com/questions/44479826/how-do-you-set-... That question is about interpreting backslash escape sequences for bytes in an environment variable. All this person wants is `os.environb` (and look, its existence highlighted a Windows incompatibility, saving them from subtle bugs like every other Python 3 improvement). https://docs.python.org/3/library/os.html#os.environb https://docs.python.org/3/library/os.html#os.environb
- mixmastamyk 8y agoThanks, never noticed environb. I’m still learning new things about Python 3 ten years later.
- SoulMan 8y agoI just came from pycon India 2018. This is exactly what the keynote was about.(it was by author of Flask)
- codedokode 8y agoThe author has files with invalid names and complains that Python refuses to accept them. Maybe he should fix the names first?
- gpderetta 8y agoIf all tools he had access to behaved in the same way, he wouldn't be able to fix these "wrong" file names.
- flohofwoe 8y agoIMHO the whole python3 string mess could have been prevented if they had chosen UTF-8 as the only string encoding instead of adding a strict string type with a lot of under-the-hood magic. That way strings and byte streams could remain the same underlying data, just as in python2. The main problem I have with byte-streams vs strings in python3 is that it adds a strict type checking at runtime which isn't checked at 'authoring time'. Some APIs even make it impossible to do upfront type checking even if type hints would be provided (e.g. reading file content either returns a byte stream, or a string, based on the content of a string parameter in the file open function). Recommended reading: http://utf8everywhere.org/ http://utf8everywhere.org/
- marcosdumay 8y ago> IMHO the whole python3 string mess could have been prevented if they had chosen UTF-8 as the only string encoding instead of adding a strict string type with a lot of under-the-hood magic. That is basically what Python2 does, and it is completely wrong.
- flohofwoe 8y agoCan you give any reasons why this is completely wrong? The web seems to work just fine with UTF-8. The advantage is that you can pass string data around as generic byte streams without even knowing about the encoding. You'll only have to care about the encoding at the end points.
- toyg 8y agoYou are joking, right? Have you ever seen non-English webpages? More often than not, a multitude of ??? and Chinese characters pop up at some point or another.
- flohofwoe 8y agoI'm from Germany so I've seen a few non-English webpages. I can't remember having seen any text rendering problems since the late 90's or so.
- ptx 8y agoText encoding in general is a mess, and Python 2 Unicode support was a mess, but Python 3 makes it much less of a mess. I think the author has a mess on his hands because he's trying to do it the Python 2 way – processing text without a known encoding, which is not really possible, if you want the results to come out right. To resolve the mess in Python 3, choose what you actually want to do: 1. Handle raw bytes without interpreting them as text – just use bytes in this case, without decoding. 2. Handle text with a known encoding – find out the encoding out-of-band from some piece of metadata, decode as early as possible, handle the text as strings. 3. Handle Unix filenames or other byte sequences that are usually strings but could contain arbitrary byte values that are invalid in the chosen encoding – use the "surrogateescape" error handler; see PEP 383: https://www.python.org/dev/peps/pep-0383/ https://www.python.org/dev/peps/pep-0383/ 4. Handle text with unknown encoding – not possible; try to turn this case into one of the other cases. Also, watch Ned Batchelder's excellent talk, Pragmatic Unicode, or, How do I stop the pain?, from 2012: https://pyvideo.org/pycon-us-2012/pragmatic-unicode-or-how-do-i-stop-the-pain.html https://pyvideo.org/pycon-us-2012/pragmatic-unicode-or-how-d...
- zeroname 8y ago> To resolve the mess in Python 3, choose what you actually want to do... The thing is that this is not actually going to happen. Programs are simply broken across the board, because few people can be bothered to deal with all these peculiarities. The difference is, in Python 2, output would be corrupted in some edge cases, but generally it would "just work". In Python 3, the program falls flat on its face even in cases that would've ended up working fine in Python 2. I don't think there's a general answer on which behavior causes less real-world problems total, but the idea that Python 3 makes less of a mess is not something I can agree with.
- talltimtom 8y agoNy experiance being from a non english language was the exact opposite. Python 2 would fail in horribly weird ways and you constantly needed to add weird tricks to get simple functions working. I. Python 3 I haven’t even encounters any similar issues everything just works. Of cause sometimes you need to specify some encodings but I don’t view that as a failure of the language. I think a lot of people have a biased view because tons of issues where just not apparent in English, but if you want a language to be viable for the entire world you have to look outside that limited set of characters.
- zzzeek 8y agoDon't think of python Unicode as a "string". Think of it as "text". I don't really understand the issues the author is having with things like sys.stdout and such because he did not provide complete examples. He should cite actual examples and bug reports that he has posted for these things, ive had no such issues. There's a lot of things we need to do to accommodate for non-ascii text but they are all "right" as far as I've observed.
- snicker7 8y agoThere are lots of comments indicating that the programmer is doing things wrong. But what is the right way to deal with encoding issues? Wait for code to break in production? Whatever "best practices" there are for dealing with unexpected text encoding in Python, they do not seem to be widely well-known. I bet a large % of Python programmers (myself included) made the exact same errors the author had, with little insight as to how avoid them in the future.
- wParser 8y agoFilenames are a good example to show people why forcing an encoding onto all strings simply doesn't work. The usual reaction from people is to ignore that and they'll shout: "fix your filenames!" Here is another example: Substrings of unicodestrings. Just split a unicodestring into chunks of 1024 bytes. Forcing an encoding here and allowing automatic conversions will be a mess. People will shout: "Your're splitting your Strings wrong!" The first language I knew that fell for encoding aware strings was Delphi - people there called it "Frankenstrings" and meanwhile that language is pretty dead. As a professional who has to handle a lot of different scenarios (barcodes, Edifact, Filenames, String-buffers, ...) - in the end you'll have to write all code using byte-strings. Then you'll have to write a lot of GUI-Libraries to be able to work with byte-strings... and in the end you'll be at the point where the old Python was... (In fact you'll never reach that point because just going elsewhere will be a lot easier)
- joshuamorton 8y agoSub strings of Unicode strings are fine. Byte level chunking of a Unicode string requires encoding this string as bytes, then working with bytes, then deciding the text. Splitting a piece of Unicode text every 1024 bytes is like splitting an ascii string every 37 bits. It doesn't make sense.
- jessaustin 8y agoTFA is short and to the point. A few examples, a few links to other examples. Py3's insistence on shoving Unicode into every API it possibly could maybe fit, is often inconvenient for coders and for users. This thread has 100 comments, mostly disagreeing in the same fingers-in-ears-I-can't-hear-you fashion. Whom are we struggling to convince, here?
- TimJYoung 8y agoGetting Unicode right, especially with various file systems and cross-platform implementations is hard, for sure. But, I think this quote: "And, whatever you do, don’t accidentally write if filetype == "file" — that will silently always evaluate to False, because "file" tests different than b"file". Not that I, uhm, wrote that and didn’t notice it at first…" shows a behavior that, to me, is inexcusable. The encoding of a string should never cause a comparison to fail when the two strings are equivalent except for the encoding. For example, in Delphi/FreePascal, if you compare an AnsiString or UTF-8-encoded string with a Unicode string that is equivalent, you get the correct answer: they are equal.
- kbumsik 8y agoYeah, this behavior might be because Python doesn't store unicode string as it is. AFAIK Python always store string as an array of fixed-sized bytes for random access. In other words, the size of an element of the array is the same as the maximum size of characters in the string, meaning that even a character of 1 byte ASCII can be stored as 4 bytes. So when one side is bytes (filetype in this case) and the other side is a string, the underlying byte representation can be different even if they represent as the same string in higher level.
- ubernostrum 8y agoPython as of 3.3 chooses an internal representation on a per-string basis. This encoding will be either latin-1, UCS-2, or UCS-4, and the choice is made based on the widest code point in the string; Python chooses the narrowest encoding capable of representing that code point in a single unit. This does mean that a string which contains, say, some English text and an emoji will "blow up" into UCS-4, but the overhead isn't that severe; most such strings are not especially large. It also means that strings containing only code points < U+00FF are smaller in memory on Python 3.3+ than previously, since prior to 3.3 they would be using at least two bytes per code point and now use only one.
- mikezter1 8y ago> The encoding of a string should never cause a comparison to fail when the two strings are equivalent except for the encoding. You'll have to admit that the encoding is a property of a string, just like the content itself. As always, you as a programmer are bound to know both of these properties to have predictable results. To compare two strings of different encoding to one another, you'll have to find a common ground for interpreting the data contained in the string. If you don't want or need that, then all you have is a "string" of bytes.
- madrox 8y agoDealing with string encoding has always been the bane of my existence in Python...going back over 10 years when I first started using it. I've never had such wild issues with decoding/encoding in other languages...that may be my privilege, though, since I was dealing with internal systems before Python, and then I got into web scraping. Regardless, string encoding/decoding in Python is hard, and it doesn't feel like it needs to be.
- mikezter1 8y agoThe encoding is a property of the string, just like the content, just as with any other object. If you want to compare strings with different encodings, you'll have to convert at least one of them. I was never forced into encoding hell again, after reading this excellent post: https://www.joelonsoftware.com/2003/10/08/the-absolute-minimum-every-software-developer-absolutely-positively-must-know-about-unicode-and-character-sets-no-excuses/ https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
- Walkman 8y agoIt's just a not very well explained rant of some shitty libraries and a lot of legacy code. If you want to read about REAL complaints, read Armin Ronacher thought about it instead: http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/ http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/
- loeg 8y agoI agree Python3 is an awful mistake and that straight-up Unicode is not well suited for storing arbitrary byte strings from old disk images. However, Python 3.1+ encode disk names as WTF-8 (aka utf-8b): https://www.python.org/dev/peps/pep-0383/ https://www.python.org/dev/peps/pep-0383/ .
- jlarocco 8y agoI can't believe there are still people whining about this in 2018. Those problems with gpodder, pexecpt, etc. aren't due to Python 3, they're due to the software being broken. Without knowing the encoding, UNIX paths can't be converted to strings. It's unfortunate, but that's the way it is, and it's not Python's fault.
- singularity2001 8y agoI invest some karma to point out how I'd love for str to just use UTF-8 by default, and print as UTF-8 by default: print(b'DONT b"EVERYTHING!"') print(str(b'SAME!')) print(str(b'I DONT WANT TO add ,"UTF-8" everywhere!','UTF-8')) line="ום עולם" output.write(line) # TypeError: a bytes-like object is required, not 'str' fp.write(output.getvalue()) # TypeError: write() argument must be str, not bytes Please at least allow us to set a global option via sys.setdefaultencoding('UTF8') as before to automatically encode/decode as UTF-8 by default!
- gspetr 8y agoThis post barely scratches the tip of the iceberg. For a more comprehensive discussion of unicode issues and how to solve them in Python, "Let’s talk about usernames" does this issue more justice than I could write in a comment: https://news.ycombinator.com/item?id=16356397 https://news.ycombinator.com/item?id=16356397
- andrewstuart 8y agoThat's not Python's fault - those are programmer errors. Having said that, Python really has something to answer for with "encode" versus "decode" - WTF? Which is which? Which direction am I converting? I still have to look that up every single time I need to convert. Why the heck are there not "thistothat" and "thattothis" functions in Python that are explicit about what they do?
- burntsushi 8y agoThis is something I see folks trip on. I think encode/decode are fine names actually. The problem it's that Unicode strings have decode defined at all, and similarly, that byte strings have encode defined. Byte strings should only have a decode operation and Unicode strings should only have an encode operation. Depending on your input, the wrong operations can actually succeed!
- guitarbill 8y agoNot sure what you're talking about mate, on Python 3.6: >>> "hello world".decode("utf-8") Traceback (most recent call last): File "<stdin>", line 1, in <module> AttributeError: 'str' object has no attribute 'decode' >>> b"hello world".encode("utf-8") Traceback (most recent call last): File "<stdin>", line 1, in <module> AttributeError: 'bytes' object has no attribute 'encode'
- burntsushi 8y agoThat's good. I guess it's only in Python 2 then.
- luckystarr 8y agoIn Python 2 it works almost the same, except you only get an error when encoding/decoding doesn't work out. So I see this as an improvement.
- dcbadacd 8y ago
- Aardappel 8y agoGoing of on a tangent a bit here, but I think there are 2 important related issues: * API design should fit the language. In a "high on correctness" language like Haskell or Rust, I'd expect APIs to force the programmer to deal with errors, and make them hard to ignore. In a dynamically typed language like Python where many APIs are very relaxed / robust in terms of dealing with multiple data types (being able to see numbers/strings/objects generically is part of the point of the language), being super strict about string encoding sounds extra painful compared to a statically typed language. I'd expect an API in this language to err on the side of "automatically doing a useful/predictable thing" when it encounters data is only slightly incorrect, as opposed to raising errors, which makes for very brittle code. Most Python code is the opposite of brittle, in the sense that you can take more liberties with data types before it breaks than in statically typed languages. Note that I am not advocating incorrect APIs, or APIs that silently ignore errors, just that the design should fit the language philosophy as best as possible. * Where in a program/service should bytes be converted to text? Clearly they always come in as bytes (network, files..), and when the user sees them rendered (as fonts), those bytes have been interpreted using a particular encoding. The question where in the program should this happen? You can do this as early as possible, or as late as possible. Doing it as early as possible increase the code surface where you have to deal with conversions, and thus possible errors and code complexity, so that doesn't seem so great to me personally, but I understand there are downsides to most of your program dealing with a "bag of bytes" approach too.
- slavik81 8y agoPart of the problem is that encoding is treated as something that must be explicitly handled in the string API, but it's something that's just assumed by default in the IO API. Python just guesses what the encoding is, and it often guesses wrong. The design of the API leads people to do the wrong thing. Encoding should be a required argument for `open` in text mode.
- dan-robertson 8y agoI don’t think Haskell is a very good example to promote for string handling. Things are mostly strict and well behaved once they make it into the Haskell program but before then they either need to satisfy the program’s assumptions before being input or the program will be buggy/crash unless it is carefully written such that it’s assumptions are right.
- perlgeek 8y agoThe real problem here is that * UNIX file systems allow any byte sequence that doesn't contain / or \0 as file and directory names * User interfaces have to render that as strings, so they must decode * There is no meta data about what the file encoding is Many programs use the encoding from the current locale, which is mostly a good assumption, but the way that locales scope (basically per process) has nothing to do with how file names are scoped. So, many programs make some assumptions. Some models are: 1) Assume file names are encoded in the current locale 2) Assume file names are encoded in UTF-8 3) Don't assume anything The "correct" model would be 3), but it's not very useful. People want to be able to sort and display file names, which generally isn't very useful with binary data. Which is why most programs, including python, use 1) or 2), and sometimes offer some kind of kludge for when the assumption doesn't hold -- and sometimes not. IMHO a file system should store an encoding for the file names contained in it, and validate on writes that the names are correct. But of course that would be a huge POSIX incompatibility, and thus won't happen. People just live with the current models, because they tend to be good enough. Mostly.
- makecheck 8y agoA file system doesn’t fix the problem either because there are regularly problems at file system boundaries. For example, moving from a case-sensitive file system to a case-insensitive file system, you will eventually find two distinct bags of bytes that something thinks are the same; if the original disk contains files with both names, you have to pick one or error out. Even if you know the first file system is UTF-8 and the 2nd is ISO-8859-1, the case difference can still be there. It seems to me that the only correct solution is to error out when there are at least two viable versions of a file. Even if you’re trying to “sort and display file names”, you would need to acknowledge at that point that the exact file name isn’t clear. And once your program is doing that, it doesn’t really matter why the two versions are different: maybe it’s a Unicode problem, maybe it’s not.
- goerz 8y agoWould it really be POSIX-incompatible? Does the standard mandate that a filesystem can place no such restriction on top of "filenames are unencoded bytes"? If not, then it's just that tools cannot blindly assume filenames are decodable. Isn't MacOS guaranteeing UTF8 these days, while still being POSIX-compliant?
- deleted 8y ago[deleted]
- dan-robertson 8y agoPart of the issue is to do with bytes and strings being considered totally different by python but confusingly similar to people. The error from "file" != b"file" is particularly bad. It makes sense if you realise that a == b means a,b have the same type and their values are equal. But there is no way a even a reasonably careful programmer could spot this without super careful testing (and who’s to say they would remember to test b"file" and not "file"). Other ways this could be solved are: 1. String == bytes is true iff converting the string to bytes gives equality (but then can == become non transitive) 2. String == bytes raises (and so does string == string if encodings are different) 3. Type-specific equality operators like in lisp. But these are ugly and verbose which would discourage their use and so one would not think to use bytesEqual instead of == 4. A stricter/looser notion of equality that behaves as one of the above called eg === but this is also not great
- xapata 8y ago> The error from "file" != b"file" is particularly bad. It makes sense if you realise that a == b means a,b have the same type and their values are equal. But there is no way a even a reasonably careful programmer could spot this without super careful testing (and who’s to say they would remember to test b"file" and not "file"). I'm of the opposite opinion. I appreciate that b'a' != 'a'.
- dan-robertson 8y agoI don’t think it’s a problem that they aren’t equal. This is reasonable. The problem is that it is hard for one to foresee this error. The mental model of bytes and strings is likely to be either their equal-looking literals or a mental concept of “bytes and strings are basically the same except for some exceptions.” One cannot reasonably trace every variable to figure out whether it is a bytes or a string. The comparison a == b being false comparing strings to bytes makes sense when a and b could be anything. However when b is already definitely a (byte) string, it is more useful to get an error when a has a different type. What is your opinion on numbers: Should 1 == 1.0? What about 1+0j? Or 1/1 (the rational number, although I’m not sure this can be constructed)?
- vfclists 8y agoIf python programmers think they are the only ones with UTF problems, try Lazarus and Freepascal development mailing lists. The debates have going since forever, and I am sure issues will be popping up every now and then. Try Elixir. According to their docs they've had it right from the word go - I think.
- andrewstuart 8y agoIs the author saying that the Python programming language handles this badly, and all other (relevant) programming languages do not? Or is that that Python's attention to detail means that issues that would be glossed over or hidden using ther languages are brought to the fore and require addressing?
- deleted 8y ago[deleted]
- godman_8 8y agoThis is mostly why PHP6 wasn't a thing.
- luckystarr 8y agoAuthor doesn't seem to care that there is a difference between Unicode the standard and utf-8 the encoding. While the changes on the fringes to the system are debatable, they are also in a way sensible. Internal to your application everything should be encoding independent (unicode objects in Py2, strings in Py3) while when talking to stuff outside your program (be it network, local file content or filesystem names) it has to be encoded somehow. The distinction between encoding independent storage and raw byte-streams forces you to do just that! Stop worrying and go with the flow. Just do it as it is supposed to be done and you'll be happy.
- franga2000 8y agoIf you're storing files with non-Unicode-compatible names, you should really stop. Even if on Unix, you can technically use any kind of binary mess as a name, doesn't mean you should. And this applies to all kinds of data. All current operating systems support (and default to) Unicode, so handling anything else is a job for a compatibility layer, not your application. If you write new code to be compatible with that one Windows ME machine set to that one weird IBM encoding sitting in the back of the server room, you're just expanding your technical debt. Instead, write good, modern code, then write a bridge to translate to and from whatever garbage that one COBOL program spits out. That way, when you finally replace it, you can just throw away that compatibility layer and be left with a nice, modern program. In EE terms, think of it like an opto-isolator. You could use a voltage divider and a zenner diode, but thats just asking for trouble.