4 ms·
Getting Unicode right, especially with various file systems and cross-platform implementations is hard, for sure. But, I think this quote: "And, whatever you
by TimJYoung 8y ago
Getting Unicode right, especially with various file systems and cross-platform implementations is hard, for sure. But, I think this quote:
"And, whatever you do, don’t accidentally write if filetype == "file" — that will silently always evaluate to False, because "file" tests different than b"file". Not that I, uhm, wrote that and didn’t notice it at first…"
shows a behavior that, to me, is inexcusable. The encoding of a string should never cause a comparison to fail when the two strings are equivalent except for the encoding. For example, in Delphi/FreePascal, if you compare an AnsiString or UTF-8-encoded string with a Unicode string that is equivalent, you get the correct answer: they are equal.
- kbumsik 8y agoYeah, this behavior might be because Python doesn't store unicode string as it is. AFAIK Python always store string as an array of fixed-sized bytes for random access. In other words, the size of an element of the array is the same as the maximum size of characters in the string, meaning that even a character of 1 byte ASCII can be stored as 4 bytes. So when one side is bytes (filetype in this case) and the other side is a string, the underlying byte representation can be different even if they represent as the same string in higher level.
- ubernostrum 8y agoPython as of 3.3 chooses an internal representation on a per-string basis. This encoding will be either latin-1, UCS-2, or UCS-4, and the choice is made based on the widest code point in the string; Python chooses the narrowest encoding capable of representing that code point in a single unit. This does mean that a string which contains, say, some English text and an emoji will "blow up" into UCS-4, but the overhead isn't that severe; most such strings are not especially large. It also means that strings containing only code points < U+00FF are smaller in memory on Python 3.3+ than previously, since prior to 3.3 they would be using at least two bytes per code point and now use only one.
- mikezter1 8y ago> The encoding of a string should never cause a comparison to fail when the two strings are equivalent except for the encoding. You'll have to admit that the encoding is a property of a string, just like the content itself. As always, you as a programmer are bound to know both of these properties to have predictable results. To compare two strings of different encoding to one another, you'll have to find a common ground for interpreting the data contained in the string. If you don't want or need that, then all you have is a "string" of bytes.
- TimJYoung 8y agoSure, but you can have defined rules about what happens when you compare values with disparate encodings, similarly to how you have to have rules about how column expressions are compared in SQL with regard to their collations. The way such things are done is typically to coerce the second value into the encoding of the first value, and then compare the two values. What the Delphi compiler does is issue warnings when there might be data loss or other issues with such coercions so that the developer knows that it might not be safe and that they might want to be more explicit about how the comparison is coded.
- deleted 8y ago[deleted]
- Terr_ 8y ago> The encoding of a string should never cause a comparison to fail when the two strings are equivalent except for the encoding. If you mean comparing "file" == b"file" that's not possible on several levels. Firstly, even if you say "just compare the bytes", the computer doesn't know what byte-format you want for "file". Sure, it's "Unicode", but is it UTF-8 or UTF-16 or what? Those choices will produce different results, and the computer cannot accurately guess the right one for you. Secondly, that violates Python's normal rules by introducing type juggling. It's equivalent to asking for expressions like ("15"==15) or ("True"==True) to work, and involves all the same kinds of long-term problems. (Don't believe me? Work in PHP for a few years...)
- TimJYoung 8y agoAs I said, Delphi/FreePascal have been handling this for years without issue, so "not possible" sounds like you're giving up a little too early. The encoding of a string is not the same thing as its type, and shouldn't be treated as such. Python has to know the encoding of the "file" string by the encoding off the source file. It then also has to know what the encoding of the b"file" string is because it is explicitly specified. That's all of the information that it needs to make the comparison, so it should either a) issue a compilation/runtime error if its an invalid comparison, or b) return a proper comparison result. Returning an invalid comparison result is the worst of all possible outcomes. As for character sets/code points: https://stackoverflow.com/questions/130438/do-utf-8-utf-16-and-utf-32-differ-in-the-number-of-characters-they-can-store https://stackoverflow.com/questions/130438/do-utf-8-utf-16-a... A byte string is simply a string that is using the lower ASCII characters (< 127). The code points for "file" map cleanly to the same code points in any Unicode encoding.