3 ms·
This is a neat little optimization which can help you avoid needing to convert everything into utf-8 underneath the hood - leaving things in their raw/encoded f
by mmerickel 12y ago
This is a neat little optimization which can help you avoid needing to convert everything into utf-8 underneath the hood - leaving things in their raw/encoded forms. This may help in certain cases but would probably also make the C-api slightly more complex.
The real problem with Python2 is that it attempts to auto-decode bytes into unicode by guessing the encoding. This actually works most of the time, but not always. Unfortunately the fact that it works at all causes people to ship code that they think is fine... until later when a byte-string comes along with a non-standard encoding and it blows up. Python3 fixed this by making it blow up every time.
- mborch 12y agoI think HTTP request and response headers make a great example. You read a raw stream of bytes line by line (8-bit fixed), and split on newlines. In this protocol we know that each string is "latin-1", so we can extract and decode each header with this encoding. At virtually zero cost, because we don't copy, or transcode the actual string data. And what's more, we stay true to the protocol. These headers were never unicode-encoded (any variant), so it is awkward to suddenly have unicode strings to deal with in the rest of the program. The alternative would be bytes, but in Python 3, those are simply impossible to work with as strings. Why is Python 3 unpopular? It's because it does not really advance Python as a language. It's a bit cleaner around the edges, but the cost was very high for very little gain.