3 ms·
If you dive head-first into Python's string behaviour, you'll eventually learn the hard UnicodeDecodeError-way what the difference is between a stream of bytes/
by derpadelt 9y ago
If you dive head-first into Python's string behaviour, you'll eventually learn the hard UnicodeDecodeError-way what the difference is between a stream of bytes/octets and a text made of unicode code points.
Much the same as learning that a timestamp without a timezone is not worth much, a text as a stream of bytes is not much worth without the encoding it is in.
PHP also has nice footguns in that area.
- mschuster91 9y ago> PHP also has nice footguns in that area. PHP would deal with uploaded files by itself and write them correctly and directly to disk (unlike some Java implementations like Nexus which buffers in RAM, you can guess what happens). As for ordinary POST/GET parameters, it stores them in a string aka a byte stream which you can then post-process e.g. by translating to UTF-8 based on the browser encoding header. So basically the only way to shoot yourself is if you're doing substr and friends on user input instead of using the mb_ variants.
- wahern 9y agounlike some Java implementations like Nexus which buffers in RAM, you can guess what happens I can't guess. It's 2017 and everybody knows that RAM is unlimited and nobody need ever worry about an OOM condition. Assuming, for the sake of argument, that RAM isn't unlimited, it's unrecoverable, anyhow--modern language designers and Linux kernel architects have made sure of it.
- greenshackle2 9y agoYeah especially in Python 2 you can get into fun messes if you don't really understand what you're doing. Python 2 lets you encode bytes: >>> 'abcd'.encode('ascii') 'abcd' And decode unicode: >>> u'abcd'.decode("ascii") u'abcd' It's nonsensical. Thankfully Python 3 has removed all this madness: >>> b'abcd'.encode('ascii') AttributeError: 'bytes' object has no attribute 'encode' >>> 'abcd'.decode('ascii') AttributeError: 'str' object has no attribute 'decode' (For those not familiar, Python 3 switched around the notation for unicode/bytes. In Python 2 "abcd" is a bytes literal, adding u makes it unicode, in Python 3 "abcd" is a unicode literal, adding b makes it bytes.)
- duskwuff 9y ago> PHP also has nice footguns in that area. PHP, generally speaking, doesn't do Unicode at all. Outside of functions which explicitly do encoding conversions (mbstring, iconv, etc), all "strings" are just handled as a bag of bytes. The main footguns I'm aware of are the "utf8_encode" and "utf8_decode" functions, which actually do lossy UTF8 <-> ISO8859-1 conversions.
- photojosh 9y agoCoincidentally, I just read this 2-yo post on Swift's string handling. Relevant. https://news.ycombinator.com/item?id=10519711 https://news.ycombinator.com/item?id=10519711