5 ms·
Strings – Dive Into Python 3
- emiljbs 13y agoApparently, all I know about strings is correct.
- sluu99 13y agoMy thought exactly too haha
- Roboprog 13y agoAnd, he didn't even touch on the horror that is EBCDIC. Once you've had to touch that, the idea of "code point" for a character is something you can't ignore, hoping that things just work "most of the time" -- ASCII A != EBCDIC A.
- sluu99 13y agoAre there still (many) of people using EBCDIC?
- Roboprog 13y agoYes. High volume printing is often done using IBM's AFP/MODCA print language, which typically has the text in IBM's EBCDIC encoding. (disclosure: I once worked at a company that made tools to port code and data from IBM minicomputers to Unix & MS platforms, and also at the largest [format,] print & mail shop in the US)
- Roboprog 13y agoOr the 6 bit funky set that the old CDC Cyber mainframes used to use! ("What is this lower case 'a' you speak of???")
- reeses 13y agoEBCDIC A != EBCDIC A :-) It's easier to read text on punchcards in EBCDIC of whatever variation, though.
- ableal 13y ago"Thank you, Mark Pilgrim". (Unless you already knew it ten years ago. Back when he wrote "Everything you thought you knew about strings is wrong." it was quite true of most every programmer, and it was thanks to this piece and similar ones, by Spolsky and others, that the information got spread around. There must be some phrase for this opposite of the "self fulfilling prophecy": the cautionary phrase that causes itself to become false in the future ;-)
- ygra 13y agoAs for me, I found Spolsky's article lacking, too. But lurking for years on the Unicode ML is probably not something most people do. You learn a lot there, though.
- sluu99 13y agoTL;DR: think of string as tuple of numbers. some are bytes, some are integers. if you want to transform those numbers into a particular encoding (e.g. UTF-8, CP-1252) then that's a different story. EDIT: I know the article went all out about character abstraction, that why i said "some are bytes, some are integers"
- morpher 13y agoThis is entirely the wrong take-away message from this article. The point is that strings are not sequences of numbers, but are, rather sequences of characters. Characters are abstracted from the underlying byte representation which is unimportant when dealing with strings. For situations where a concrete byte representation is needed, you can get one by encoding the string.
- BorgHunter 13y agoEven this definition can get hairy, though. What is a character? Is 'á' one character or two? Most human beings would say one, but in actuality I formed it with an 'a' (U+0061) and a combining acute accent (U+0301): Two separate code points. But you can also get the same result with 'á' (U+00E1); this is not true of all combining character combinations. In the past, I've had to deal with horrible mashups of fixed-byte-length columns in flat text files with UTF-8 bolted onto it. In Java, no less. Trying to figure out how to deal with all the edge cases (how do you truncate a string when the boundary is between a "normal" character and a combining character?) was an endless parade of the bizarre. Strings are hard, fundamentally.
- deleted 13y ago[deleted]
- gruseom 13y agoThe epigraph to that chapter is brilliant.
- moreati 13y agoA Friday challenge: In Python when is u'ß'.upper() equal to u'SS'? I discovered one case today, there may be others. Answer: https://twitter.com/moreati/status/332910618858364928 https://twitter.com/moreati/status/332910618858364928
- kzrdude 13y agoAnd will it equal 'ẞ' in a later update? http://opentype.info/blog/2013/04/22/capital-sharp-s-in-use/ http://opentype.info/blog/2013/04/22/capital-sharp-s-in-use/
- safod 13y agoActually, this is a bug. Unicode codepoint U+1E9E is LATIN CAPITAL LETTER SHARP S and should be the result of u"ß".upper(). This is especially so because otherwise u"Maße".upper() (Maße means measures) returns "MASSE", which could be confused with u"Masse".upper() (Masse means mass). In such cases, where confusion is possible and no uppercase ß is available, the German dictionary Duden actually suggests using SZ instead. Therefore, u"Maße".upper() would have to return "MASZE". However, since the Python string processing routines can hardly carry a dictionary around just to check whether there is a similar word that would have the same uppercase spelling, this is obviously not feasible. U+1E9E would be the way to go.
- ygra 13y agoAccording to Unicode ß gets converted to SS in uppercase. This is by definition and doesn't change (stability policies, as far as I recall). Even in German you'll never see ß capitalised as SZ (except when I do it, but I'm a very, very small minority – and now I'm more likely to use ẞ).