4 ms·
> if you’re preparing your code for others to read—whether on screen or on paper—skip the ligatures. The article should have led with this. Instead this main c
by ketzu 3y ago
> if you’re preparing your code for others to read—whether on screen or on paper—skip the ligatures.
The article should have led with this. Instead this main caveat is buried in a paragraph, where it is preceded by, what boils down to, "do what you want, but you will recognize its wrong someday." Of course people will focus on that part and on the most prevalent use of ligatures in programming: In their own local editor. (Where people usually have strong opinions about what's the right thing to do.)
The tangent in (1) on how they contradict unicode could have been skipped as well, just go straight to the "the problem is" part. Why? Unicode already "contradicts" itself: △ is WHITE UP-POINTING TRIANGLE (trine) or U+25B3, which I am sure everyone can easily distinguish from the mentioned Δ Greek Capital Letter Delta U+0394 especially when used outside a clear context. (There are other symbols that are easier confused, but I find them close enough for the point, especially because fonts might render them differently.)
For private use, I find ligatures make spotting negations and reading comparisons easier more often than they are wrong. But I see why ligatures might not be a good choice for a presentation or code blocks on a website.
- Fatnino 3y agoVscode highlights characters that might be confused for other characters if you use anything other than the common variant. I'm not sure if this is default behavior or something I did to it over the years. I can see it being helpful but it feels way too trigger happy. It's always notifying me about Hebrew letters that only look similar to other things if you take off your glasses, squint, and change the font.
- 3836293648 3y agoYou don't even have to go as far as a triangle. There are two sets of the latin alphabet in unicode and two sets of the greek. The normal one and the maths one. And the mu on your keyboard is not the same on a windows machine and a linux machine. (Alt Gr + m, at least in my language)
- hakre 3y agoFor portable code, use UTF-8 only, and only the part in ASCII. This also helps to distinguish digits from alphas etc. Unicode is a different code. (and as I learned today, Julia code might be exclusive here)
- Macha 3y agoUTF-8's using only characters from the ASCII subset is just ASCII, this was kind of the point of UTF-8 at inception.
- hakre 3y agoYes, exactly, and Unicode has more encodings than UTF-8 so stating it explicitly makes clear which encoding of Unicode that is. Writing only ASCII would not be relaying on Unicode, but on ASCII only. IMHO a bit shortsighted today.
- Macha 3y agoUsing UTF-8 but only the ASCII bits, and using ASCII are the exact same activity, however. I don't see how the first is any different to the latter, other than which document you look up to tell you that 00110000 is "0". The difference between an "ASCII subset of UTF-8" parser and ASCII parser is whether the error message on encountering a high bit set to 1 is "Invalid character" or "Character not in permitted ranges". If your point is that your program should use Unicode internally, this is already true for most programming languages but is independent of your input, plenty of languages routinely converting UTF-16 to UTF-8 to work with Linux or the web when they use UTF-16 internally, or UTF-8 to UTF-16 to work with Windows when they use UTF-8 internally. But I'd argue if you've done that work already, what harm is there in allowing your users to type € or á or 風?
- hakre 3y ago> Using UTF-8 but only the ASCII bits, and using ASCII are the exact same activity, however. Which was my point. The encoding might be different. For ASCII 7 bit is fine, for UTF-8 and only ASCII, 8 bits are required and, as you also already point out, the high bit must be set to zero with it. So the point is to name the encoding (UTF-8) of the ASCII characters. And the input was portable code, not user input of a program.
- kps 3y agoΔ vs ∆ is a slightly better example, since traditional printed mathematics just used the former. (Not making you look: U+0394 GREEK CAPITAL LETTER DELTA vs U+2206 INCREMENT)
- mananaysiempre 3y ago> The tangent in (1) on how they contradict unicode could have been skipped as well Not only because confusables already exist, but also because (as I said[1] the previous time this was posted) covering all ligatures used in all typographical styles is very much a non-goal of Unicode. The official position is that the font shaping layer[2] sits atop Unicode’s semantic representation and is free to ligate, spindle, or mutilate it for display however it prefers (at least for Latin, Greek, and Cyrillic it’s a preference; other scripts can’t be rendered at all without doing it, such as Arabic—barring the legacy presentational forms—or Burmese[3]). The only reason Unicode even has those ligatures is that some IBM encodings (which were more presentational in nature) encoded them, and the IBM employees who wrote a large part of the early standard (based on the decades of i18n experience they had at that point) wanted roundtripping. [1] https://news.ycombinator.com/item?id=29639966 https://news.ycombinator.com/item?id=29639966 [2] https://github.com/n8willis/opentype-shaping-documents https://github.com/n8willis/opentype-shaping-documents [3] https://r12a.github.io/scripts/mymr/my.html#compositeV https://r12a.github.io/scripts/mymr/my.html#compositeV