6 ms·
Disclaimer: as a heavy user of unicode, using it for both French and Japanese, I love it and see how important it is in the world. I just ranted about why ASCI
by Iv 7y ago
Disclaimer: as a heavy user of unicode, using it for both French and Japanese, I love it and see how important it is in the world.
I just ranted about why ASCII is important to programmers: https://news.ycombinator.com/item?id=21760540 https://news.ycombinator.com/item?id=21760540
and this is a perfect example. ASCII has almost a 1-to-1 mapping between screen representation and byte representation. Once you know your font will differentiate between 1 i L | l and 0 O you still need to know a bit about control characters and you are good to go.
Unicode has tons of pages, control characters, diacritics and rendering oddities that make it hard to use like a tool.
I think your approach of considering like bytes to render is spot on. I consider these strings as something like SVG fragments: you can do text operations on them but you have to be confident that you know what you are doing.
- Avamander 7y agoI'm personally incredibly annoyed by just the idea of "Unicode is hard, let's do ASCII", most of the world is non-ASCII, it's just annoying and sad to still see systems that fail when people try to use their native languages. UTF-8 should be the default pretty much everywhere, there are quite easy ways to avoid homograph attacks, those attacks are a poor excuse to discriminate against non-anglosphere.
- antris 7y agoYes, all user-facing text should be Unicode, always. But ASCII has its use cases as well, as the parent comment mentioned in programming. I am natively non-ASCII compatible, but I am glad to program in ASCII and English. The only things that should not be English in written code are domain-specific terms that do not have an official unambiguous English translation. That's the only case to be made for non-ASCII characters in programming that I can think of, but I think romanization can take care of it with most languages.
- ksec 7y agoThis, I think people are arguing for ASCII in programming, not going back to ASCII in general.
- squiggleblaz 7y agoBut aside from APL and a few Haskell coders, who isn't programming in ascii? I mean, sometimes i chuck an emoji into a comment but I don't think that really counts. And how could programming in ascii have solved this bug? The problem is that the strings were compared without giving any thought to what "are these strings equal" is supposed to mean. The only general solution to this kind of bug - that is, to considering distinct emails to be identical - is to have a special function that can check if two emails are identical according to the spec. This function will be different than determining if two names are the same in an American database where dotless and dotful i are the same, and that function will also be different than a function determining if two names are the same in a Turkish database where dotless and dotful i are different. Although I don't know why you're trying to compare strings for identity, you're probably just after similarity anyway - how often do we convert strings to lowercase before we compare them. Generic string equality functions are always the source of bugs, since there's no useful, single uniformally applicable definition of string equality.
- elldoubleyew 7y ago>I think romanization can take care of it with most languages Except for the ones that it really can't. When you try and romanize standard written Chinese it becomes nigh impossible to read for most speakers, often it would be more confusing than an English translation.
- simias 7y agoI'm very much in favor of UTF-8 everywhere, I think the pros outweigh the cons, but: >there are quite easy ways to avoid homograph attacks I'd like to hear about those because as far as I can tell it's still very much an unsolved problem.
- jcranmer 7y agoThe simplest way to avoid most of them is to ban script mixing, especially mixing Latin, Greek, Cyrillic, and Cherokee.
- squiggleblaz 7y agoSo in a secure environment I can't write words in IPA? IPA requires mixing Greek beta and theta with Latin letters (but there's a Latin/IPA phi just for fun). It also doesn't help ɡ vs g. It sounds like any genuine solution is either not simple or excludes legitimate use cases.
- cdirkx 7y agoYes, the Unicode Technical standard identifies the Latin, Greek and Cyrillic (most of IPA characters) scripts as confusable, and recommends against mixing them in identifiers. Use in general text is fine though.
- cdirkx 7y agoThe Unicode Technical Standard [1] (different from Unicode Standard) recommends treating identifiers (filenames, variable or fuction names, email adresses, usernames, etc.) different from normal text. There is a special class of 'identifier characters' which already exludes a lot of sneaky characters like invisible punctuation and obscure scripts that are not in modern use. Additionally there are 5 additional restriction levels for identifiers depending on your specific situation: 1. ASCII only 2. single script 3. single script or Latin+{Japn, Hanb, Kore} 4. single script or Latin+{any, excluding Cyrillic, Greek} 5. Not containing any characters in the recommended blacklist of characters for use in secure contexts 6. No restrictions other than normal identifier restrictions For any specific strings that might be confused, there is an algorithm to compute the visual 'skeleton' of a string and match it against that of another string to test if they are confusable. [1] https://www.unicode.org/reports/tr39/ https://www.unicode.org/reports/tr39/
- iforgotpassword 7y ago> there are quite easy ways to avoid homograph attacks, those attacks are a poor excuse to discriminate against non-anglosphere. No. There are these kind of attacks every now and then to this day. Maybe of you're not following itsec they fly under your radar, but getting this right is exceptionally hard. And besides these kind of attack, every major os had multiple bugs just in the processing of Unicode that could at least be used for DoS attacks. So saying it's easy to avoid any sort of abuse of Unicode seems quite ridiculous. Go ahead and support it for messages, display names and whatnot, but for the love of god, limit the login name of users to ASCII. Don't assume that your Python/Go/JavaScript lib for Unicode handles sanitizing and canonicalization properly. It doesn't. And even if it has only a minor bug that doesn't lead to direct issues, the next update of the lib might fix the problem and now you have to deal with the fact that your db might contain data that was processed with the old faulty lib and now gets compared to the properly processed output of the new version. Just don't. Use it as opaque data for displaying, as GP said, but never as an identifier for anything.
- knolax 7y agoI swear software's biggest problem is that developers prioritize their own ergonomics over actually producing functional code. Imagine if auto engineers just decided "safety is hard, let's just make deathtraps". I know the current narrative is that the 737 Max failed because of MBAs but it's a really big indictment that out of all the complex and hard to engineer systems in an aircraft it was poor software that caused the crash.
- bregma 7y agoIt was a bad corporate ecosystem that caused the crash. The people at fault were the accountants that tweaked lines un Excel until they got the numbers they wanted amd then pressed reality into the service of those numbers. Bad software didn't eliminate pilot training on the new systems. Bad software didn't shortcut the required FAA safety certification on the system. Bad software didn't reduce the number of redundant sensors that fed the system. Bad beancounters did all that. C-levels in suits who got million-dollar bonuses for killing 346 people to gain marginally increased quarterly results. Don't blame the software.
- jsjohnst 7y ago> really big indictment that out of all the complex and hard to engineer systems in an aircraft it was poor software that caused the crash While I frequently point out how bad we are as an industry at making fault tolerant code, this part is just flat wrong. The software portion of the 737, while definitely flawed in a catastrophic way, would not have came to be if the aeronautics engineers had done their job and designed a flight worthy plane without software hacks. Not to imply the aeronautics guys are the root cause either though, the 737 Max fiasco is a top to bottom completely failure of Boeing as a whole, virtually every department involved in the Max has a significant reason to share in the blame.
- rsync 7y ago"I'm personally incredibly annoyed by just the idea of "Unicode is hard, let's do ASCII", most of the world is non-ASCII, it's just annoying and sad to still see systems that fail when people try to use their native languages." You are correct that most of the world is non-ascii and I am enthusiastic about recognizing that diversity and I am willing to pay certain costs in return for the richness it provides. However, I will point out that certain systems have been deemed crucial and in need of deliberate (and brutal) dumbing-down. Specifically, I speak of the global Air Traffic Control system that is English only[1]. Tagalog/Flemish/Satsugu is hard. Let's (land airplanes with) English. [1] https://en.wikipedia.org/wiki/Aviation_English https://en.wikipedia.org/wiki/Aviation_English
- Iv 7y agoI want to be able to collaborate with Japanese, Chinese, Indians, Russian, Arabian, Malay programmers. Even if we somehow live in a dystopia where all the english-speaking world disappears in a magic puff, I think I will continue to shut down my native french and communicate with other fellow programmers in English ASCII. I wish more programmers understood the value of having a lingua franca is. UTF-8 should be used everywhere on the user-facing side, but under the hoods, you want unambiguous representation of strings and code. It is easy to avoid homograph attacks if you do have a non-unicode representation available to you.
- tim333 7y agoI was thinking that. Unicode is great for presenting text to users but for file names, email addresses, code and the like ASCII has a lot going for it.
- BoppreH 7y agoIt's not just Unicode weirdness, it's the whole concept of string operations. Examples: 1) Naive CSV libraries that break when given a field containing a comma itself. Same for injections. 2) Indexing things by strings (file names, user names, etc) relies on humans typing and reading with 100% accuracy. 3) Control characters are still present. Try to generate a file name containing every ASCII character (from 0x01 to 0x7F), then try to delete it in different ways. It's really frustrating. 4) Even string templates can cause problems. MMO's used to have game masters identified with "[GM] Character name", until scammers started using the same pattern. The sin here was to concatenate "[GM] " with the character name, which is spoofable, instead of using a badge or different color.
- ChrisSD 7y agoFor 3), that's not the case on Windows filesystems as these characters are not allowed in file and folder names. On Linux, in addition to control characters, the shell interprets the `-` character specially. This means that if a filename starts with `-` you may have surprising results even if the filename is quoted.
- nybble41 7y ago> On Linux, in addition to control characters, the shell interprets the `-` character specially. It's not the shell that interprets the leading `-` character specially, but rather the program receiving the string. The convention of marking the end of option processing with `--` helps, for programs that support it; you can also prefix relative paths with `./`.
- ComputerGuru 7y agoNo, it doesn’t. What’s a typical corporate email address? First.Last@...? F.Last@...? More than half the world’s population does not use the Latin alphabet.
- atoav 7y agoThe longer I program, the more I am convinced that falling back to ASCII in 90% of (non-embedded) use-cases is just a programmer avoiding to do the extra brainwork to deal with encodings. That is why I like the Rust approach: make these issues front and center, implement clear solutions for common use cases and enforce them. In to many languages encodings feel like an afterthought rather than something that has been considered from the start. When I started using Rust I e.g. learned that the Strings a OS uses in its filenames are not necessarily valid UTF-8. Rust forces you to handle this explicitly. Strings are complex and hiding this can be dangerous
- jiggawatts 7y agoThe issue I see with rust is that they (partially) repeated the mistake older operating systems made: they assumed that the Unicode of that time will be Unicode forever. While UTF8 is just an encoding, a Rust UTF8 "String" is actually a UTF8-encoded Unicode 11.0 string. Or version 12.0, possibly 12.1. Maybe 13 soon! Who knows. A moving target, certainly. If you're "agile", you can just recompile and you will be fine, the next Rust version will surely support any changes in the Unicode standard. But once a Rust program or crate is no longer actively maintained, many small assumptions will be baked into its Unicode handling at a much lower level than, say, the typical Windows C++ program that uses the OS-provided dynamic libraries for string processing. Who knows, maybe the consortium will never introduce breaking changes, but I suspect that in a decade or two people will be cursing Rusts too-strong integration with the Unicode of the 2010s...
- jcranmer 7y agoThe Unicode consortium has guaranteed that there are many properties they won't change once they're released. Even if the property value is objectively wrong. More to the point: the standard library of Rust relies on no properties of Unicode that will change--there's no builtin normalization, case folding, grapheme cluster support, etc. So there's no data tables in the standard library that need to be updated. The only assumption Rust makes of Unicode is that no assigned codepoints will be changed, and that the space will not grow beyond the current limit of U+10FFFF.