4 ms·
While disappointing, it's perhaps not surprising that the umlaut in "Schrödinger's Cat" caused trouble. The fact that the apostrophe was ruled too risky as well
by h2s 14y ago
While disappointing, it's perhaps not surprising that the umlaut in "Schrödinger's Cat" caused trouble. The fact that the apostrophe was ruled too risky as well, however, is an indictment on software engineering as a profession.
If people are scared to put basic punctuation marks in the names of things, out of fear that badly-written software might break as a result, then that is a sign of just how far we still have left to go.
- heja2009 14y agoI'd argue it is worse that almost all languages can not be used to name things. A pain that many computer users in non-English speaking countries feel regularly. As a German I set all my devices to US English for this reason -although this causes other problems. The apostrophe is really just related to shells and scripting and should be a bit easier to fix.
- praptak 14y agoI agree and disagree. If everything had unicode support, we'd still probably want to limit the codepoint set used for identifiers that are supposed to be unique. Lots of characters have visually indistinguishable glyphs ('x' could be cyrillic kha as well as plain old x) and this already causes problems - see homoglyph attacks.
- estebank 14y agoNotice that the fix was to use "Schrödinger’s Cat" instead of "Schrödinger's Cat". The fix was removing the single quote, as the script checked if there were characters that would run afoul of string escaping. The problem as it is has nothing to do with Unicode and was actually fixed by using a Unicode character. That doesn't mean that there aren't any Unicode bugs in other tools, but that wasn't the issue in this particular case.
- j-kidd 14y agoHandling of special character is usually a bigger problem than unicode. Just today, I tried to name a file with both a single quote and an exclamation point from bash shell. Ended up doing that with a GUI file manager.
- qb45 14y agomv foo \!\'
- pyre 14y agoBetter: $ touch 'j-kidd'\''s file!' $ ll 'j-kidd'\''s file!' -rw-rw-r-- 1 uid gid 0 Apr 11 13:30 j-kidd's file! History expansion will happen on the ! in double quotes: $ ls "!$" ls "'j-kidd'\''s file!'" ls: cannot access 'j-kidd'\''s file!': No such file or directory It won't happen on single-quotes: $ ls '!$' ls: cannot access !$: No such file or directory The only issue is that you can't escape a single-quote within single quotes, so you have to do one of these '\'' (escape a single-quote block, a literal single quote, start a new single quote block).
- qb45 14y agoYep, that's an important and maybe non-obvious behavior of the shell: "directly adjacent strings which are double quoted, "'single quoted or even'\ unquoted\ and\ possibly\ full\ of\ escape\ sequences\ 'get concatenated and count as a single parameter.' Running touch on the above will create one file with very long name.
- kzrdude 14y agoI was going to question you about that, but then I realized I, too, want computers to handle natural language filenames. Or, "don't be your computer's tool!" -- the operator should accept no arbitrary limitations.
- btilly 14y agoI usually wind up calling Perl for that kind of stuff. That also lets you write confusing filenames like "foo\rbar". Which can be really irritating to figure out without a GUI.
- pavlov 14y agoUnfortunately your comment is a sign of just how far we have left to go to get rid of the notion that computers are for use by English speakers primarily, and the rest of the world is an afterthought at best. On 99% of the keyboards I've seen in my life, the middle row reads like this: ASDFGHJKLÖÄ' The apostrophe is considered significant enough to be on that row, but so are Ö and Ä. It's reasonable for users to expect that software can accept the letter that is right next to L on their keyboard, yet there remain software engineers who assume that users won't be surprised that things break if they dare to touch this key.
- h2s 14y agoMy point was that Unicode is newer than ASCII, and that we can't hope to deal with Unicode (Ö) properly if we can't even cope with ASCII (') yet. Nothing to do with anglocentrism at all. I agree that there are lots of annoying computing problems for non-English speakers though.
- pavlov 14y agoWhat I tried to say is that dealing with ASCII is meaningless; it's not even a useful starting point. For the majority of people, a string format that accepts 100% of ASCII but 0.1% of Unicode is just as useless as one that only accepts 95% of ASCII. Therefore the goal should never be to get your ASCII coverage from 95% of 100%.
- kbolino 14y agoThere are two issues here: one is not accepting Unicode properly, and the other is making incorrect assumptions about the content of strings. Both need to be resolved, and all this cultural butthurt is not productive to solving either of them. Resolving the Unicode issues is undeniably a higher priority for speakers of foreign languages, but there are still plenty of languages and libraries whose support for Unicode ranges from nonexistent, to antiquated, to limited, to just plain broken. There's nothing application developers can do in the short term to fix those problems, but by examining their own code and removing fallacious assumptions they can better facilitate the proper handling of Unicode once it becomes available in their underlying technologies. In the mean time, though, they may only have ASCII, or ISO 8859-X, or KOI8-R, or Shift-JIS, etc. to test with.