13 ms·
Mimic – abusing Unicode to create tragedy
- r721 11y agoIn cases like those I use unicodelookup.com to list suspicious characters :)
- rbinv 11y agoI guess someone should develop an IDE/editor plugin that marks non-ASCII characters outside of string literals.
- lazyjones 11y ago> marks non-ASCII characters outside of string literals. Many programming languages support non-ASCII variable name characters now.
- rbinv 11y agoGood point, didn't consider that. Although any good IDE/editor should catch the use of "undeclared" variables and functions.
- TorKlingberg 11y ago> Many programming languages support non-ASCII variable name characters now. Just because you can do something doesn't mean you should. It is usually worth keeping variable names and such in English in enable international collaboration. Also non-ASCII source files can get mangled in transit.
- kuschku 11y agoWell, it does happen, though – look at this weather data from a large German newspaper, it is in a custom format ('|' separated values) and in German: http://wetter.bild.de/data/meinwetter.txt http://wetter.bild.de/data/meinwetter.txt It happens all the time, everywhere, that people write code and stuff in their native language.
- Dylan16807 11y agoThat's data. The suggestion is about variable names.
- kuschku 11y agoWell, the variable names of Bild.de (for example HTML class names) are also in German. It happens all the time, everywhere.
- halostatue 11y agoRuby now allows (some) Unicode glyphs as names (allowing for things like Δv). 08:11:32 >> Δv = 3 => 3 08:11:39 >> p Δv 3 => 3 My solution when I have problems like this is to start building a negative regexp in vim: /[^-a-zA-Z0-9 \[\]] I then add other symbols as I find them. I can usually find the illegal characters in about 30 seconds this way—and I can add the non-ASCII glyphs that I expect to be present to my regexp.
- metasean 11y agoJavaScript allows a wide range of Unicode characters - http://stackoverflow.com/a/9337047 http://stackoverflow.com/a/9337047
- archimedespi 11y agoVim plugin: https://github.com/vim-utils/vim-troll-stopper https://github.com/vim-utils/vim-troll-stopper
- Svenstaro 11y agoWow, now that's just pure evil.
- torgoguys 11y agoYes, seriously. This is why we can't have nice things.
- cstross 11y agoAlso GREAT if you're trying to identify untaken phishing domain names to register for your next scam!
- sheraz 11y agowouldn't you end up with the 'xn--' ascii expansion in the url window?
- sschueller 11y agoMost modern browser will show you the unicode version.
- deleted 11y ago[deleted]
- germanier 11y agoNowadays most modern browsers will revert to the punycode ("xn--") if there is any chance of confusion, cf. https://en.wikipedia.org/wiki/IDN_homograph_attack https://en.wikipedia.org/wiki/IDN_homograph_attack
- treerock 11y agoIn Chrome, probably [1]. Other browsers don't seem to be as strict. [1]https://www.chromium.org/developers/design-documents/idn-in-google-chrome https://www.chromium.org/developers/design-documents/idn-in-...
- mikewilliams 11y agoHyperlink i.e http://.com http://.com How many people will double check the url bar and notice the url is actually http://www.xn--m3haa.com/ http://www.xn--m3haa.com/? Not all of them. :edit+seems like the three umbrella unicode symbols are not supported on hn, are they supported in e-mails?
- Tepix 11y agoA lot of unicode characters are blocked for domain names for this exact reason.
- sheraz 11y agoThere is a special place in hell for anyone doing this. I'm going to watch this repo and blacklist pull requests from anyone who forks it :-)
- creshal 11y agoThey share the place with coding blogs that use instead of spaces for code snippets.
- minikomi 11y agoOr “real” quotes in code examples..
- scintill76 11y agoOr en-dashes instead of double hypens for command line flags... I can feel my blood pressure rising just thinking about it.
- zaptheimpaler 11y agoDamn you Google docs! they also do this :(
- TorKlingberg 11y agoI don't think people do this intentionally. Either the code snippet has passed through MS Word (why?) or their blog tool is being "helpful".
- harryc2011 11y agoIt happens if you paste the code into an Outlook email (which used Word as the rendering engine IIRC)
- jff 11y agoI once had to do a team project with another student who did all his coding in Wordpad, god knows why. His indentation was more or less random. I wanted to murder him.
- adrianN 11y agoOn a Mac you (used to?) get a non-ascii space when you hit the space bar while holding Alt or something like that. Easy to fat-finger it in any case and looks the same in most text editors. It's a great source of fun for novice Mac-using programmers to find out why the compiler complains.
- amadahy 11y agoThis is still happening as of today: ps aux | grep foo zsh: command not found: grep It happens to me at least every other day.
- mcculley 11y agoThe Commodore 64 (or some other machine from my childhood) would generate a non-ASCII space if you held down control (or shift maybe) when pressing it. To this day I'm careful about that. I didn't know it was still a possible problem.
- kps 11y agoIt's not terribly difficult to define custom keyboard layouts for OS X. Make a copy of your preferred layout and get Ukelele [sic] from SIL¹ to remove NBSP from Option-Space. (Or just hand-edit the XML changing " " to " ".) ¹ http://scripts.sil.org/ukelele http://scripts.sil.org/ukelele
- naggie 11y agoIt still does! You've just explained something that has been annoying me for about a year now. Thanks!
- thousande 11y agoA good IDE would pick this up. I had the habit of pressing the ALT key a bit ahead of time before an OR operator. if (foo || bar)
- mattlondon 11y agoI had something similar happen in the wild to me. I work for a "major search engine" that does a lot of advertising & marketing stuff. To get the most out of it, we need customers to implement some javascript on their ecommerce sites. As is often the case, javascript code that needs to get implemented on an ecommerce site often gets copy-pasted or emailed around a lot internally within a customer before it reaches the right person who can add it to the site's pages. In this example somewhere along the way, a normal javascript snippet got all of the semi-colons changed from ; to ;. In case you've not already spotted it, ; is not a ; but is actually "Greek Question Mark" (http://www.fileformat.info/info/unicode/char/037e/index.htm http://www.fileformat.info/info/unicode/char/037e/index.htm). It was very confusing why Chrome was moaning about a semi-colon an illegal token. I had a genuine "Am I going mad? Seriously?" moment before I realised what was happening.
- slowmotiony 11y agoI would probably have to quit my job before I could figure out that problem. May I ask how did you spot it?
- slig 11y agoSomething similar happen to me, but with that fancy quotes. I spotted by doing a "binary-search weird bug hunt". Cut half of the code off and see it it's still complaining, if it's, cut the other half, and so on.
- jprince 11y agoAh, that old stand-by. Always warm and ready for the odd unicode hell-bug.
- sethammons 11y agoI had a similar problem. Copied a code snippet. Ruby started complaining about an undefined function. After nearly going mad, and then looking at the source through a hex editor, you could see Unicode whitespace. I have yet to forgive ruby or Unicode whitespace. Or the chat utility from whence I copied.
- gnud 11y agoThis can actually be used productively, to see how your app reacts to weird input :)
- n-gauge 11y agoJust chuck the code into an XML validator. Any character > 127 will be flagged as invalid.
- pierrec 11y agoI can foresee a new phenomenon arising in stackoverflow-style sites and coding discussion forums: "My simple piece of code looks perfect and should work without problems. Yet it won't compile! Help!" Answer: "Try running `./mimic --reverse` on your source."
- po1nter 11y agoThat should be a comment.
- pierrec 11y agoTrue, except in the rare situation where the question really features a code sample containing some evil homographs. Edit: Oh hey, I actually found one that fits the bill: http://stackoverflow.com/questions/14925894/trouble-with-argc-and-argv http://stackoverflow.com/questions/14925894/trouble-with-arg...
- probably_wrong 11y agoI actually almost submitted something in that vein once. I'd type > ls | wc -l and get > bash: wc: command not found As it turns out, I need Alt+1 to type a pipe character in my keyboard. If I'm not quick enough releasing the Alt key, I'll type Alt+Space instead of just Space, which inserts a Non-breaking space[1] in Mac. This character is not a space, and therefore it gave me a weird "command not found" error. This lasted for months until I found out what the problem was - given that it was a combination of my keyboard settings and OS, finding the root of the error took quite some time. The hint? The "command not found" error had an extra space in front of the unknown command. [1] https://en.wikipedia.org/wiki/Non-breaking_space https://en.wikipedia.org/wiki/Non-breaking_space
- TazeTSchnitzel 11y agoThis bit me as well, as I mentioned in a comment above. The solution I found best was to make OS X not produce an NBSP on alt+space.
- wodenokoto 11y agoSo, one might wonder why these homo-graphs have different code points. After all the French A and the English A share the same code point. It's really difficult to do the right thing here. If Greek question marks share code point with semi-colon, it obstructs search and replace for question marks. Subtle differences in how Japanese and Chinese are written has led to differently written characters sharing the same code point. It's nice that you can easily look up most Japanese characters in a Chinese dictionary and see how they are used in China, but it has become frustratingly hard to get subtleties in their written form right. The Chinese version may have the line strike through another line, while the Japanese only has it touching. I honestly don't know how to go about posting how to same code points have different written forms! But it seems like it would be nice if code editors warned about text outside ascii. You usually only want that in strings and comments.
- ant6n 11y agonow that we gave up on ucs-2, we could re-encode those overloaded Japenese/Chinese characters as separate Japenese and Chinese characters on astral planes like the supplemental multilingual one.
- rspeer 11y agoThey're doing that. That's what a lot of plane 2 is.
- lazyjones 11y ago> It's really difficult to do the right thing here. If Greek question marks share code point with semi-colon, it obstructs search and replace for question marks. Context is the key here. Greek text doesn't use the semicolon for other purposes and searching/replacing such single characters in source code is a terrible idea anyway (think comments, string literals...). So what is the prohibitive failure scenario here? Indistinguishable (for humans) characters with different code points were a stupid idea, it's fine to abuse it in order to point out that fact.
- wodenokoto 11y ago
- deleted 11y ago[deleted]
- motti 11y agoThis sort of stuff can be the basis for many XSS attacks, see http://websec.github.io/unicode-security-guide/character-transformations/ http://websec.github.io/unicode-security-guide/character-tra... For instance, \u2329, \uFE64, \uFF1C and \u3008 can be best-fitted automatically to \u003C (the regular '<' mark in HTML)
- lisivka 11y agoIt is also good tool to check is Unicode supported well: just convert all user visible messages and then check interface of the program for <?> or [].
- ant6n 11y agoOne could name variables and functions to later identify whether code was copied (e.g. to find out whether somebody copied some GPL code).
- Kristine1975 11y agoNote to self: Run mimic --reverse on GPL code I copy.
- pmlnr 11y agoThis somewhat reminds me if this little entry on how "tolerant" JavaScript is... https://mathiasbynens.be/notes/javascript-identifiers https://mathiasbynens.be/notes/javascript-identifiers
- sly010 11y agoIronically I have weird OCD where I always assume I made a typo, so I keep deleting and retyping code a few dozen character at a type, often in lines where I see nothing wrong. Over time this has just become something my hands do whenever my brain needs time to think about something else. So in a way I developed natural immunity to said unicode tricks ;)
- jobigoud 11y agoI think you're not alone. A common error I've noticed is when you make a typo somewhere (that compiles) and copy and paste it in a different place where you have the correctly named symbol. It's often hard to see the typo because the eye fly over the word. So you erase and type it manually.
- SCHiM 11y agoHey I do the same! Only with variables, and almost always when using array indexes that are not 'i' or 'x'. Sometimes is annoys me, and I've noticed that it's worse when I use CamelCasing and not as bad when i_do_this.
- deleted 11y ago[deleted]
- cevaris 11y agoSome people just want to see the world burn...
- austinjp 11y agoI'm reminded how very useful I've found Text::Unidecode in the past. http://search.cpan.org/~sburke/Text-Unidecode-1.27/lib/Text/Unidecode.pm http://search.cpan.org/~sburke/Text-Unidecode-1.27/lib/Text/...
- avian 11y agoAuthor of Python port of Unidecode here. I wrote a comment previously, pointing out that Unidecode does the reverse of Mimic. But then I actually checked the tables of characters that Mimic uses and deleted my comment. Mimic chooses replacement characters solely based on their visual similarity with ASCII. Unidecode, while still doing character-by-character replacements without deeper analysis, tries to optimize the replacement tables for transliteration of natural languages. For example, mimic will replace Latin capital H with Greek capital eta (U+0397), because they look similar. However, Unidecode will replace U+0397 with Latin capital E, because Latin E is typically used in place of Greek eta when transliterating Greek text to Latin.
- Drdrdrq 11y agoI have used the php port long ago when creating a simple website search engine... Great project!
- AUmrysh 11y agoI think the line about "Mimic substitutes common ASCII characters for obscure homographs" has it backward. Shouldn't it say Mimic substitutes obscure homographs for common ASCII characters?
- pascalmemories 11y agoNever occurred to me before, but here "substitutes" reads to me as being commutative. I read both as having the same meaning. (i.e. you end up with unicode homographs replacing your ascii) Just me?
- dmd 11y agoIn that case, you won't mind if I substitute poison for your favorite tasty beverage.
- kps 11y agoTechnically, my favorite tasty beverage is poison.
- pbhjpbhj 11y agoSubstitute works IMO the same way replace does[1]. Substitute poison for healthy food. Substitute poison with healthy food. The first means you take away healthy food and give poison. The later means you take away poison and give healthy food. [1] Except that "replace X for Y" sounds weird, except in the common phrase "replace like for like" (and probably some others!).
- mafro 11y agos/for/with
- lstamour 11y agoTechnically, it does both ;-)
- hollerith 11y agoNot to my ear. To my ear, substituting X for Y is the same thing as replacing Y with X.
- grabcocque 11y agoYOU ARE A TERRIBLE PERSON AND I LIKE YOU
- acdha 11y agoMac users might appreciate the great UnicodeChecker: http://earthlingsoft.net/UnicodeChecker/ http://earthlingsoft.net/UnicodeChecker/ It offers a convenient utility to diff arbitrary strings, which is also quite handy for e.g. detecting normalization discrepancies, and installs a service so you can highlight a character in any app and use “Display character information” to see what it actually is. I have Python command-line version in my PATH which displays the character info for arbitrary input strings: https://github.com/acdha/unix_tools/blob/master/bin/unicode-characters.py https://github.com/acdha/unix_tools/blob/master/bin/unicode-...
- acdha 11y agoUsing the Taylor Swift example from https://news.ycombinator.com/item?id=10438363 https://news.ycombinator.com/item?id=10438363 in the comparison window looks like this: https://www.dropbox.com/s/9j9h5rjt4gu22hb/Screenshot%202015-10-23%2014.49.30.png https://www.dropbox.com/s/9j9h5rjt4gu22hb/Screenshot%202015-... Each hex value shown can be clicked to open the Unicode character info for that codepoint
- ehosca 11y agoi smell a Notepad++ extension
- b0ner_t0ner 11y agoTaylor Swift? Never heard of her: https://www.google.com/search?q=Τаylοr+Ѕwіft https://www.google.com/search?q=Τаylοr+Ѕwіft :D
- pierrec 11y agoKind of surprised at how poorly Google handles this (I would have expected at least a correction suggestion)! Heck, it might open the door for an obscure blackhat/phishing technique...
- DannoHung 11y agoMaybe some sort of extortion scheme? Send an email to a small business person that isn't very technically savvy, say you have just erased all the search results for their business from Google, provide link, demand a Bitcoin to return the results. Maybe a low hit rate, but if you could automate it, you could run the scam on a lot of places.
- mfoy_ 11y agoSimilar in theme to that trick of sending strangers that "link to your facebook page" (http://facebook.com/profile.php?=73322363 http://facebook.com/profile.php?=73322363)
- omgtehlion 11y agoSlightly tangential: In Russia there is a government procurement portal. Where gov organizations have to post their requests to enforce competetion and best prices. The usual tactics [1] of corrupt officials was replacing cyrillic (russian) letters with respective latin homoglyphs so only affiliated companies can find and win this contract. [1] http://www.bbc.com/russian/rolling_news/2013/04/130409_rn_state_auctions_improve http://www.bbc.com/russian/rolling_news/2013/04/130409_rn_st...
- thomasfl 11y agoNow that you have revealed the secret, Hacker News will be banned in russia forever.
- florian-f 11y ago> var ﷺ = 1; < undefined
- Procrastes 11y agoThis seems like a useful tool for fuzztesting your dev ops person, or if you are the dev ops person, for fuzz testing development. Fuzz for all!
- Kristine1975 11y agoPiping the result through TTS creates weird results (on OS X): echo "hello world" | mimic --me-harder 100 | say
- archimedespi 11y agoCan anybody provide an audio snippet for those of us who use Linux?
- roryokane 11y agoI don’t have an audio snippet, but I can transcribe what the voice says on different runs. It usually pronounces random letters individually, but sometimes pronounces syllables with letters missing: “L-W-R-D”, “L-L-W-R-L”, “hell-erl”, “H-L-er-D””, “eor-D”, “H-L-L-erl”, “L-L-W-L-R-D”, “H-L-W-R-L”, “hell-W-R”
- webXL 11y agoHmm... I wonder if this can be used in browser source maps.
- foolfoolz 11y agois anyone aware of the reverse of this, a homoglyph normalization library? id love to be able to take strings that visually look the same and compare them against one master list, such as for spam detection
- TazeTSchnitzel 11y agoIn some languages which allow non-ASCII but aren't Unicode-aware (PHP, for instance), you can add significant, invisible zero-width spaces to identifiers.
- tucif 11y agoSpotify used to have a security problem with this kind of characters: https://labs.spotify.com/2013/06/18/creative-usernames/ https://labs.spotify.com/2013/06/18/creative-usernames/
- reinderien 11y agoMimic author here... sorry, humanity...
- ChrisArgyle 11y agoAnd now I know what I'm doing for April 1st next year.
- cruise02 11y ago> Replace a semicolon (;) with a greek question mark (;) in your friend's C# code and watch them pull their hair out over the syntax error I'm not sure how frustrating this would be. Wouldn't most people just delete the character immediately and type a new one?
- patal 11y agoIf faced with a linter error, I don't typically delete the marked stuff, write it anew and hope fingers crossed that the error would be gone. I would try to make sense of the message, how it applies, and what the error is. At some point though, I definitely would pull my hair over a greek question mark.
- cruise02 11y agoOnly one character is going to be marked in this case, not a whole line or section of code. Deleting it and retyping it costs one second. I guess I've seen more than my fair share of encoding issues. I used to tutor at a university, so students were constantly coming in with code they'd copy/pasted out of their assignment (usually a Word doc) or from a web site.
- patal 11y agoI think that's a great argument. If someone mails the code, I hope to have the cleverness to suspect the encoding. However, I thought about a code repository or similar where this may be an issue, but most often is not. And I have seen some code where a wrong language character did not provoke a reasonable error, but some arbitrary parser error that went off in another line altogether (not necessarily C#).
- CUViper 11y agoI don't know about C# compilers, but gcc gives me two errors, "stray ‘\315’ in program" and "stray ‘\276’ in program", which I suppose are the two utf-8 bytes. Rust says, "unknown start of token: \u{37e}". Either way, you get a pretty strong clue that there's a funny character present.
- Animats 11y agoThere's a set of rules used on domain names to stop homoglyph abuse there.[1][2] Applying those rules to language identifiers would prevent this problem. It's also useful to apply those rules to login names for forum/social systems. The rules prevent mixed language identifiers, mixed left to right and right to left text, and similar annoyances. [1] https://tools.ietf.org/html/rfc5893 https://tools.ietf.org/html/rfc5893 [2] http://unicode.org/reports/tr46/ http://unicode.org/reports/tr46/
- Induane 11y agoThese dang democrats done banned Ben Carson from google man! https://www.google.com/search?q=Ben+Сarѕоn https://www.google.com/search?q=Ben+Сarѕоn
- bradbeattie 11y agoAdd the following to your ~/.vimrc to always highlight non-ascii characters: au BufWinEnter * let w:matchnonascii=matchadd('ErrorMsg', "[\x7f-\xff]", -1)
- mrzool 11y agoSome men just want to watch the world burn.
- cammsaul 11y agoThe repo's README mentions a vim plugin to highlight Unicode homoglyphs. As an Emacs user, I did a quick M-x package-list-packages, thinking I'll find at least half a dozen equivalent Emacs packages. To my dismay, there were none. So I spent the rest of my afternoon correcting this glaring deficiency. Fellow Emacs users, protect yourself from Unicode trolls and grab it here: https://github.com/camsaul/emacs-unicode-troll-stopper https://github.com/camsaul/emacs-unicode-troll-stopper
- perlancar2 11y agoMade a perl port: https://metacpan.org/pod/mimic https://metacpan.org/pod/mimic (currently 50% faster)