12 ms·
Watch out: ɢoogle.com isn’t the same as Google.com
- ergot 10y agoFor me it just redirects to http://money.get.away.get.a.good.job.with.jack.ilovevitaly.com The actual domain is http://xn--oogle-wmc.com/ http://xn--oogle-wmc.com/ Which is an Internationalized domain name[1] in punycode transcription [1] https://en.wikipedia.org/wiki/Internationalized_domain_name https://en.wikipedia.org/wiki/Internationalized_domain_name The G in question here is https://en.wiktionary.org/wiki/%C9%A2 https://en.wiktionary.org/wiki/%C9%A2 OR http://charcod.es/#%C9%A2/610 http://charcod.es/#%C9%A2/610
- underyx 10y ago>ilovevitaly.com This Vitaly guy… I got tons of referral header spam (that shows up in e.g. Google Analytics) for all sorts of social media buttons and EU cookie law scare tactic sites. And then there was Vitaly who just spammed me with ilovevitaly.com, which if I recall correctly actually was a site about himself at the time.
- ergot 10y agoWow what an odd site
- cdubzzz 10y agoInteresting, this domain now redirects to: http://money.get.away.get.a.good.job.with.more.pay.and.you.are.okay.money.it.is.a.gas.grab.that.cash.with.both.hands.and.make.a.stash.new.car.caviar.four.star.daydream.think.i.ll.buy.me.a.football.team.money.get.back.i.am.alright.jack.ilovevitaly.com/#.keep.off.my.stack.money.it.is.a.hit.do.not.give.me.that.do.goody.good.bullshit.i.am.in.the.hi.fidelity.first.class.travelling.set.and.i.think.i.need.a.lear.jet.money.it.is.a.secret.%C9%A2oogle.com/#.share.it.fairly.but.dont.take.a.slice.of.my.pie.money.so.they.say.is.the.root.of.all.evil.today.but.if.you.ask.for.a.rise.it%27s.no.surprise.that.they.are.giving.none.and.secret.%C9%A2oogle.com
- Entangled 10y agoWeb browsers should have an option to show non-ascii chars in urls in red.
- blacktulip 10y agoAnd on by default
- nailer 10y agoThey should already only show punycode for characters inside your locale. 'ɢ' is obviously an exception since (I imagine) it's considered to be in your locale, but maybe it shouldn't be.
- Freak_NL 10y agoThat would be very confusing for multilingual users. Just because my OS is configured to use a certain locale, doesn't mean I don't read text in scripts not considered part of it.
- nailer 10y agoYour OS (and browser) support multiple languages, so if you speak a language they should in the list.
- Freak_NL 10y agoThey are of course, but if you use that list instead of a single locale, you end up with a solution that only highlights 'strange' characters when they are not part of your language/locale set. So for someone who speaks only Latin character based languages you could highlight all Cyrillic characters, but for someone who speaks Russian you still have the original problem (it's not as if you can just highlight all Latin characters in their case!).
- mcv 10y agoThis would be a great solution. Allowing unicode characters in domain names is just inviting trouble. I understand that people with non-Latin scripts want domain names in their own language and alphabet, but there are way too many unicode characters that will confuse people about legitimate-looking domain names. Showing non-ascii in red would be an easy solution for everybody.
- donquichotte 10y agoSome time ago I registered http://www.goolge.io/ http://www.goolge.io/. Still haven't done anything with it, I guess at some point I'll just redirect it to duckduckgo. [EDIT: now it's redirected to duckduckgo.] This can of course be used in a malicious way. I thought about rebuilding the homepage of the bank Credit Suisse on www.credit-siusse.ch, but that's probably illegal.
- 75j 10y agoI registered http://www.4ppl3.com http://www.4ppl3.com a while back. No potential for abuse really, but I just thought it was fun to have a l33t-speak version of the domain name of one of the world's most litigious companies. That said, I haven't done anything with it, and I'm not a domain squatter, so if anyone wants it I can hook you up!
- drzaiusapelord 10y agol33t-speak is so far off my radar that I was wondering if there was some hot startup called '4 people 3' or somesuch. I doubt anyone at Apple remotely cares about l33t-speak from a branding perspective.
- erelde 10y agoTo be fair, I don't think I know anyone who "cares" about l33t speak. H4xx0r j0k3s is all it is. I never seen anyone go out of their way to defend or use leet speak all day long.
- Symbiote 10y agoIn some countries, it's the best that can be done for vanity car registration plates. e.g. a site selling them in the UK is promoting "JO66 ERX", which is probably supposed to be read as "Jogger X". Current bid £750, for some reason.
- pbhjpbhj 10y agoThat looks a lot like the sort of trademark use that authorities have deemed infringing, I'd expect the registrar to "recover" that unless you've got a clear explanation (like Goolge is your name, even then ... (remember Nissan, Mike Rowe, etc.)).
- chaz6 10y agoI thought there were supposed to be registry rules preventing similar looking names to be registered as an idna. I guess not.
- shshhdhs 10y agoI believe they aren't preventative measures, but responsive. So if Google contacts ICANN, then they may do something about it
- darkr 10y agoSome registries do this automatically. Some don't.
- talideon 10y agoYes and no. One of the problems is that Verisign's handling of IDNs wasn't exactly the best conceived, which left them with silly IDN codepoint tables like this: https://www.iana.org/domains/idn-tables/tables/com_latn_1.2.txt https://www.iana.org/domains/idn-tables/tables/com_latn_1.2....
- underyx 10y agoIt was a pretty nice surprise that when sending this URL in Slack it was automatically converted to `xn--oogle-wmc.com`.
- seagreen 10y agoThe fact that we need application-specific security measures against this just emphasizes the problem. There are a lot of applications.
- Fiahil 10y agoSlack is not doing anything. It's Google chrome filling up your clipboard with the "extended" version of the url.
- underyx 10y agoBut when I paste it in the Slack message box it shows the ɢoogle.com version.
- pvdebbe 10y agoI haven't used slack, but I think both are doing the best practices around there: Chrome copies the punycoded URL to clipboard, Slack will decode pasted punycode-URLs into a nicer presentation.
- vbezhenar 10y agoWhat about Googlé.com and infinite number of other variations?
- StavrosK 10y agoWhy is everyone thinking so small? What about https://www.goоgle.com https://www.goоgle.com? Or how about the word "gullible" isn't in the dictionary? http://www.dictionary.com/browse/gulliblе http://www.dictionary.com/browse/gulliblе
- koliber 10y agoWould it be possible to register a .xn--cm-fmc TLD and have a .cоm registry all of your own?
- tlrobinson 10y agoNot sure why you're getting downvited, people seem to have missed your clever use of the Cyrillic "o".
- vbezhenar 10y agoI think, it's impossible to register this domain.
- cesis 10y agoWhy Google analytics isn't filtering out this referral spam?
- akerro 10y agoIt's literally not their job to filter referrals... they do the opposite, they collect referrals.
- TazeTSchnitzel 10y agohttps://en.wikipedia.org/wiki/IDN_homograph_attack https://en.wikipedia.org/wiki/IDN_homograph_attack
- talideon 10y agoMost registries did a better job on constructing their IDN tables than Verisign did. :-(
- alessioalex 10y agoThis just redirects me to http://xn--oogle-wmc.com/ http://xn--oogle-wmc.com/ so I know it's not the real google (using Chrome).
- joncrocks 10y agoI believe now that browsers have support for non-ascii URLs, each of them have schemes for anti-phishing. See https://www.w3.org/International/articles/idn-and-iri/ https://www.w3.org/International/articles/idn-and-iri/ and https://wiki.mozilla.org/IDN_Display_Algorithm https://wiki.mozilla.org/IDN_Display_Algorithm plus http://www.chromium.org/developers/design-documents/idn-in-google-chrome http://www.chromium.org/developers/design-documents/idn-in-g...
- 77pt77 10y agoBrowsers have supported this for almost a decade.
- reacweb 10y agoMaybe browser should have a security option to whitelist characters in URL. When a URL uses another character, there would be popups with explanations and choices.
- orbitur 10y agoThis is something that's been bugging me for years. Why are there multiple representations of alphabet characters in Unicode? It seems reasonable to include accent marks, but what's the benefit in having a Cyrillic 'o' alongside a standard 'o' or the 2 or 3 other ASCII-lookalike sets of characters?
- leeoniya 10y agothe font metrics and hinting/kerning are likely language or dialect-specific
- kalleboo 10y agoOne goal of Unicode has been lossless round-tripping between legacy encodings (to encourage adoption). If such an encoding contains both Latin and Cyrillic, they must be separate to enable that.
- kps 10y agoCompatibility with ISO8859. For example, for Cyrillic, the first 128 characters U+40xx match ISO8859-5.
- jstimpfle 10y agoThere will never be agreement what's the set of distinct characters (also, what characters should be included, bitcoin logo, facebook logo?)). I see Unicode as a necessary evil. Due to its complexity most applications should treat Unicode text as black boxes. I never rely on Unicode for computation. When receiving Unicode I always make sure it's in the ASCII range. It could be argued that there should never have been Unicode domain names but I guess Western people are very lucky that ASCII includes most of their characters...
- user5994461 10y ago> When receiving Unicode I always make sure it's in the ASCII range. [...] Western people are very lucky that ASCII includes most of their characters... Please don't spread the myth of Western languages being encodable in ASCII, and don't pretend to support Unicode (or anything else than English) if you filter everything to ASCII. The _only_ Western language that is encodable in ASCII is English. Corollary: English is the only language that can be encoded in ASCII. The other western languages have endless issues with text being encoded/stripped down to ASCII. e.g. French, Spanish, Portuguese, German...
- Kenji 10y agoUnicode URLs are the devil. Too many indistinguishable characters. URLs should stay full ASCII imho. And I say that as someone whose language requires non-ASCII symbols. Or, in Bruce Schneier's words: "Unicode is just too complex to ever be secure."
- rurban 10y agoBut think about the poor underrepresented folks using foreign character sets? You really need to support this 'sub café {} café()' => Undefined subroutine café in your friendly and social programming language, otherwise you will be accused of discrimination. Of course the two é are not normalized. Which unicode-friendly language does really check for mixed script confusables? Java only is my guess. Even perl6 falls into this trap. http://unicode.org/reports/tr39/#Mixed_Script_Confusables http://unicode.org/reports/tr39/#Mixed_Script_Confusables
- palunon 10y agoWhen it is just accents, it's ok. But when your users have a language that uses à radically different alphabet, sometimes they can't even read Latin transliterations.
- rurban 10y agoagree. but then you need to declare your exoting encoding somehow, such as in perl via use encoding 'greek'; and then the parser does not need to guess about mixed scripts encodings on every identifier. there's only latin and greek valid, everything else invalid. how many languages even check for mixed script confusables? for dynamic languages this check is much too expensive, but they are leading the "good cause", allowing everything, and checking nothing.
- rurban 10y agoWhat about google.com which is really <U+202E>goog<U+202C>le.com :) TR36 bidi spoofs are usually worse than TR39 confusables. Move over with your cursor over it. http://www.unicode.org/reports/tr36/#Bidirectional_Text_Spoofing http://www.unicode.org/reports/tr36/#Bidirectional_Text_Spoo... That's why browsers or dns tools use libidn, just programming languages not.
- SamWhited 10y agoThere has been talk at the IETF of redefining IDNA2008 (the current way you prevent issues like this) in terms of the PRECIS framework (RFC 7564). This wouldn't exactly "solve" the problem, but it would mean that IDNA could be more agile with respect to Unicode versions and would make it easier to react to new problems, new confusable characters, etc. as they happen.
- Programmatic 10y agoI'm not sure how feasible this is, but wouldn't it make sense for .com/.net/etc to be latin alphabet only and allow other domains to be localized with unicode? I wouldn't really have a problem with 新浪首页.cn, and I doubt I would confuse ɢoogle.ru or whatever with google.com
- barkingcat 10y agoThat defeats the purpose of an internationalized dns system. The whole point of getting unicode into domain names is so we can have 新浪首页.com so that it's no longer a latin alphabet centric system.
- Programmatic 10y agoDoesn't that yield a whole class of problems though that we're trying to solve with obtuse solutions such as "let's make that character set in red so people don't get phished"? How is that any more international and/or easy to use? It seems that putting the allowed character set into the tld would be a pretty user-friendly way of doing that. Edit: As an added bonus, tlds are centrally managed, and are already western/latin encoded. So why not customize it with a localized abbreviation for the language or tld type?
- hyperhopper 10y agoOne is a matter of international standardization of a protocol. Another is a matter of client side security for a certain type of user.
- hannele 10y agoAhh, the old classic, PayPaI: https://en.wikipedia.org/wiki/PayPaI https://en.wikipedia.org/wiki/PayPaI (uppercase 'i')
- hannele 10y agoI'm curious, why is it allowed to register domain names with mixed character sets? I am behind allowing Unicode characters in domain names for the obvious reasons, but are there compelling use cases for allowing them to be mixed?
- klodolph 10y agoTechnically, Unicode is only one character set. If you want to disallow mixing, you have to disallow it on some other basis, like script. There are many edge cases to consider, though, and many legitimate reasons to mix scripts.
- Roboprog 10y agoCool! I want a cool non-alpha unicode domain. I guess "square-root" is already taken, but there must be some cool domains left (even though nobody can actually type them in). Actually, some of these would probably be nice aliases for some math / science oriented sites. E.g. - .com
- Roboprog 10y agoMeh. Markup ate my "radioactive pie" (9762 dec / 2622 hex) symbol :-(
- a3n 10y agoThis is strange to me. This is clearly meant, in unicode, to be 'G' that we all know and love. It has uselessly expanded "the alphabet" (to be western-centric) in a confusable way. Unicode maybe should have been three dimensional, with "concept of G" in the 2D space, and "ways of representing G" behind G, along the third axis. All ways of representing G, whether little capital, capital, lower case, would or at least could equate to conceptual G in the 2D space.
- drewmate 10y agoThat's a really interesting proposal, but I'm afraid it would be difficult to implement in practice. If this third dimension were actually encoded into the number that represents each character, you'd end up with a lot of wasted bits (since most characters probably wouldn't even need the 3rd dimension, or at least as much of it as the heaviest users.) Another option would be to supplement the metadata that already accompanies Unicode characters (which block it is in, the name of the character/block, etc...) This could be done in practice now, but the information would almost certainly just be ignored if it needed to be looked up in a supplemental table. Furthermore, it's difficult to agree on just about anything in Unicode, and classifying all the characters based on concept seems like a Herculean task for a slow-moving body. Any ideas for how to accomplish this in practice?
- a3n 10y agoI'll get to that as soon as I make email secure by design.
- jahewson 10y agoThis already exists in Unicode, it's called "Variation Selectors" and they have their own block and are used to select emoji skin tones amongst other things. But it would be wrong to use them in this case because an IPA G and the letter G are semantically different things and should not be unified into a single character just because they look similar.
- stevenbedrick 10y agoIt actually does do something along those lines, with the "canonical" and "compatible" equivalence rules: https://en.wikipedia.org/wiki/Unicode_equivalence https://en.wikipedia.org/wiki/Unicode_equivalence As mentioned by others on this thread, the real issue is not with Unicode per se, but rather with the ways that web browsers handle it (or fail to handle it, as the case may be).
- cjrd 10y agoProud owner of http://gïthub.com http://gïthub.com checking in...
- yamaneko 10y agoAwesome site, by the way. I'm just checking out your tutorial on LDA.
- y4mi 10y agothe visiblend screenshot on your projects page is dead because of an unresolveable dns href. the screenshot on your kmap repo[1] was dead as well, until i actually opened it. i'm guessing the jpg isnt generated until somebody clicks on it. enough cyberstalking for me this evening :p [1] https://github.com/cjrd/kmap https://github.com/cjrd/kmap
- transfire 10y agoOh, you mean Unicode Sucks(TM)? Yes. Yes it does.
- jahewson 10y agoBrowsers already blacklist many visually similar characters, it seems that the IPA characters need to be added to that list.
- mikebay 10y agoWhy all news are so Google related lately ? Is this something to do with Trump and draining the government IT swamp ? I hope google will not have part in government at all. They have broken long time peoples privacy. Google ? - No thanks!