10 ms·
Show HN: DNS over Wikipedia
- aaronjanse 6y agoHey HN, I saw a thread a while ago (linked in README) discussing how Wikipedia does a good job keeping track of the domains of websites like Sci-Hub or The Pirate Bay. Someone mentioned checking Wikipedia to find links to these sites, so I thought this would be a fun thing to automate! To try it out, install an extension or modify your hosts file, then type in the name of a website with the TLD `.idk`. For example: scihub.idk -> sci-hub.tw Cheers!
- Polylactic_acid 6y agoIts incredible how insane this seems from the title but how practical it sounds from the readme..
- mathieubordere 6y agohehe yeah, was thinking the exact same thing
- basch 6y agoRight? Basically a modern im feeling lucky meets meta-dns.
- Vinnl 6y agoI created whereisscihub.now.sh a while ago for exactly this purpose (but limited to the subset of Sci-Hub, of course, and it used Wikidata as its data source). It has since been taken down by Now.sh. Just as a heads-up of what you could expect to see happening :)
- CapriciousCptl 6y agoCould use some sort of verification since Wiki can be gamed. 1. Look at past wiki edits combined with article popularity or other signals to arrive at something like a confidence level. 2. Offer some sort of confirmation check to the user.
- upgoat 6y agoWoah this is hecka cool!! Nice work to the authors.
- erikig 6y agoInteresting idea but: - How do you handle ambiguity? e.g what happens when sci-hub.idk and scihub.idk differ? - Aren’t you concerned by the fact that Wikipedia is open to editing by the public?
- captn3m0 6y agoMaybe use WikiData? The slower rate of updates might work in your favour to avoid vandalism.
- trishmapow2 6y agoIt also shows all the historical links. Sample: https://www.wikidata.org/wiki/Q21980377 https://www.wikidata.org/wiki/Q21980377
- hobofan 6y agoI was going to say "The Wikipedia page uses the data from Wikidata", as I thought I had seen that in the past. Turns out that it's not the case, and after picking a few samples, it looks like Wikidata is barely put in use in Wikipedia (aside from the inter-wiki links).
- skissane 6y agoA lot of Wikipedia templates now pull data from Wikidata – https://en.wikipedia.org/wiki/Category:Templates_using_data_from_Wikidata https://en.wikipedia.org/wiki/Category:Templates_using_data_... In the case of the Sci-Hub article, it actually would pull the website address from Wikidata, except that the article has been configured to override the Wikidata website address data with its own. Heaps of articles do this – https://en.wikipedia.org/wiki/Category:Official_website_different_in_Wikidata_and_Wikipedia https://en.wikipedia.org/wiki/Category:Official_website_diff... – but I understand the aim is to try to reduce the number of those cases over time.
- hobofan 6y agoYeah, via the overriding is how I came to the conclusion that it's not being used. I looked at the pages (specifically the infoboxes) and basically all values shown in the article were spelled out in the source. At the same time no Wikidata ID was present in the source (though I now realize that the template can automatically query that based on the page it's being used on). > A lot of Wikipedia templates now pull data from Wikidata It looks okay for the English wiki, but the others seem to trail behind quite a lot (though obviously number of best templates isn't a perfect metric). English: 539 German: 128 Chinese: 93 French: 52 Swedish: 37 Spanish: 19 Given that only ~2200 English Wikipedia articles have their infobox completely from Wikidata (https://en.wikipedia.org/wiki/Category:Articles_with_infoboxes_completely_from_Wikidata https://en.wikipedia.org/wiki/Category:Articles_with_infobox...), it looks like Wikidata integration still has a long long way to go.
- leoh 6y agoNice work! Sometimes I seem to be directed to a wikipedia page as opposed to a URL. For example, with `aaronsw.idk` or `google.idk`. I wonder why that's the case?
- O_H_E 6y agoI think it directs to the correct link when it is labeled `URL` in wiki. In the other cases the link is labeled `Website`.
- aaronjanse 6y agoThis was exactly the issue! I just pushed fixes for this problem.
- cooper12 6y agoI've written a userscript[0] before regarding official websites and I feel this is the hierarchy you should be using: 1. Try getting the Wikidata "official website" property 2. Then any link inside of a {{url}} template or |website= in an infobox 3. And if you really want to try to get something to resolve to, the first site wrapped in {{official website}} If you need code to reference: https://en.wikipedia.org/wiki/User:Opencooper/domainRedirect.js https://en.wikipedia.org/wiki/User:Opencooper/domainRedirect... [0]: https://en.wikipedia.org/wiki/User:Opencooper/domainRedirect https://en.wikipedia.org/wiki/User:Opencooper/domainRedirect
- yreg 6y agoIs this inconsistency intended?
- renewiltord 6y agoThis is hecka cool. What a clever concept! I like the idea of piggy-backing on top of a mechanism that is sort of kept in the right state by consensus.
- abiogenesis 6y agoNitpicking: Technically it's not DNS as it doesn't resolve names to addresses. Maybe CNAME over Wikipedia?
- usmannk 6y agoNitpicking nitpicking: "Technically" CNAME is DNS insofar as DNS is "technically" defined at all.
- datalist 6y agoIt is not even a CNAME. It is a JavaScript redirect based on the the response of an HTTP request to Wikipedia.
- stepanhruda 6y agoNo one would click on “I’m feeling lucky for top level pages over Wikipedia” though
- jneplokh 6y agoAwesome idea! It could be applied to a lot of different websites. Even ones where I'm too lazy to type out the whole URL :p Regardless, having a system where you can base it off a website could definitely be expanded beyond Wikipedia. Great work!
- oefrha 6y agoPretty cool, although legally gray content distribution sites like Libgen, TPB, KAT, etc. are often or often better thought of as a collection of mirrors where any mirror (including the main site, if there is one) could be unavailable at any given time.
- jaimex2 6y agoAlternatively just don't use Google.
- Sabinus 6y agoWhat search engines don't censor?
- jaimex2 6y agoYandex and Duckduckgo are good.
- newswasboring 6y agoAnd find the site through clairvoyance?
- gbear605 6y agoOne concern is that you can’t always trust the Wikipedia link. For example, in this edit [1] to the Equifax page, a spammer changed the link to a spam site. They’re usually fixed quickly, but it’s not guaranteed. So it’s a really neat project, but be careful about actually using it, especially for sensitive websites. [1]: https://en.wikipedia.org/w/index.php?title=Equifax&diff=945519521&diffmode=source https://en.wikipedia.org/w/index.php?title=Equifax&diff=9455...
- edjrage 6y agoTrue, seems pretty risky. Maybe the extension could take advantage of the edit history and warn the user about recent changes? Edit: Unrelated to this issue, but I have a more general idea for the kinds of inputs this extension may accept. It could be an omnibox command [0] that takes the input text, passes it through some search engine with "site:wikipedia.org", visits the first result and finally grabs the URL. So you don't have to know any part of the URL - you can just type the name of the thing. [0]: https://developer.chrome.com/extensions/omnibox https://developer.chrome.com/extensions/omnibox
- yreg 6y agoThe user should exercise caution, but in the use cases provided (a new scihub/tpb domain) that applies regardless.
- frei 6y agoPretty neat! Similarly, I often use Wikipedia to find translations for specific technical terms that aren't in bilingual dictionaries or Google Translate. If you go to a wiki page about a term, there are usually many links on the sidebar to versions in other languages, which are usually titled with the canonical term in that language.
- iakh 6y agoSelf plugging a quick page I wrote to do exactly this some time ago: http://adamhwang.github.io/wikitranslator/ http://adamhwang.github.io/wikitranslator/
- nitrogen 6y agoOut of curiosity, how well does Wiktionary fare in this regard?
- matsemann 6y agoI often do it as well. It's not perfect, but it's nice for things not directly translateable. For instance events known by different names in different languages, where translating the name of the event with google just does a literal translation.
- greenpresident 6y agoI use it primarily for cooking ingredients. The names on some unconventional grains and vegetables are easy to translate using this method and not always available in conventional dictionaries. It would also be useful for identifying cuts of meat, as US cuts and, for example, Italian cuts differ not only in name but in how they are made. Compare the images on this article for an example of what I mean: https://en.wikipedia.org/wiki/Cut_of_beef https://en.wikipedia.org/wiki/Cut_of_beef
- carlob 6y agoI own an illustrated encyclopedia of Italian food: there are 9 pages of regional cuts of beef! That's what you get in a country that unified 160 years ago...
- jakear 6y ago> If you Google "Piratebay", the first search result is a fake "thepirate-bay.org" (with a dash) but the Wikipedia article lists the right one. — shpx How interesting. Bing doesn't do this, which leads me to believe it's not a matter of legality. Is Google simply electing to self-censor results that it'd prefer it's used not to know about? Strange move, especially given the alternative Google does index is almost definitely more nefarious.
- jonchurch_ 6y agoI'm not sure how long that's been the case. The actual site at their normal domain seems to have been down for a few months, with a 522 cloudflare timeout. I'm curious if that's the case for you as well, or if it's my ISP blocking (I wouldn't expect to see the cloudflare error if my ISP was blocking but I don't know). I bring this up because if the site is unresponsive from wherever you're searching (or perhaps unresponsive for all, idk) then maybe it got de-ranked on google.
- onion2k 6y agoFor me the address on Wikipedia times out with a 522 in exactly the same way. Bing's top result of the .party address works fine. I strongly suspect this is an ISP issue, but it is interesting that Google seems to have no knowledge of the .party domain.
- el_nino 6y agoJust tested over Tor and it works: piratebayztemzmv.onion P.S. Yes, it's their new official "readable" onion site link.
- shp0ngle 6y agoNote that this is their official website, not just another fake Tor proxy https://torrentfreak.com/the-pirate-bay-moves-to-a-brand-new-onion-domain-191206/ https://torrentfreak.com/the-pirate-bay-moves-to-a-brand-new...
- 6y ago
- snek 6y agoThe extension has nothing to do with DNS, a more accurate name would be "autocorrect over wikipedia". The rust server set up with dnsmasq is a legit DNS server though.
- MatthewWilkes 6y agoIt isn't autocorrect either. It's domain name resolution.
- segfaultbuserr 6y agoThere's a risk of phishing by editing Wikipedia articles if the plugin gets popular. Perhaps it's useful to crosscheck the current URL against the 24-hour earlier and 48-hour earlier versions of the same article. Crosscheck back in time, not back in revision, since one can spam the history by making a lot of edits.
- BillinghamJ 6y agoNice idea! Maybe should involve some randomised offsets so it can't just be planned ahead of time
- pishpash 6y agoAnd what would you do if there was a difference?
- s_gourichon 6y agoReturn an error code. Also, since the DNS protocol allows ancillary informations, perhaps return additional informations in fields that would seem fit, else in comments. Edit: this is not DNS over wikipedia. As other pointed out, there is no DNS involved in the linked artifact. One option would be to show alternatives with dates and let user choose.
- cxr 6y agoI jotted down some thoughts about this very thing last year. Here's the part that argues that it could work out to be fairly robust despite this apparent weakness: > Not as trivially compromised as it sounds like it would be; could be faked with (inevitably short-lived) edits, but temporality can't be faked. If a system were rolled out tomorrow, nothing that happens after rollout [...] would alter the fact that for the last N years, Wikipedia has understood that the website for Facebook is facebook.com. Newly created, low-traffic articles and short-lived edits would fail the trust threshold. After rollout, there would be increased attention to make sure that longstanding edits getting in that misrepresent the link between domain and identity [can never reach maturity]. Would-be attackers would be discouraged to the point of not even trying. https://www.colbyrussell.com/2019/05/15/may-integration.html#wikipedia-name-system https://www.colbyrussell.com/2019/05/15/may-integration.html...
- sm4rk0 6y agoNice hack, but you can do it much easier with DuckDuckGo's "I'm Feeling Ducky", which is used by prefixing the search with a backslash: https://lmddgtfy.net/?q=%5Chacker%20news https://lmddgtfy.net/?q=%5Chacker%20news That's especially useful if DDG is default search engine in your browser. (I'm not affiliated with DDG)
- kelnos 6y agoThat just takes you to the first DDG result, no? The purpose of this seems to be to treat Wikipedia as a trusted, reliable source of truth about the canonical URL for websites (debatable, of course). The idea is that you don't trust the search engines, perhaps because you live in a country where your government has required search engines to censor results in some way, but (for some reason?) lets you go to Wikipedia.
- 29athrowaway 6y agoMany Wikipedia articles can be edited by anyone. This is not secure.
- rootsudo 6y agoThis is cool!
- HugoDaniel 6y agoDNS translates a name into an IP address. This is not DNS per-se, it is just a search plugin for the url bar. If an analogy was needed with a network service perhaps this is more like a proxy redirector than DNS. Keep in mind: with this you will still be misdirected if your DNS/hosts file is pointing the name into a different IP than it should be.
- capableweb 6y agoIndeed. Even the GitHub repositories description has this error. > Resolve DNS queries using the official link found on a topic's Wikipedia page @aaronjanse: you probably want to correct this. "Resolving DNS records" carry a specific meaning in that you have a DNS record and you "resolve" it to a value, which actually. You're kind of doing, in a way, I suppose. I was convinced when I started writing this comment that calling this "resolve dns queries" is wrong. But thinking about it, DNS resolving is not necessarily resolving a "name into a IP-address" as @HugoDaniel in the comment I'm replying to is saying (think CNAME records and all the others that don't have IP addresses). It's just taking something and making it into something else, traditionally over DNS servers. But I guess you could argue that this is resolving a name into a different name, that then gets resolved into a IP address. So it's like a overlay over DNS resolving. Meh, in the end I'm torn. Anyone else wanna give it a shot?
- nulbyte 6y agoArguably, the plugin and resolver "resolve" domains under the top-level domain idk. However, the primary service provided is not DNS (which is not offered at all via the plugin), but HTTP redirection. DNS, on the other hand, serves a variety of applications, not just HTTP clients.
- parhamn 6y agoI agree. DNS in conversation = K/V mapping pair for routing somewhere. TXT/MX/CNAME/A/WIKI etc. For the sake of this repo and what they're trying to get across this seems fair. I'm confused that I felt compelled to write this though.
- penagwin 6y ago
- jrockway 6y agoWhy does Google censor results, but not Wikipedia? It seems like you can DMCA Wikipedia just as easily as Google. Overall this is a nifty hack and I like it a lot. Wikipedia has an edit history, and a DNS changelog is something that is very interesting to have. People can change things and phish users of this service, of course, but with the edit log you can see when and potentially why. That kind of transparency is pretty scary to someone that wants to do something malicious or nefarious.
- jhasse 6y agoGoogle also sells copyrighted content, Wikipedia doesn't.
- BubRoss 6y agoWouldn't dns over github make more sense than this?
- lizardmancan 6y agoname server over everything
- nomanlaghari 6y agoplease vist below website for poetry https://bit.ly/2yErlmt https://bit.ly/2yErlmt
- LinuxBender 6y agoThis may be a little off topic, but has anyone ever considered a web standard that includes a cryptographic signed file in a standard "well known" location that would contain content such as - Domains used by the site (first party) - Domains used by the site (third party) - Methods allowed per domain. - CDN's used by the site - A records and their current IP addresses - Reporting URL for errors Then include the public keys for that payload in DNS and in the APEX of the domain? Perhaps a browser add-on could verify the content and report errors back to a standard reporting URL with some technical data that would show which ISP is potentially being tampered with? Does something like this already exist beyond DANE? Similar to HSTS maybe the browser could cache some of this info and show diffs in the report? Maybe the crypto keys learned for a domain could also be cached and warn the user if something has changed (show diff and option to report)? Maybe more complex would be a system that allows a consensus aggregation of data to be ingested by users so they may start off in a hostile network and some trusted domains populated by the browser in advance, also similar to HSTS?
- andrekorol 6y agoThat's a good use case for blockchain, in regards to the "consensus aggregation of data" that you mentioned.
- Spivak 6y agoWhy would you need a blockchain for this? This would just be a text document sitting at $domain/.well-known/$blah and verifiable by virtue of being signed with a cert that's valid for $domain.
- blattimwind 6y agoWouldn't this be an excellent use case for Wikidata? For example looking up "sci hub" on Wikidata leads to https://www.wikidata.org/wiki/Q21980377 https://www.wikidata.org/wiki/Q21980377 which has an "official website" field.
- hk__2 6y ago> Wikipedia keeps track of official URLs for popular websites This should be Wikidata. Wikipedia does that, but this is more and more moved into Wikidata. This is a good thing, because Wikidata is much easier to query, and the official website of an entity is stored at a single place, that is then reused by all articles about that entity in all languages.
- snorrah 6y agoDoes this comply with the terms of service? I know this won’t be a popular reply and that’s fine, but I just want to know whether your admittedly intriguing concept isn’t taking the piss :)