10 ms·
Show HN: Fully-searchable Library Genesis on IPFS
- ramon 5y agoNice, is it open source? Can we learn from this implementation? Is there any documentation?
- sixtyfourbits 5y agohttps://libgen-crypto.ipns.dweb.link/source.tar.gz https://libgen-crypto.ipns.dweb.link/source.tar.gz It makes use of the sql.js-httpvfs library, previously discussed on HN here: https://news.ycombinator.com/item?id=27016630 https://news.ycombinator.com/item?id=27016630
- deleted 5y ago[deleted]
- ajvs 5y agoAwesome project, now all we need is the SciHub version of this.
- antegamisou 5y agoMost of SH papers are accessible from libgen too.
- sayonaraman 5y agothere is an open source project scihub-p2p (written in Go) indexing scihub articles on libgen torrents, but it also seems to work with IPFS https://www.reddit.com/r/scihub/comments/ovp5c0/make_scihub_live_on_p2p/ https://www.reddit.com/r/scihub/comments/ovp5c0/make_scihub_...
- ducktective 5y ago> decentralized Web I thought the web and internet was decentralized already? Interesting project, btw! Many thanks to the devs.
- superkuh 5y agoI agree. IPFS is more centralized in that 99% of the people that attempt to view data on the IPFS network go through the actual web proxies not IPFS. And when that happens they're easy to take down with legal or traffic attacks. Additionally, IPFS's devs have already stated on their community forums that content like sci-hub is not welcome there.
- sixtyfourbits 5y agoThe gateways are indeed centralized (though there are several different public gateways). But anyone who installs IPFS on their computer can access the content directly. Also some browsers support IPFS natively, e.g. https://brave.com/ipfs-support/ https://brave.com/ipfs-support/ From what I understand, the notion of sci-hub/libgen "not being welcome" was only about discussion on the official forums. See https://discuss.ipfs.io/t/mirror-of-sci-hub-in-ipfs/1613 https://discuss.ipfs.io/t/mirror-of-sci-hub-in-ipfs/1613 and https://news.ycombinator.com/item?id=25209246 https://news.ycombinator.com/item?id=25209246. But IPFS is a protocol just like bittorrent or HTTP, and the software is open source; it doesn't and can't enforce copyright restrictions.
- superkuh 5y ago>But IPFS is a protocol just like bittorrent or HTTP, Yes, but it's a protocol with a centralized single group doing development who can change whatever they want without the users' consent. Take a look at what is happening to the Tor ecosystem on Oct 15th this year: all tor v2 routing support is being dropped from the main client and infrastructure (for security reasons). Entire communities built on onionland and other tor v2 features, as well as all URLs/links, search engine databases, etc, will just go poof when the devs drop support. Unfortunately being a protocol isn't enough. It has to be a community protocol, not a proprietary one where everyone follows one group's code. HTTP and bittorrent are safe from these kinds of attacks. IPFS isn't (yet) and that's why their butt-covering anti-sci-hub/libgen stance is worrying.
- gzer0 5y ago
- c54 5y agoThis is great! I've been half-barely-following IPFS development for a few years now but I think this is a salient use case that I could actually see myself using. I think also with IPFS i can share files with peers pretty easily? It's nicer than uploading to a filesharing site, and easier than setting up a torrent. So, what's next @sixtyfourbits? Is there a read-only wikipedia on ipfs yet? edit: found it, but I think it's not searchable https://en.wikipedia-on-ipfs.org/wiki/ https://en.wikipedia-on-ipfs.org/wiki/
- sixtyfourbits 5y agoNext up is the scimag collection (also maintained by library genesis), which is a backup of all articles on sci-hub. The torrents are already widely replicated, but whether or not IPFS is going to be able to scale to the level required for 85M articles is still an unknown.
- lettergram 5y agoI thought the new release allows for an exabyte of data. Not sure it can handle bandwidth though
- 8K832d7tNmiQ 5y agoIt's probably just a few steps away to finally build an IPFS version of sci-hub. Sadly, I'm not a fan of the site's method of searching the libgen index by using sqlite's partial load feature [1] mainly because of the possible limited available storage issue. [1]:https://phiresky.github.io/blog/2021/hosting-sqlite-databases-on-github-pages/ https://phiresky.github.io/blog/2021/hosting-sqlite-database...
- morsch 5y agoI remember the sqlite via static host discussion, but I don't understand what you mean by "the possible limited available storage issue". Can you explain?
- dudehere 5y agoIt's a feature which actually makes fast decentralized search real-time. If no partial function is involved, it won't work before downloading the entire database. It's a feature and beauty, not something to dislike. It's just given.
- hugoroussel 5y agoThe website looks down :( I love libgen will switch to your version for sure :)
- fabianhjr 5y agoIt is available over IPFS. Do you have a client installed and running? EDIT: if you do have IPFS and the companion extension then libgen.crypto should resolve to something like `http://libgen.crypto.ipns.localhost:8080/ http://libgen.crypto.ipns.localhost:8080/` which currently works as advertised.
- dudehere 5y agoIt's not a Web-site. It is more suitable to call it antisite since it has no single location where its content is hosted. To access it, you need to use software given in the description https://libgen.fun/dweb.html https://libgen.fun/dweb.html
- madars 5y agoTech details from the Getting Started guide: > How does this work? > SQLite compiled into WebAssembly fetches pages of the database hosted on IPFS through HTTP range requests using sql.js-httpvfs layer, and then evaluates your query in your browser. The same guide, https://libgen-crypto.ipns.dweb.link/ https://libgen-crypto.ipns.dweb.link/, also explains how you can also download the page to search locally without constant internet access. sql.js-httpvs was previously discussed on HN here: Hosting SQLite databases on GitHub Pages or any static file hoster (1812 points) https://news.ycombinator.com/item?id=27016630 https://news.ycombinator.com/item?id=27016630
- easrng 5y agoThis is cool, but more centralized than it needs to be. The update check resolves libgen.crypto using @unstoppabledomains/resolution[0] with its default Ethereum provider, Infura[1]. That means that if Infura disables the default API key, goes down, or starts censoring responses, the update check will fail and users will be stuck on an older version of the site. Using the .crypto domain for updates is unnecessary, a simple IPNS[2] lookup (Not to be confused with DNSLink[3]) would've sufficed. [0]: https://www.npmjs.com/package/@unstoppabledomains/resolution https://www.npmjs.com/package/@unstoppabledomains/resolution [1]: https://github.com/unstoppabledomains/resolution/blob/HEAD/README.md#default-ethereum-providers https://github.com/unstoppabledomains/resolution/blob/HEAD/R... [2]: https://docs.ipfs.io/concepts/ipns/ https://docs.ipfs.io/concepts/ipns/ [3]: https://docs.ipfs.io/concepts/dnslink/ https://docs.ipfs.io/concepts/dnslink/
- miohtama 5y agoPOKT Network offers P2P Ethereum API nodes. It is early though, and paying is difficult (consumers do not want to pay) https://www.pokt.network/ https://www.pokt.network/
- dvcrn 5y agoIPNS is so slow that using it for anything just makes for a very unpleasant experience
- dudehere 5y agoYou can switch to literally any other resolver in the browser settings or use alternative URLs to access the antisite. Also, in future there can be other access paths to this search. All of them go through different infrastructures and may or may not work for a particular user.
- nynx 5y agoPretty neat. Too bad searches take such a long time.
- dudehere 5y agoTry all the various URLs in different browsers given in the intro, not just one combination of URL and browser. Some combinations are much faster. Use them then. The fastest you find is usually just within a few seconds. Also, consider pinning as described on the libgen.crypto antisite itself. If you pin it, the search is going to be instantaneous.
- isoprophlex 5y agoAnd it's fast as hell too! This is fantastic yo, mad props. Thanks for liberating our collective knowledge, ipfs style. Keep it up!
- severine 5y agoYou could add a link to IPFS Lite ("IPFS Lite node with modern UI to support standard use cases of IPFS ") on F-Droid. Seems actively updated, can anone vouch for its quality? Source: https://gitlab.com/remmer.wilts/ipfs-lite https://gitlab.com/remmer.wilts/ipfs-lite F-Droid: https://f-droid.org/en/packages/threads.server/ https://f-droid.org/en/packages/threads.server/
- namibj 5y agoWorks, but last I tried, not even a local gateway, so you're limited to the built-in web-view.
- cmeacham98 5y ago> Put http:// http:// as the protocol prefix instead of various https:// https://, ipfs://, or ipns://. The most universal format is http:// http://, since transport-level security (TLS) is not yet fully adopted in dWeb systems (neither it is essential for our applications where a domain name resides on a blockchain and SSL traffic encryption is employed further on) I understand that this isn't really their fault because they'd have to get a CA to issue a cert for the non-standard .crypto TLD, but has to be untrue, assuming I understand correctly that the HTTP version is just hosting a JS IPFS client? And therefore the non-IPFS link is suceptible to MITM attacks if I understand correctly.
- dudehere 5y agoNo difference ICANN or blockchain TLD for making an SSL certificate. Where can this MITM stick his foot in? The blockchain domain record is encrypted and contains the publicly visible target IP-address. Even if accessed via http, it can be instantaneously checked. And upon visiting that IP address, the site forces https encryption. Nobody can know from mere traffic analysis what exactly you are doing within the site, since it's encrypted by SSL. IPFS encryption isn't necessary either, since content is addressed by its hash which automatically guarantees its authenticity. It doesn't matter if you use https or http. It's only imperfect in the transition during DNS resolution. A 3rd party can know what site you go to, but not more than that. Suspecting a public generic blockchain-domain resolving node set up for freedom in defacing would be a bit too much. Usually those are OpenNIC servers or huge DNS providers, unless you specify a custom DNS resolver in your settings. For a book site it's a pretty decently protected transition, actually maximal available for a non-expiring (unmanned) service which it is. Legit certificates expire or cost money. The system behind libgen.crypto is fully unmanned, i.e. eternal, except the files themselves which need hosting.
- cmeacham98 5y ago> And upon visiting that IP address, the site forces https encryption. This isn't true if there is a MITM attacker. When you visit an HTTP website, the website doesn't get to redirect you to HTTPS or anything else, because the game is already lost. When you visit a website over HTTP the attacker goes first. The legit response never made it to the client, because it was replaced in transit with a redirect to the attacker's scam/phishing/malware website.
- mijailt 5y agoIt's not working for me on Firefox 92.0. It hangs on "Initializing...", with a single error on the console which says Loading Worker from “http://libgen.crypto.ipns.localhost:8080/dist/257fb50677e11621f8a0.js” http://libgen.crypto.ipns.localhost:8080/dist/257fb50677e116... was blocked because of a disallowed MIME type (“text/plain”)
- dudehere 5y agoIn any browser you may set a custom DNS in privacy and security settings to https://resolver.unstoppable.io/dns-query https://resolver.unstoppable.io/dns-query This will resolve https://libgen.crypto/ https://libgen.crypto/ via DNS-over-https. Follow this manual, if it's a bit new to you: https://thewindowsclub.com/access-blockchain-domains-oi-browsers https://thewindowsclub.com/access-blockchain-domains-oi-brow...
- mbStavola 5y agoImmutability is a blessing and a curse for IPFS. It's cool for preventing things like censorship. Something like SciHub would really benefit from it. However, for "real world" use cases, many people want to be able to remove or modify what they've uploaded. With IPFS, as far as I'm aware, doing either doesn't really change the underlying data but just creates a new object in IPFS instead which you'd point to via IPNS. Anyone who still wanted to view the old content still could, provided they had the right content id. God forbid you accidentally upload a "personal" photo, your only hope is that someone never comes across the content id of that image. There is no way to undo it!
- eric__cartman 5y agoFrom my understanding if you accidentally upload a personal file, as long as no one downloaded it in the time you took to realize your mistake taking down the only node that has the file (your computer in this case) should effectively "erase it" in the sense that unless the node comes back up, even if someone has the id of that file they are SOL.
- unknownOrigin 5y agoOk, I have to ask, what is an actual difference between torrents and ipfs? I don't care for technical details, I mean the business logic, so to speak. - Both use DHTs to search for sources of fingerprinted content. - Both use nodes (seeds in BT terminology) that actuallu store the content. - Both don't have an "archive" system, and so if at least one node doesn't have the file, it may as well not exist at all. - Both can have content censored by going after the node operators. Am I getting any of these wrong?
- easrng 5y agoBitTorrent is easier to understand (IMO) and can sometimes be faster at lookups than IPFS. IPFS dedups content, whereas if you have an identical file in 2 torrents but only one is seeded, you can't download it from the other one. As far as I know those are the only differences. Edit: Actually BitTorrent v2 dedups files so it seems like IPFS and BitTorrent are now functionally identical.
- dr_dshiv 5y agohttps://libgen-crypto.ipns.dweb.link/ https://libgen-crypto.ipns.dweb.link/ Wow, this is way more user friendly than normal! Nice work.
- betwixthewires 5y agoBravo. This is big stuff. There are some good critiques and ideas in this thread I won't go over since it's been done, I have been putting off building something akin to this (not the same thing but somewhat similar) hoping someone more motivated would do it, and I'm very excited to see it happen.
- DantesKite 5y agoWhat does this mean? That if any legal entity ever tried taking Library Genesis down, they'd be unable to?
- dudehere 5y agoIt's a composite design, not everything placed in one location. Even from this primitive explanation it is clear that taking it down is a nontrivial task. And with time this system will only grow stronger without applying more effort, just because the number of supporting participants will grow when they pin the antisite. This will make the blocking formidable.
- DantesKite 5y agoThank you for taking the time to explain.
- dudehere 5y agoYou are very welcome.
- unraveller 5y agoIt's a fantastic day for digital monks! I just wish the url took a searchable ?query=aristotle on load so you could add it to your browser's search engine list
- dudehere 5y agoYeah, there has been a discussion on that matter. It would take the same time as the entire bootstrapping and the search query processing to open such an URL. This is doable but may exhibit inadequate performance. I support you, though, that independent of the search time it is valuable to have a linking standard like that.
- stavros 5y agoWouldn't #query=Aristotle work and be just as fast?
- dudehere 5y agoThe databases still needs to bootstrap first, then the code can search on it. Initialization phase should prepend showing a record.
- stavros 5y agoSure, that's why the anchor would be faster, because you can use the same initialization code and still accept the query in the URL.
- dudehere 5y agoAnchor does not solve the bootstrapping issue. You may click an URL with an anchor or without, but the new page in the browser will need to load the database fragment and search over it before showing you the retrieved record. There's no faster way known yet in this implementation.
- snvzz 5y agoPretty nice. I have used it without issue to download a few public domain books. Relative to the usual website, it is a bit crude. A lot of metadata is missing, thus it is hard to decide which book to download. Particularly annoying are the lack of ISBNs and such, and the inability to click the author and see other books by them. The worst point is that files come with no proper filename (just a hash!), thus inviting everyone to rename them in a non-standard manner, rather than offer a filename people won't have to rename.
- dudehere 5y agoIt was necessary to crop the meta info that much to host for free. It's a feature, not a real limit. The clickability problem is explained down this thread, it's the same as making an URL addressing a book. Not that nice. It's a technological peculiarity. The naming may have a solution a bit later. It wasn't clear initially how to approach it.
- snvzz 5y ago>It was necessary to crop the meta info that much to host for free. It's a feature, not a real limit. Maybe have a step of indirection (an extra page) showing the metadata, with the ability to download directly still in the index should the metadata "page" not be fetchable. >The clickability problem is explained down this thread, it's the same as making an URL addressing a book. Not that nice. It's a technological peculiarity. It's good as long as libgen is aware. To be clear, it's good to be up at all, and it doesn't need to be perfect on the first iteration. >The naming may have a solution a bit later. It wasn't clear initially how to approach it. Showing a filename somewhere to allow downloaders to manually rename the file to a standard form would be a step in the right direction.
- dudehere 5y agoYes, I think issuing multiple free opportunities might give a more complete solution without compromising the unmanned service. My browsers actually do offer to rename the files. It might be that yours is set not to prompt.
- bscphil 5y agoCouple of thoughts: 1. > Show HN Did you @sixtyfourbits make this? Any stories about how you came to be involved in the project? IPFS seems like a pretty ideal way to handle sharing documents like this, I'm surprised LibGen hasn't used it before (previously, you would get redirected to one of many constantly dying domains that may have ads and frequently 404 on the actual book you're looking for). 2. Also, this interface frequently doesn't work in Firefox for me. It hangs while trying to load the file. Fortunately, you can check the browser dev tools and find the actual IPFS gateway link (which uses ipfs.io), and go to that link directly. My experience is that the direct link not only works far more frequently, it's actually faster as well. So this raises an obvious question: rather than load the file in a fancy interface, why doesn't the link just take you directly to the IPFS gateway? 3. Is there any concern that systematically using a legitimate service like IPFS to share illegal material will create a situation similar to that of Bittorrent, which is similarly often presumed illegal until proven otherwise? That seems like a shame. I suspect the only reason why rights holders have not cracked down on IPFS is that it's not yet big enough to be on their radar.
- dudehere 5y ago2. It's random for a random user. I'd suggest to try every suggested natively supporting browser with the list of given URLs to find which combination works the best for you. 3. An arbitrary IPFS gateway can be set up our rented, it's not a taboo. They are usually $10/Mon.
- bscphil 5y ago2. I don't understand what you mean? Are you saying the gateway is random? I tried several different browsers and got ipfs.io every time. Are you saying it's random whether it works or not? If so, that seems ... bad. 3. It's not about whether it's difficult to act as an IPFS node, it's about whether doing so will (in the future) bring you under legal scrutiny the same way running a node serving copyrighted content on the Bittorrent network will do now. DMCA against the major gateways will probably work to make files difficult to access, and IPFS necessarily reveals the IP address of the node you connect to, if you don't reveal a gateway. Similar techniques are used to get the IP addresses of Bittorrent users, and send them demands for financial compensation or sue them in court for distributing copyrighted material. If the same becomes common for IPFS, it would not be unlikely to see college networks come under pressure to ban access to IPFS, and this would limit access to LibGen's database in a significant way.
- CuriouslyC 5y agoIPFS is pretty cool. We just need to come up with a really good solution to distributed search, and get some big content creators to sign on, and it'll take off.
- therealcamino 5y agoHonest question: what do you envision happening if Library Genesis becomes widely known and people looking for e-books make that their first stop instead of Amazon or another store or the public library? In this thread I see discussion about the technical aspects, scientific research aspects, and censorship aspects, but nothing about what the economic effects will be if you're successful.
- dudehere 5y agoIt's a very good question. LG should fit in the abyss for the poor, but let the business evolve. A rebalancing from the legal entities will be required, but then a global balance can be established. Even widely known, it should take its place, and businesses their place. The two sides aren't mutually exclusive, but rather complementary. Business cannot offer what LG does, and in this frame it is pointless to battle LG.