5 ms·
Quite a neat way to crawl websites using a browser extension. That by itself is a form of donation to the search engine. Maybe in the future you can have dedica
by tyropita 4y ago
Quite a neat way to crawl websites using a browser extension. That by itself is a form of donation to the search engine. Maybe in the future you can have dedicated software for self-hosted clients that users can run to crawl and index websites for mwmbl? Kinda like folding@home.
How are the batches of URLs to be crawled generated/discovered and posted at your API?
How do you deal with duplicate crawls?
- thatwasunusual 4y agoI have also thought that distributed crawling with the help of browser extensions, and/or clients like folding@home, could be a good idea. But how to deal with "spam injections"?
- 6510 4y agoYaCy just crawls the results [again] locally before showing them to you.
- thatwasunusual 4y agoMaybe I misunderstand, but doesn't that mean you lose the benefit of having distributed crawlers if everything has to be crawled (again) locally somewhere?
- nix23 4y agoYaCy can do distributed crawling and exchange the Indexes (in Peer to Peer mode). I have some node's who just receives and send indexes without crawling (much less storage intensive).
- tyropita 4y agoAfter a certain scale I think you can let clients do double-work and let the most common crawl data, among different clients, win. And since you control what URLs need to be crawled, you protect yourself against rogue clients sending arbitrary URLs. There certainly are a lot of elegant ways to reduce spam for this particular problem imo.
- thatwasunusual 4y ago> And since you control what URLs need to be crawled, you protect yourself against rogue clients sending arbitrary URLs. I'm not worried about the URLs, but the content of the URLs sent back. Say the server tells a client to crawl a CNN article. The "hacked" client sends a fake CNN article back.
- happymellon 4y agoGet 3 people to scrape it and see if there are significant differences. Some might, because of A/B testing or news updating, but even updating news will get a positive similar page and those that don't should probably fall into an exceptions category until it can be determined what can be done about it. Maybe a flag in the URL to give you a static page or just accept that it changes often enough that even faked pages won't last long?
- zaarn 4y agoThen I'll just add 3 million bots to the network (or just enough to have about 50%) and I can guarantee to win the A/B test against an honest client most of the time.
- closedloop129 4y agoThen OP has to do things that don't scale: Review some pages and identify a subset that can be trusted. Then OP can compare their downloads to new accounts and mark the bots.
- zaarn 4y agoThen the botnet will just be honest for like a year before it abuses the network. Even better because now honest new clients can be kicked as they disagree with the bot majority. So now the network bleeds users.
- 4y ago
- daoudc 4y agoRoughly this https://github.com/mwmbl/mwmbl/wiki/Crawler-Phase-2-Design https://github.com/mwmbl/mwmbl/wiki/Crawler-Phase-2-Design It changed a bit in the implementation.