3 ms·
Excerpt worth pulling out for HN readers: Crawling incidentally I think is the biggest issue with making a new search engine these days. Websites flat out refu
by mceoin 5y ago
Excerpt worth pulling out for HN readers:
Crawling incidentally I think is the biggest issue with making a new search engine these days. Websites flat out refuse to support any crawler than Google, and cloudflare and other protection services and CDN’s flat out deny access to incumbents. It is not a level playing field. I would actually like to see some sort of communal web crawl supported by all web crawlers that allows open access to everyone. The benefits to websites would be immense as well, as they could be hit by a single crawler, rather than multiple, and bugs could be ironed out.
- lifthrasiir 5y agoYour comment has been misplaced, it should be under https://news.ycombinator.com/item?id=28665395 https://news.ycombinator.com/item?id=28665395 (I was pretty confused until that realization).
- infogulch 5y agohttps://commoncrawl.org/ https://commoncrawl.org/ is a big crawler with all data made public. Probably a good resource for new search engines.
- boyter 5y agoYes and no. It works for the deep web, but for anything that needs more frequent crawling no. They don't do it often enough. I have mentioned this elsewhere, but id love there to be a single crawl that everyone was able use and build indexes off. It would be especially good for webmasters too as you would have one bot to worry about and optimise for, reducing the load on systems, and allowing for innovation in the search space that cloud flare and other cdn's are stifling because they understandably are blocking crawlers that aren't Google or Bing. For example id love to see something added to robots.txt that allows you to say "crawl at 2am, hit has hard as you want, don't come back for a week" or some such. Having a single crawl could allow for this.
- marginalia_nu 5y agoWhile it's an obstacle, I don't think this is as big of an issue as they make it out to be, and I say this as someone who runs a DIY crawler and a janky homebrew search engine. You can get listed as a good bot with cloudflare &c if you ask nicely enough. Even if you aren't, their IP ranges are public so you can crawl them really slowly if you absolutely must have their stuff. In the end, you are never going to get completeness, that's a bad mark to aim for, and not even Google does that. That goal requires you have a model of the Internet that's as old as Vannevar Bush. The Internet isn't a bunch of static files on a souped up FTP server anymore, most pages are generated request-time, increasingly client-side. There are websites that don't just have infinite documents, they have infinite wildcard subdomains as well. There are websites that give you different results depending on what time of day it is, or where you are from. The hard problem of search engines is to find a good subset of the internet for some definition of good; and the first step toward doing that (in a timely fashion) is to rein in the crawling. Search engines are always always going to present a selection of web sites, and there are always going to be pages you can't find because they didn't make the cut; either because they were excluded, or because they excluded themselves.
- KirillPanov 5y agoHey, your project is awesome, but you seem to be the only non-big3 crawler who doesn't encounter this problem and call attention to it [1]. Maybe because you (awesomely) focus on low/no-javascript, text-heavy sites? Much of the most obnoxious antibot stuff (reCAPTCHA, CDN "AI" throttling, a lot of what cloudflare does, etc) is javascript-based. Maybe I'm totally wrong here. If so, you really should make your crawl corpus publicly available. Trickle it into one of the public clouds and charge people a few microcents per query. You'll get to retire early, knowing you've done a good deed for human technological progress. > There are websites that give you different results depending on what time of day it is, or where you are from. That has been true since at least 1993 [2], before there were search engines! > The hard problem of search engines is to find a good subset of the internet for some definition of good This really is a case of the tech industry trying to redefine "search" in a way that gives them editorial power. Librarians don't do that sort of thing. We could learn a lot from their integrity. Or we could just stop calling it "search", because that's not what it is anymore. For example, that box at the top of amazon.com is most definitely not a search box anymore; it's a sales-lead-generator box. The button should be marked "try to sell me shit which is not necessarily related to these words" instead of "search". [1] off the top of my head: archive.is/archive.today, gigablast. I'm actually having trouble remembering any still-operating whole-web crawlers besides Google, Bing, and Yandex. [2] https://en.wikipedia.org/wiki/Common_Gateway_Interface https://en.wikipedia.org/wiki/Common_Gateway_Interface
- snidane 5y agoWell, then change your user agent to 'googlebot' when crawling the web, since it now became the synonym for search crawlers?
- jeroenhd 5y agoThat trick doesn't work because people also filter out Google's IP addresses to verify that the "googlebot" crawler is actually a real Google bot.
- deleted 5y ago[deleted]
- jasonladuke0311 5y agoCan you run a crawler from GCP or do they explicitly filter for other Google AS?