4 ms·
If Google can scrape my site, am I allowed to scrape Google results? Could I create a Google clone by scraping? If I scraped the most common search results fro
by kyle_morris_ 6y ago
If Google can scrape my site, am I allowed to scrape Google results? Could I create a Google clone by scraping?
If I scraped the most common search results from Google, front page only, and removed all the ads what would Google's argument against that be?
On one hand, so many sites make finding information difficult, on the other it feels pretty scuzzy that Google prevents searchers from clicking through to the site that put the work into generating content.
- wombatmobile 6y agoIt is scuzzy that Google steps on other sites' air hoses. Your idea to scrape Google's search results is pithy, ironic counter-innovation at its dastardly best. All you need to pull this off is funding for a top legal team, and deep reserves of emotional energy. Go for it!
- deleted 6y ago[deleted]
- anonymousab 6y agoI would assume Google respects robots.txt, so you should be fine non-abusively scraping their site insofar as you respect their robots.txt
- tempestn 6y agoYeah, the tricky thing for those scraped by google is that given google's search monopoly, the sites can't block their scraping entirely, since they need to be shown in search results.
- 1vuio0pswjnm7 6y agoGoogle's robots.txt does not tell the full story. For example, if you include a User-Agent header and put certain strings in it, e.g., "curl/7.47", you will be blocked. echo -e 'GET /search?q=robots.txt HTTP/1.1\r\nHost: www.google.com\r\nUser-Agent: curl/7.47\r\nConnection: close\r\n\r\n' |socat -,ignoreeof ssl:www.google.com,verify=0 The problem with the robots.txt "standard", e.g., ones like Google's with no "crawl-delay" directives, is that it does not define what is a "robot". The query above is obviously not a "robot", but Google, with all it resources, still treats as such. Google probably does more (abusive) scraping than any other entity. Web scraping is in their DNA. It is in their web pages, too. curl https://www.google.com/search/static/gs/animal/m05py0.html|grep scrape
- volume 6y agoIf you succeed then make something that is a proxy for gmail, and then for sheets and docs and chat! google-nextgen.com is not taken... yet.
- gscott 6y agoGoogle was pretty unhappy with Bing for doing just that. https://www.wired.com/2011/02/bing-copies-google/ https://www.wired.com/2011/02/bing-copies-google/
- TAForObvReasons 6y agoFrom the article, Genius lost this case because: > Genius isn’t the copyright holder for these lyrics, it just licenses them itself. Both Google and Genius are licensing the lyrics. Ironically, Genius ended up having to settle a case years ago because they were using lyrics without the appropriate licensing [1]. [1] https://www.nytimes.com/2014/05/07/business/media/rap-genius-website-agrees-to-license-with-music-publishers.html https://www.nytimes.com/2014/05/07/business/media/rap-genius...
- bynormous 6y agoI think the legal argument services like serpapi make is that as long as you don't create a google account and/or accept google's terms then you are free to scrape and clone what is publicly accessible (at least in the US). I have no idea though.
- 1vuio0pswjnm7 6y ago"If Google can scrape my site, am I allowed to scrape Google results." You are alllowed. Google would not likely try to sue you. They will try to block you however. Bing was created by copying Google results. Google did not sue Microsoft, but they did try to expose the copying.
- bransonf 6y ago> Bing was created by scraping Google results. Do you have a source? Just curious about the back story.
- asdfaoeu 6y agoNot GP but from memory this was the same incident: https://www.wired.com/2011/02/bing-copies-google/ https://www.wired.com/2011/02/bing-copies-google/
- deleted 6y ago[deleted]
- mcny 6y ago2011 https://searchengineland.com/google-bing-is-cheating-copying-our-search-results-62914/ https://searchengineland.com/google-bing-is-cheating-copying... https://www.wired.com/2011/02/bing-copies-google https://www.wired.com/2011/02/bing-copies-google
- mthoms 6y agoApparently, the copying allegation is false. There's more to it: https://news.ycombinator.com/item?id=24112418 https://news.ycombinator.com/item?id=24112418 Not excusing MS, but it seems they were not wholesale copying Google results.
- 1vuio0pswjnm7 6y ago"Scraping" was the wrong word. I was remembering incorrectly. Microsoft did not have to scrape. They were apparently using search data captured through Internet Explorer. There was a time before Google had control of the browser. This illustrates how companies with browsers can gather data about users' web activity and how far they can go. Today, even Firefox is gathering data about users' activity with "telemetry", and users typing things into the browser's search box are by default sending this data to Google.
- oefrha 6y agoFrom https://www.google.com/robots.txt https://www.google.com/robots.txt: User-agent: * Disallow: /search Allow: /search/about Allow: /search/static Allow: /search/howsearchworks