3 ms·
Interestingly: - Bing: Goodreads is third result - DDG: Paywalled NZ Herald article appears to be the only result. Of course what you've really uncovered is
by BeefWellington 2y ago
Interestingly:
- Bing: Goodreads is third result
- DDG: Paywalled NZ Herald article appears to be the only result.
Of course what you've really uncovered is that Google search is the only thing that respects robots.txt: https://www.goodreads.com/robots.txt https://www.goodreads.com/robots.txt
- deleted 2y ago[deleted]
- Abishek_Muthian 2y agoNo, you've uncovered that. I gave a thought on whether this could be some copyright issue with Amazon but didn't bother to check the robots.txt file, nicely done!
- atombender 2y agoThe robots.txt does not prevent Google from scraping the quote page, which is the one that matches for me: https://www.goodreads.com/quotes/11541816-few-hundred-years-of-western-society-that-we-have-lost https://www.goodreads.com/quotes/11541816-few-hundred-years-... The robots.txt seems to exclude things which are genuinely useful to exclude, like RSS feeds.
- BeefWellington 2y agoWhile true, it seems like every individual quote page also includes: <meta content='noindex' name='robots'> Worth noting as well, their sitemaps all 404.