4 ms·
The whole point of using an headless browser is to work around web sites that attempt to block simple "curl" style scraping (or where you need to execute JavaSc
by devit 9y ago
The whole point of using an headless browser is to work around web sites that attempt to block simple "curl" style scraping (or where you need to execute JavaScript to scrape).
So making it detectable (intentionally, even, right there in the user agent!) is really absurd.
Or actually, it makes one wonder about Google's motives.
- williamdclt 9y agoThat's definitely not the whole point of headless browsers, that's more of a side-effect. The whole point of headless browsers is rather automation and testing.
- dx034 9y agoSame as torrents are for the distribution of legal content. That was the original thought and it's still used for that but I'd bet the majority of headless browser requests crawl websites not owned by the scraper.
- heipei 9y agoThat's one use-case for Headless browsers. Most people actually use Headless browsers to test their website, i.e. for functionality / performance / rendering.
- ericd 9y agoMaking the web harder to crawl would make it harder to create a Google competitor. I doubt that that's their intention, though.