5 ms·
> Scraping things that don't want to be scraped If all else fails, no website can withstand OCR-based screen scraping. It is slow(er), but fast enough for many
by eastendguy 5y ago
> Scraping things that don't want to be scraped
If all else fails, no website can withstand OCR-based screen scraping. It is slow(er), but fast enough for many use cases.
- elorant 5y agoAssuming that you eventually manage to load the page somehow. Which in some edge cases may entail simulating mouse movements and random delays.
- eastendguy 5y agoAgreed. -> I use the ui.vision extension to simulate native mouse movements.
- mkl 5y agoA browser extension is probably an easier way to extract text than OCR (unless you're targeting a wide range of sites, I suppose).
- timwis 5y agoHave you tried on a page protected by cloudflare captcha?
- dec0dedab0de 5y agoI have not had to deal with that, but I have idly thought that it might be easier to pipe the audio version into google assistant or something, and see what it comes up with.
- 1vuio0pswjnm7 5y agoIts funny I never seem to hit these infamous Clouflare captchas. The only impediment I encounter with Cloudflare is they require plaintext SNI to read their blog, https://blog.cloudflare.com https://blog.cloudflare.com. Unlike almost all other Cloudflare, ESNI will not work.
- eastendguy 5y agoIt seems to be no problem if you automate a real browser as opposed to a headless browser. I think they test for that.