4 ms·
Very interesting. Can't wait to give it a shot. I personally use a combination of xpath, basic math and regex, so this class/id security solution isn't a major
by anyfactor 4y ago
Very interesting. Can't wait to give it a shot.
I personally use a combination of xpath, basic math and regex, so this class/id security solution isn't a major deterrent. Couple of times, I did find it to be an hassle to scrape data embedded in iframes, and I can see the heap snapshots treat iframes differently.
Also, if a website takes the extra steps to block web scrapers, identification of elements is never the main problem. It is always IP bans and other security measures.
After all that, I do look forward using something like this and making a switch to nodejs based solution soon. But if you are trying web scraping at scale, reverse engineering should always be your first choice. Not only it enables you a faster solution, it is more ethical (IMO) as you are minimizing your impact to it's resources. Rendering full website resources is always my last choice.
- woodpanel 4y ago> is more ethical How do you deal with pages that use JS to load their content (e.g. SERPs) and restrict those endpoints to be called from within that page? I'm lucky if I can use cheerio to just traverse the DOM on a given page, but increasingly I have to render the page and that "scales" as well, at least in terms of maintainability since I can more or less use the same API to traverse the (then JS-modified) DOM
- deleted 4y ago[deleted]
- 10000truths 4y agoYou can observe the network requests that the JS makes under the Network tab of the Developer Tools console. The restrictions you mention can be bypassed by setting the Origin and Referer HTTP headers to whatever satisfies the server.
- timtom39 4y ago> But if you are trying web scraping at scale, reverse engineering should always be your first choice. Not only it enables you a faster solution, it is more ethical (IMO) as you are minimizing your impact to it's resources. Rendering full website resources is always my last choice. I find my time is by far the most limited resource. I am usually scraping huge corporations at scale and don't care/doubt I will impact their resources. If they would open their APIs I would use those. That being said, I often end up reverse engineering to preserve my own resources. I can and do run thousands of instances of chrome but it isn't cheap. Also, related to IPs, carrier grade NAT has been a blessing ;)
- anyfactor 4y ago> carrier grade NAT Are you using something you have built or a service?