3 ms·
Can I ask you what you are using to scrape with? I am finding that Amazon's anti-bot strategies change every so often, so unless I use phantom.js or Selenium,
by laaph 10y ago
Can I ask you what you are using to scrape with? I am finding that Amazon's anti-bot strategies change every so often, so unless I use phantom.js or Selenium, I have to change odd things like headers and various things every so often.
I finally reworked my code to just download the html files using phantom.js and parse it with perl so that I can run this on a headless server, and this method hasn't broken lately. I find trying to use phantom.js is kind of crazy (mainly I have a hard time with the documentation) and javascript is not my native language. I was using curl, LWP, and wget in the past.
- nevster 10y agoI'm not using any libraries. All hand coded to just get the bits it needs in the html using lots of indexOf. With regards to the anti-bot problem, this isn't running on any server, it's running on the user's computer. The number of requests they make per second doesn't seem to trigger any anti-bot measures. I'm just using plain Java URLConnections.