4 ms·
One tip I would pass on when trying to scrape data from a website, start by using wget in mirror mode to download the useful pages. It's much faster to iterate
by VBprogrammer 6y ago
One tip I would pass on when trying to scrape data from a website, start by using wget in mirror mode to download the useful pages. It's much faster to iterate on scraping the data once you have it locally. Also, less likely to accidentally kill the site or attract the attention of the host.
- strin 6y agoThat works only for static page though. Many modern pages would require you to run a selenium or puppetteer to scrape the content.
- thaumasiotes 6y agoThat's never required; the data shows up in the web page because you requested it from somewhere. You can do the same thing in your scraper.
- dewey 6y ago> You can do the same thing in your scraper Rendering the page in Puppeteer / Selenium and then scraping it from there sounds like a lot easier than somehow trying to replicate that in your scraper?
- thaumasiotes 6y agoSure. How does that relate to the claim that your scraper is actually unable to make the same requests your browser does?
- dewey 6y agoHow are you going to deal with values generated by JS and used to sign requests?
- thaumasiotes 6y agoIf they're really being generated client-side, you're free to generate them yourself by any means you want. But also, that's a strange thing for the website to do, since it's applying a security feature (signatures) in a way that prevents it from providing any security. If they're generated server-side like you would expect, and sent to the client, you'd get them the same way you get anything else, by asking for them.
- dewey 6y agoI'm not sure what's your point. Of course you can replicate every request in your scraper / with curl if you want to if you know all the input variables. Doing that for web scraping purposes where everything is changing all the time and you have more than one target website is just not feasible if you have to reverse engineer some custom JS for every site. Using some kind of headless browser for modern websites will be way easier and more reliable.
- pocket_cheese 6y agoAs someone who has done a good bit of scraping, how a website is designed dictates how I scrape. If it's a static website that has consistently structured HTML and is easy to enumerate through all the webpages I'm looking for, then simple python requests code will work. The less clear case is when to use a headless browser vs reverse engineering JS/server side APIs. Typically, I will do like a 10 minute dive into the client side js and monitor ajax requests to see if it would be super easy to hit some API that returns JSON to get my data. If reverse engineering seems to hairy, then I will just do headless browser. I have a really strong preference for hitting JSON apis directly because, well, you get JSON! Also you usually get more data then you even knew existed. Then again, if I was creating a spider to recursively crawl a non-static website, then I think Headless is the path of least resistance. But usually, I'm trying to get data in the HTML, and not the whole document.
- shiyason 6y agoI’ve been doing web scraping for the past 5 years and this is exactly the approach I take as well!
- edmundsauto 6y agoFor these sites, I crawl using a JS powered engine, and just save the relevant page content to disk. Then I can craft my regex/selectors/etc., once I have the data stored locally. This helps if you get caught and shut down - it won't turn off your development effort, and you can create a separate task to proxy requests.
- alephu5 6y agoI did web-scraping professionally for two years, in the order of 10M pages per day. The performance with a browser is abysmal and requires tonnes of memory so not financially viable. We used them for some jobs, but rendered content isn't a problem, you can also simulate the API calls (common) and read the JSON, or regex the script and try to do something with that. I'd say 99% of the time you can get by without a browser.
- inovica 6y agoFully agree. It takes some thought :)
- johtso 6y agoOr just use scrapy's caching functionality. Super convenient.
- spsphulse 6y agoDefinitely! Scrape the disk, not the web.