4 ms·
Or you could use the API that the web app has to inevitably use.
by lapser 4y ago
Or you could use the API that the web app has to inevitably use.
- chrsig 4y agothis is assuming that it's documented...otherwise you're just hand evaluating javascript to figure out what it would call...and then you get to thinking that you should just embed a javascript interpreter and evaluate it...and at that point, you've gone down the path of implementing a headless browser.
- lapser 4y agoNot really. The API these days tends to be JSON so you can just figure out how to it works and what represents what. For example I've been able to reimplement xmltv scrapers for several sources in less than a 100 lines with Scrapy. It's not hard, just requires a little discretion.
- chrsig 4y agoThe difficulty isn't in making a scraper for a single site, but rather in the general case. That is, making a scraper that can be pointed at an arbitrary site not known at the time of development.
- vertere 4y agoI assumed most people were talking about dealing with single sites. From your previous comment about API documentation and "hand evaluating" Javascript I gathered that you were too. How would those things help one solve the general case?
- hombre_fatal 4y agoAlmost everything uses simple JSON APIs which are far more trivial than html scraping. You also don't need to evaluate the Javascript to figure out what it's doing (something I can't even imagine doing, do you really do this? and where have you done it?), just browse the website normally with your network tab open and look at the endpoints and you're basically done. Obfuscated APIs like Pokemon GO and Netflix are in the tiny minority.
- camgunz 4y agoI just wrote ~20 scrapers and maybe 3 ended up being able to grab data from a JSON API. Mostly what I ran into was (an HTTP API that returns) templated HTML, and wacky Sharepoint stuff. For the Sharepoint stuff, I found I often had to grab tokens out of script tags. Sometimes they were in hidden inputs, but either way is kind of the same thing. I was ready to break out a JS parser, but fortunately I didn't need to. I did run into Cloudflare DDoS protection and Incapsula, which I will say is pretty irritating and IMO antithetical to the web. Incapsula is so bad I get captcha'd just browsing around in a Firefox private window. If I were polling every few seconds or something I'd get it, but denylisting all AWS IPs or looking for "headless" in the User Agent (or looking at navigator params, testing TLS fingerprints, etc.) is bonkers. It's the laziest kind of upselling from web developers where you're making the site harder to use, but not actually keeping real scrapers out, because they're doing even more JavaScript interventions ahead of the HTTP request and using residential IP proxies.
- hombre_fatal 4y agoInteresting, I also have a big scraping project and the breakdown is probably like 70% HTML parsing, 25% JSON APIs, 5% weird APIs. I can of course imagine that this pie chart simply depends on the sites/genre/industry you're scraping. DDoS protection does throw a wrench into the mix, though I don't blame anyone for using it. DDoS protection might seem antithetical to the web, but... so is DDoS and abuse. Kind of like how being an asshole is antithetical to getting along as a society but you still have to address the reality that there will always be abusers and bad actors. I also think being able to do what you want with your service is a fundamental part of the web incl putting it behind a captcha. It's just part of the beautiful chaos.
- camgunz 4y agoIn the large I don't disagree: people have the right to protect themselves. But I think you've made an argument for caching, not for captchas. I'd even be fine with changing cache-control interpretations to "cache this, because if you come back before it's changed we won't serve it to you again". But this stuff is obvious web developer upsell. For what it's worth, it didn't even work. Headless Chrome and some editing of the JavaScript environment was all it took. So it's definitely a ripoff.
- matheusmoreira 4y agoLove this approach. We can just bypass all the normal web scraping and get the structured data straight from the source. These APIs are usually no less stable than the ever changing HTML structure anyways. Case study: YouTube.js https://news.ycombinator.com/item?id=31021611 https://news.ycombinator.com/item?id=31021611 https://github.com/LuanRT/YouTube.js https://github.com/LuanRT/YouTube.js