10 ms·
I admit being a bit disappointed that a well-known disadvantage of web scraping was not mentioned: Web scraping is fragile! Web sites change, web frameworks ev
by b6z 6y ago
I admit being a bit disappointed that a well-known disadvantage of web scraping was not mentioned: Web scraping is fragile!
Web sites change, web frameworks evolve, and just some subtle reordering of some <divs> or renaming of CSS classes, and your perfect scraping code from yesterday will break tomorrow -- maybe not leaving you empty-handed, but probably missing some data or delivering the wrong one.
If there is an API you can use, use it. If your budget allows to pay for API access, buy it. APIs tend to be more stable than scraping, and the data provider will probably inform you if it changes. Contacting them might even get you more interesting data, as not every column they have in their database might become published on the web site.
- danso 6y agoFor the purposes of their research and study, web scraping is stable enough, given that many websites (especially government sites), aren’t overhauled frequently. And many government sites, like those run by coroners, aren’t likely to have APIs. Contacting the agency for the data is always a good step, but even if they are responsive, they may not be willing to email you on a daily basis with data updates.
- anilakar 6y agoWe used to despise web scraping and tried to dump data from an industrial system by listening to raw Modbus RS-485 traffic. Turned out that there was no way to get the register maps from the vendor nor the firmware. We ended up writing a scraper for the main unit. Once they have been installed and configured, they're likely not going to be updated unless absolutely necessary (such as when adding new, unsupported controllers to the bus).
- deleted 6y ago[deleted]
- b6z 6y agoThat is a good argument, and I should have mentioned it, yes. For a one-off job, web scraping will probably be the best choice, and maybe even the fastest to implement. I have done my own share of web scraping for personal projects (and thus know about the fragility), but I didn't care much about broken results in the long run. But in the article they mentioned re-running their program to update their data, so it could be a long-term effort. And anyone planning a long-term project reading that article and taking their advice should at least be warned about this possible problem.
- andrew_eit 6y agoI've actually had the opposite experience. After 'scraping' some forums via their APIs for weeks, I ended up realising that the data and metadata given to me by the API was so restricted (e.g. providing 'recent comments' instead of all comments) that a pure vanilla web scraping approach became the preferred option. I agree with all the points you mention about the shortcomings though and your argument is sound. This is my opinion in the other direction, APIs come with an element of trust.
- konjin 6y agoYou both are right. Apis are extortion and scraping is fragile. I once amused myself by crating random invisible divs when generating a server side html page. It made scraping resulting files impossible, and made no difference to the look of the page.
- Buttons840 6y agoI doubt it was impossible. For example, the first step a scraper might perform is to delete all invisible divs.
- barrkel 6y agoRandom invisible divs aren't likely to defeat a moderately motivated scraper, though. Depending on what they're looking for, it could be as simple as getting the inner text of a sufficiently high-up element and matching a regex. More complex scrape defeating measures I've seen are blobs of JS that need evaluating in order to generate URL parameters (all that needs doing is extract the JS and run it in a JS engine, if you don't want to drive a headless browser, with care of course!) or that need a captcha defeating (just buy some deathbycaptcha API calls).
- hansvm 6y agoI've seen this approach backfire a bit too. Rather than having to scrape web content, my work is reduced to pulling out my favorite sandboxed JS interpreter bindings, running the snippet, and extracting the rich object they just created with exactly the data I wanted. You only need a headless browser if there's a meaningful interplay between the JS and the rest of the site.
- ricardo81 6y ago'brittle' for sure, if you're requiring periodic updates.
- 1vuio0pswjnm7 6y ago"Web scraping is fragile!" Can you give any specific examples (sites and data needed from them).
- b6z 6y agoOne example, unfortunately not very scientific: Years ago, I had a bunch of Web comics I visited daily. To make this more efficient, I had a script on one of my servers which scraped the corresponding sites in the morning and produced the daily collection on a single page. Every few weeks or months, one of the comics would just not appear or not be updated (i.e. there was the same strip day after day). I had to update the code each time, checking the source code and finding the new page and location of the image. Later, some pages started using JavaScript to load the image file, and I lost interest in sophisticating my script. So, no single page comic collection for me anymore. :)
- projektfu 6y agoFWIW, I’ve generally had good luck with RSS feeds for comics.
- CallMeJim 6y agoHave you tried Dosage? https://dosage.rocks/ https://dosage.rocks/ Set up a cron job to run `dosage @` every day, it will check for new comics and download them.
- 1vuio0pswjnm7 6y agoCan you name the sites?
- djtriptych 6y agobrittle, but in a way more stable, in the sense that the provider can't cut off access suddenly without also changing their website. Also the breaks that do happen tend to be trivial to fix. The entire project needs to be managed differently. I'd look at another frequently updated project (youtube-dl comes to mind) where breaks are _expected_
- rutthenut 6y agoYep, we've had exactly this issue with the web portal on one of our applications. A customer used some 'quick and dirty' convenience tool that used web scraping to invoke the portal application pages, and wrapped that up with part of their admin workflow. When the pages changed, even slightly, it broke their admin tool. We had said not to do this, but the workflow/web-scraping tool was surely quicker and cheaper to implement than custom API development within that organisation.
- rutthenut 6y agoGoing way back in time, well before www and soap, I recall having to develop an application that would talk to another system using serial comms, effectively pretending to be a VT52-type terminal. Had to send commands, verify responses, send subsequent commands based on available options, etc. That was used to integrate between two different types of Telco exchange equipment, for provisioning of X25 circuits on the national network in the UK. Supplier of the kit did not have any form of API, just a terminal-level application, so client-side code had to pretend to be an operator. Ahh, those old days, using green-screen terminals, eh
- ajsnigrutin 6y agoBeen there, done that.... even to hack an old ciscos into a bodged-up production/testing enviroment (take config from production, change two addresses from production to testing values, and apply the whole config to a testing cisco). The good thing with ciscos and most of the old technology is, that once you write that script, it works for years... commands never change, outputs never change, some perl, a regex or five, and you're done. Doing the same with a webpage, where it's "pride month" today, "womens day" tomorrow, "day against aids" the day after, and each means a new div, a new popup, a new redirect, a backend update inbetween, to make the new banner possible, etc., is a pain in the ass.
- joshxyz 6y agoGood point. Yet a good remedy is to add real-time asserts on your scraped data and get yourself notified as soon as you encounter unexpected data formats.
- underanalyzer 6y agoThe article seems to recommend using an api if possible.
- hansvm 6y agoI've found the opposite to be true -- when an entity is maintaining an API and their website with the same data, the website is their core business. The API is prone to being incomplete, buggy, subject to sudden deprecation, unreasonably rate limited (crippling access to some objects below what a casual human user has), and so on. Conversely, overall document structure doesn't change much over time. I know it _can_; there's a social contract that APIs should change slowly while documents can change whenever, but that isn't what I observe in the wild. Even on fairly major redesigns, the overall structure has minimal edits. A technique I've used before (wasted effort in hindsight since web pages are stable and I never have to update my scrapers) is to come up with several semantically different ways of accessing a piece of data on a page. It serves two purposes; you can recover from small page changes by having the different methods vote, and you can detect most kinds of page changes by noticing discrepancies, notifying yourself that the scraper needs to be updated soon.
- achillean 6y agoI think it really depends on the type of product that the business makes money on. If one of the main products is data then I'd wager their API will have significantly more information compared to their website. If they make money via the website then yes, they're less inclined to spend resources on the API. All of our own websites are built on-top of the same public API that everybody else uses and scraping used to be a nuisance. It was also confusing because they would be able to get more data using the same free account just by using the API instead of scraping. Exactly like the OP mentioned we only show a small number of properties via the website but most scrapers never took the time to actually compare API vs website.
- Aeolun 6y agoI think this depends entirely on what you are indexing. From my experience with some 100 ish scrapers for news sites a few would break literally every day. And the only thing we really wanted was article title and date.
- taeric 6y agoMy guess is that it depends if the API is seen as customer interface or implementation detail. People are usually hesitant to constantly change how a customer interacts. All to willing to change internal details.
- kebman 6y agoSounds like job security for us web scrapers... xD
- jandrese 6y agoThere are multiple ways of doing web scraping, and some are definitely more fragile than others. I've found fully specified XPaths to be a mistake for example. It only takes one tiny change on the page to mess up the script. On the other hand, despite numerous warnings that it would be a disaster I've found I have a lot of luck maintaining regexes, even after major page reworks.
- staticautomatic 6y agoSure, but A) As implemented in code, the path doesn’t necessarily have to be fully explicit in terms of tags. You could look for a child with text containing rather than a specific class or id, for example (or get fancier with semantic similarity on tags or text). B) There’s a trade off between speed and fragility that could make a difference if your tree is deep enough that traversing it iteratively is slow enough compared to a long xpath that it becomes a limiting factor. Granted you don’t typically encounter this in a standard issue html tree but the lxml docs, for example, correctly note that xpaths can be way faster when the nesting is super deep.
- ignoranceprior 6y agoThe easy solution is to use unit testing on known URLs to ensure that all scraping cases are handled correctly.
- alfg 6y agoAlso, consider rate limits as well. Some websites may limit by IP, user-agent or by burst.
- jrumbut 6y agoKeep in mind this is a research publication. In general research systems aren't meant to be maintained long term. You do the study, or make the proof of concept, and if it has long term value then you find a long term solution. I know they mention maintenance, but having been in both fields what a web person means by maintenance and what a research person means by maintenance are an order of magnitude different.
- slowhand09 6y agoStill, para 7 excerpt. " Once your program is written, you can recapture these data whenever you need to, assuming the structure of the website stays mostly the same."