13 ms·
The State of Web Scraping in 2021
- zlib 5y agoWhat kind of stuff are people needing to scrape?
- conradfr 5y agoI have a side-project where I display the schedule of the day of 100+ French radios, like you would for TV channels. Scraping works great to get the data. I don't like node/js but I use it to do the scraping as I view the code as trash and full of edge cases and unreliable data / types and I can't complain, a dynamic scripting language is great for that.
- rmetzler 5y agoWebsites which change over time and don’t provide a simpler way of getting an update (e.g. an RSS feed or a JSON api).
- selecsosi 5y agohttps://j-archive.com/ https://j-archive.com/
- deleted 5y ago[deleted]
- ianhawes 5y agoScraping saved untold lives this past spring when large healthcare providers (i.e. Walgreens & CVS) opted to hide their vaccination appointments behind redundant survey questions. This made it more difficult to quickly ascertain when an appointment slot would become available. The elderly were less likely to look more than once a day, delaying vaccines for those that needed it the most. GoodRX built a scraping system that tapped into all the major providers. Thats what a group of vaccine hunters in my state used to get appointments for folks that had tried but were unable to.
- hyeomans 5y agoI scrape multiple government sites to fill all the data for https://www.quienmerepresenta.com.mx/ https://www.quienmerepresenta.com.mx/ It tells you who is your governor, local/federal representative, senator and municipal president. Each representative lives on a different website so I wrote scrappers for each one.
- mdaniel 5y agoI would expect it's roughly the same answers, just varying in the specifics: * those which don't offer a _reasonable_ API, or (I would guess a larger subset) those which don't expose all the same information over their API * those things which one wishes to preserve (yes, I'm aware that submitting them to the Internet Archive might achieve that goal) * and then the subset of projects where it's just a fun challenge or the ubiquitous $other As an example answer to your question, some sites are even offering bounties for scraped data, so one could scratch a technical itch and help data science at the same time: https://www.dolthub.com/repositories/pdap/datasets/bounties https://www.dolthub.com/repositories/pdap/datasets/bounties
- perlwle 5y agoBuilding a side project using python scrapy to scrape podcast shows. I use it to search by title/description etc to find interesting podcasts. Also as a way to learn different tools and frameworks.
- deleted 5y ago[deleted]
- frankfury 5y agoI love web scraping, and I used to download many images through a series of scripts that crawled throughout a certain website.
- PigiVinci83 5y agoGreat article. I've put online a discord server for sharing knowledge about web scraping, if some of you wants to join https://discord.gg/fwqhqrWhHW https://discord.gg/fwqhqrWhHW
- agustif 5y agoPuppeteer/Playwright are the easiest for me
- ardalann 5y agoself promotion: I launched my no-code scraping cloud software on ProductHunt last month after a year of testing with beta users: https://www.producthunt.com/posts/browse-ai https://www.producthunt.com/posts/browse-ai Here are a few comparisons if you're curious: - https://www.browse.ai/vs/hexomatic https://www.browse.ai/vs/hexomatic - https://www.browse.ai/vs/import-io https://www.browse.ai/vs/import-io - https://www.browse.ai/vs/octoparse https://www.browse.ai/vs/octoparse - https://www.browse.ai/vs/oxylabs https://www.browse.ai/vs/oxylabs - https://www.browse.ai/vs/parsehub https://www.browse.ai/vs/parsehub - https://www.browse.ai/vs/simplescraper https://www.browse.ai/vs/simplescraper - https://www.browse.ai/vs/webscraper https://www.browse.ai/vs/webscraper - https://www.browse.ai/vs/zyte https://www.browse.ai/vs/zyte
- anjingchi 5y agoAnyone have tried to use Cypress for scraping?
- juanse 5y agoNowadays is more and more common for websites to have some kind of rate limiting middleware such as rack attack for ruby. It would be interesting to explore the strategies to deal with it.
- hdjjhhvvhga 5y agoSame as always - proxy farms, random popular UAs with random delays etc.
- juanse 5y agoSorry, what is a UA?
- bryanrasmussen 5y agoso will Google's freezing of the UA lead to less ability to web scrape for the non big company scrapers out there?
- vivekv 5y agoSorry can you please elaborate what is Google doing?
- bryanrasmussen 5y agosorry, I thought it was a well known thing here given the various discussions over past year or so https://groups.google.com/a/chromium.org/g/blink-dev/c/-2JIRNMWJ7s/m/yHe4tQNLCgAJ https://groups.google.com/a/chromium.org/g/blink-dev/c/-2JIR... on edit: so I'm thinking as there will only be one UA floating around then, sure, older UAs can exist, but those become progressively more suspicious.
- jmnicolas 5y agoLast year I needed some quick scraping and I used a headless Chromium to render webpages and print the HTML then analyze it with C#. I don't remember exactly, but I think it was around 100 or 200 loc, so not exactly something that took long to write. In fact the most difficult thing was to figure how to pass the right args to Chromium. I wonder what does a scraping framework offer?
- Veen 5y ago> I wonder what does a scraping framework offer? HTTP requests, HTML parsing, crawling, data extraction, wrapping complex browser APIs etc. Nothing you couldn't do yourself, but like most frameworks, they abstract the messy details so you can get a scraper working quickly without having to cobble together a bunch of libraries or re-invent the wheel.
- jmnicolas 5y agoI see thanks.
- jrochkind1 5y agoJust for one example, when you have to get a form, and then submit the form, with the CSRF protection that was in the form... of course you COULD write that yourself by printing HTML and then analyzing it with C# (which triggers more requests to chromium I guess), but you're probably going to wonder why you are reinventing the wheel when you want to be getting on to the domain-specific stuff.
- jmnicolas 5y agoAh yes I see. Mine was read-only, so no need for complex stuff.
- elorant 5y agoThrottling is a prime example. If you start loading multitudes of sites in asynchronous fashion you'll have to enter some delay otherwise you run the risk of choking the server in misconfigured sites. I've DDoSed sites accidentally this way. You can of course build a framework on your own, and that's pretty much what every scraper does eventually, but it takes time and a lot of effort.
- colinramsay 5y agoIf you're familiar with Go, there's Colly too [1]. I liked its simplicity and approach and even wrote a little wrapper around it to run it via Docker and a config file: https://gotripod.com/insights/super-simple-site-crawling-and-scraping/ https://gotripod.com/insights/super-simple-site-crawling-and... [1] http://go-colly.org/ http://go-colly.org/
- IceWreck 5y ago+1 for Go. Its easy concurrency makes it an awesome language for web scraping. The go-colly framework was a bit too restrictive for my needs, but its very easy to build something on top of the standard lib's net/http, its cookiejar, and a third party library called goquery (afaik go-colly uses this too). Fun Fact: We were scraping something from an apparantly zero rate limits azure blob container, and we had to enumerate around a million URLs daily (didn't know which URLs actually existed so we guessed an offset and enumerated from there, also we had to do it at a fixed time daily). We had proxys at our disposal but didn't need them cause the blob container did not rate-limit. I wrote the scraper in Go, but a friend wrote it in Rust. Using Go was fast enough, satisfying all our requirements, but it turned out that the Rust one was 3-5 times faster. I tried to improve the Go scraper's speed by tweaking net/http transport's parameters, increasing workers, removing all NOFILE limits from SystemD and tried to profile and remove the low hanging speed issues. Nothing reduced the gap. Then I replaced the net/http client with valyala/fasthttp (another http implementation in Go) , which made it as fast as or slightly faster than the Rust one which was using the reqwest crate as http client.
- hivacruz 5y agoI used this library to get familiar with Go. It is indeed very powerful and really easy to create a scraper. My main concerns though were about testing. What if you want to create tests to check if your scraper still gets the data we want? Colly allows nested scraping and it's easy to implement but you have all your logic into one big function, making it harder to test. Did you find a solution to this? I'm considering switching to net/http + GoQuery only to have more freedom.
- colinramsay 5y ago
- amelius 5y ago> Crawl at off-peak traffic times. If a news service has most of its users present between 9 am and 10 pm – then it might be good to crawl around 11 pm or in the wee hours of the morning. How do you know this if it is not your website? Also, the internet has no time zone.
- chucky 5y agoFor sites where there is a peak usage time, it's probably obvious what that peak usage time is. A news service (their example) presumably primarily serves a country or a region - then off-peak traffic times are likely at night. The Internet has no time zone, but its human users all do.
- numeralls 5y agoIf your scraping a popular website Google Trends should be a pretty good proxy
- beardyw 5y agoI tried Python/ BeautifulSoup and Node/Puppeteer recently. It may be because my Python is poor, but puppeteer seemed more natural to me. Injecting functionality into a properly formed web page felt quite powerful and started me thinking about what you could do with it.
- abzug 5y agoOn the Ruby side both Nokogiri and Mechanize should be mentioned...
- marvram 5y agoGood call! ~ Will add them in the next version.
- gizdan 5y agoSurprised Woob[0] (formerly Weboob) isn't on the list. It's designed for specific tasks, such as getting transactions from your bank, events from different event pages, and much. [0] https://woob.tech/ https://woob.tech/
- m_ke 5y agoAnother tip, there are a few browser extensions that can record your interactions and generate a playwright script. Here's one: https://chrome.google.com/webstore/detail/headless-recorder/djeegiggegleadkkbgopoonhjimgehda?hl=en https://chrome.google.com/webstore/detail/headless-recorder/...
- sidharthv 5y agoIf you don't want to install another extension, Playwright has built in support for recording. npx playwright codegen wikipedia.org https://playwright.dev/docs/next/codegen https://playwright.dev/docs/next/codegen
- jrochkind1 5y agoI'm not familiar with "playwright", it doesn't seem to be mentioned in OP either. When I google, I see it advertised as a "testing" tool. Can I also use it for scraping? Where would I learn more about doing so?
- rmetzler 5y agoPlaywright is similar to Puppeteer, but can use different browsers not only Chrome.
- xnyan 5y agoPlaywright is essentially a headless chrome, firefox, and webkit browser with a nice API that's intended for automation/scraping. It's far more heavy than something like curl, but it has all the capabilities of any browser you want (not just chrome as with puppeteer) and makes stuff like interacting with javascript a breeze. It's similar to Google's puppeteer, but in my opinion even with chrome much more pleasant and productive. Microsoft's best developer tool IMO, saves me tons of time.
- nathell 5y agoI’ll chime in with mine: Skyscraper (Clojure) [0] builds on Enlive/Reaver (which in turn build on JSoup), but tries to address cross-cutting concerns like caching, fetching HTML (preferably in parallel), throttling, retries, navigation, emitting the output as a dataset, etc.
- throwawaysea 5y agoIs there open source software that can extract the "content" part of a given page cleanly? I'm thinking about what the reader mode in browsers can do as an example, where the main content is somehow isolated and displayed.
- specproc 5y agoI believe the main library for reader mode is called readability. I played around with a python implementation a while back. Just pipe in your raw html as part of the process. It's good, but not flawless. If I remember correctly, it included some quotes and image text as part of the body for the site I tried it on.
- stef25 5y agoThere's a PHP port of Readability and it works for some sites, for others not at all. Very far from perfect.
- tannhaeuser 5y agoYou can use SGML (on which HTML is/was based) and my LGPL-licensed sgmljs package [1] for that, plus my SGML DTD grammar for HTML5. [2] describes common tasks in preservation of Web content to give you a flavor, but you can customize what SGML does with your markup to death really; in your case, you'll probably want to throw away divs and navs to get clean semantic HTML which you can do using SGML link processes (= pipeline of markup filters and transformations), but you could also convert HTML into canonical markup (eg XML) and use Turing-complete XML processing tools such as XSLT as described in the linked tutorial. [1]: http://sgmljs.net http://sgmljs.net [2]: http://sgmljs.net/docs/parsing-html-tutorial/parsing-html-tutorial.html http://sgmljs.net/docs/parsing-html-tutorial/parsing-html-tu...
- f311a 5y agoFor Python, instead of BeautifulSoup I prefer to use selectolax which is 3-5 times faster. Also, I think very few people use MechanicalSoup nowadays. There are libraries that allow you to use headless Chrome, e.g. Playwright. It looks like the author of the article just googled some libraries for each language and didn't research the topic.
- mdaniel 5y agoLazyweb link: https://github.com/rushter/selectolax https://github.com/rushter/selectolax although I don't follow the need to have what appears to be two completely separate HTML parsing C libraries as dependencies; seeing this in the readme for Modest gives me the shivers because lxml has _seen some shit_ > Modest is a fast HTML renderer implemented as a pure C99 library with no outside dependencies. although its other dep seems much more cognizant about the HTML5 standard, for whatever that's worth: https://github.com/lexbor/lexbor#lexbor https://github.com/lexbor/lexbor#lexbor --- > It looks like the author of the article just googled some libraries for each language and didn't research the topic Heh, oh, new to the Internet, are you? :-D
- jacurtis 5y ago> It looks like the author of the article just googled some libraries for each language and didn't research the topic Yep, this seemed like an aggregate Google results page. I was initially intrigued by the article and then realized it was a list of libraries the author found via Google. There were significantly notable omissions from this list and a bunch of weird stuff that feels unnecessary. I don't think the author has actually scraped a page before.
- xnyan 5y agoI agree with your conclusion, but in any discussion about web scraping it's probably a good idea to mention BeautifulSoup given how popular it is (virtually a builtin in terms how much it's used) and given all the documentation available for it, a good starting point if perf is not going to be a concern.
- heavyset_go 5y agorequests-html is faster than bs4 using lxml. It's a wrapper over lxml. I built something similar years ago using a similar method, it was much faster than bs4, too.
- gcatalfamo 5y agoWhy no mention of selenium? Is it not cool anymore? I have never heard of mechanicalsoup: is it selenium replacement?
- nicoburns 5y agoSelenium is famously unreliable, so a lot of people have been replacing it with headless chrome where they can.
- melomal 5y agoInteresting. I was about to start on some web automation and so far I've had hammered into my head that Selenium is the 'language of the internet' or something along those lines. What would be a better solution, if you have any to recommend?
- xzel 5y agoI'd suggest Puppeteer / Playwright. Both are great. Iirc the puppeteer team largely moved to playwright.
- kjkjadksj 5y agoPuppeteer is frustrating to me. When I tried to use it I couldn’t get it to click buttons, but I did get it to hover on the button so I know I had the correct element in my code. Their click function just did nothing at all. I resorted to tabbing a certain amount of times and hitting enter.
- melomal 5y agoThank you for the suggestions! I will check them out.
- gcatalfamo 5y agoI have been using selenium with chromedriver, I mistakenly thought those were basically the same thing. Can you tell me more?
- benzible 5y agoI just needed a service to reliably fetch raw pages that I can process in my own application and so far I've been happy with this: https://promptapi.com/marketplace/description/adv_scraper-api https://promptapi.com/marketplace/description/adv_scraper-ap... $30 / month for 300K requests, rotating residential proxies, uses headless Chromium, etc.
- Jenkins2000 5y agoFor Java/Kotlin, HtmlUnit is generally pretty great.
- anjingchi 5y agoBut JavaScript not being properly executed in HtmlUnit https://stackoverflow.com/questions/19646612/javascript-not- https://stackoverflow.com/questions/19646612/javascript-not- being-properly-executed-in-htmlunit
- toastal 5y agoOCaml’s Lambda Soup (https://aantron.github.io/lambdasoup/ https://aantron.github.io/lambdasoup/) is a amazing library/, especially for those that prefer functional programming
- dec0dedab0de 5y agoScraping things that don't want to be scraped is one of my favorite things to do. At work this is usually an interface for some sort of "network appliance." Though with the push for REST APIs over the last 6 years or so, I don't have a need to do it all to often. Plus with things like selenium it's too easy to just run the page as is, and I can't justify spending the time to figuring out the undocumented API. My favorite one implemented CSRF protections by polling an endpoint, and adding in the hashed data from that endpoint and a timestamp on every request. When I hear a junior dev give up on something because the API doesn't provide the functionality of the UI, It makes me very sad that they're missing out.
- tmpz22 5y agoTo be fair selenium style scraping can take a lot of time to setup if you aren’t already familiar with the tooling, and the browser rendering apis are unintuitive and sometimes flat out broken.
- ipaddr 5y agoThat's why things like laravel's Dusk exists to put a layer over that complex experience.
- dec0dedab0de 5y agoMaybe it's because I'm using the python bindings, but it took me about an hour to go from never using it to having it do what I needed it to do. I just messed around in a jupyter notebook until I got what I needed working. Tab complete on live objects is your friend. The hardest part was figuring out where to download a headless browser from. Though I do prefer requests/bs4. I wrote a helper to generate a requests.Session object from a selenium Browser object. I had something recently where the only thing I needed the javascript engine for was a login form that changed. So by doing it this way I didn't have to rewrite the whole thing. Still kind of bothers me I didn't take the time to figure out how to do it without the headless browser, but it works fine, and I have other things to do.
- chinchilla2020 5y ago
- lucasverra 5y agoany recommendation to scrape behind google social login?
- SubiculumCode 5y agoI've not followed this space. When I did, there were a lot of questions concerning the legality of automated scraping. Have those legal issues been resolved?
- hafizhamid 5y agoAs long as you're scraping publicly available data (i.e. not going behind a login) and avoiding copyrightable content and personal data, you should be mostly fine. This article should answer most of the web scraping legality questions: https://www.crawlnow.com/blog/is-web-scraping-legal https://www.crawlnow.com/blog/is-web-scraping-legal
- rafale 5y agoWhat's the best way to get around AWS/Azure/... ip range ban and VPN ban when scrapping?
- jmuguy 5y agoThe large proxy providers operate in a sort of gray market. You pay for "residential" or "ISP" based IP addresses. In some instances these proxy connections are literally being tunneled through browser extensions running on a real world system somewhere (https://hola.org/ https://hola.org/ for instance)
- killingtime74 5y agoThere are other providers. I think big time scrapers use residential IPs
- deleted 5y ago[deleted]
- MrDresden 5y agoI've been working on a scraping project in Scrapy over the last month, using Selenium as well. My Python skills are mediocre (mostly a Java/Kotlin dev). Not only has it been a blast to try out, but also surprisingly easy to setup. I now have around 11 domains being scraped 4 times a day through a well defined pipeline + ETL then pipes it to Firebase Firestore for consumption. Next step is to write the page on top of it.
- heavyset_go 5y agoAre you using Scrapy mainly for scraping, or do you do crawling, as well?
- MrDresden 5y agoIn my case I am only using it for direct scraping.
- travisporter 5y agoI’m looking to roll my own Plaid-like service so I can download the CSVfiles from my bank account and credit card. Would selenium be the way to go?
- marban 5y agoCloudflare's protection is quite a b*tch to circumvent with any headless or python library.
- password4321 5y agohttps://news.ycombinator.com/item?id=28514998#28515629 https://news.ycombinator.com/item?id=28514998#28515629 > Cloudflare's bot protection mostly makes use of TLS fingerprinting, and thus pretty easy to bypass. https://news.ycombinator.com/item?id=28251700 https://news.ycombinator.com/item?id=28251700 -> https://github.com/refraction-networking/utls https://github.com/refraction-networking/utls Disclaimer: haven't tried it.
- alphabet9000 5y agowith node, i've had success with puppeteer-extra using puppeteer-extra-plugin-stealth
- heavyset_go 5y agoIt's a pain even when you aren't a bot. For a while there, Cloudflare's fingerprinting page would trigger Firefox on Linux to crash instantly.
- omneity 5y agoSlight aside: The most recent Cloudflare HCaptchas ask you to classify AI generated images. They don’t even look like a proper bike/truck/whatever (I don’t have an example handy). I categorically refuse to do when I’m browsing websites using it. I find this new captcha utterly unacceptable. It’s no “protection” at this point anymore. Websites are using it as an excuse to become even more user hostile. I am worried for the future of the web.
- synergy20 5y agoIn my own experience puppeteer is much better/capable than selenium but the problem is that puppeteer requires nodejs. its python-wrapper https://github.com/pyppeteer/pyppeteer https://github.com/pyppeteer/pyppeteer was not as good as selenium when you like to use python.
- novaleaf 5y agoSelf promotion: my SaaS is the lowest cost web scraping tool for high volume, and has been in business since 2016. https://PhantomJsCloud.com https://PhantomJsCloud.com My SaaS requires some technical knowledge to use (call a web api) which I suppose is why it's not ever in these lists. Some of my customers are *very* large businesses. If you are looking at evading bot countermeasures, my product isn't probably the best for you. but for TCO nothing beats it.
- lloydatkinson 5y agoIsn't phantomjs deprecated and unmaintained?
- spiffytech 5y agoYep, according to PhantomJS' README, their "development is suspended until further notice". It looks like phantomjscloud.com also supports Puppeteer.
- novaleaf 5y agoyes, bad naming on my part. While it does support PhantomJs still, the default is a Puppeteer backend.
- jamesfinlayson 5y agoFor some time now - since 2016 I think (though someone briefly tried to revive it) - headless Chrome does it faster and better now.
- mro_name 5y agoI am scraping radio broadcast pages for a decade now. Started with (ruby) scrapy, then nokogiri, then moved on to go and their html package. Currently sport a mix of curl + grep + xsltproc + lambdasoup (OCaml) and am happy with it. Sounds like a mess but is shallow, inspectable, changeable and concise. http://purl.mro.name/recorder http://purl.mro.name/recorder
- Quessked73 5y agoDoes anyone have a resource for getting into app-based scraping, if the API is obfuscated or rate limited?
- elorant 5y agoRun an http debugger, or some proxy, and find the endpoints.
- nlh 5y agoProxyMan on MacOS is quite awesome for this. It requires a bit of setup with certificates, etc., but once it's working you just fire up the target app on your phone and all the sweet sweet API requests appear on your big screen. I've scraped two apps very successfully this way. It's also fascinating to see how developers-who-aren't-me setup their APIs when they assume that nobody's looking.
- say_it_as_it_is 5y agoPyppetteer is feature complete and worth noting: https://github.com/pyppeteer/pyppeteer https://github.com/pyppeteer/pyppeteer
- marvram 5y agoThanks! Updated the blog post to include it.
- holoduke 5y agoStill using casperjs and phantomjs. Both are deprecated for many years, but I cannot find any replacement. Some of my scraping programs are running over 10 years without any issues.
- jakearmitage 5y agoIt's all fun and games until PerimeterX comes in.
- chaostheory 5y agoThe current article is better than nothing, but it missed https://playwright.dev https://playwright.dev
- marvram 5y agoThanks for the pointer! - Just included Playwright as a language agnostic tool.
- krakengerry 5y agoI think another technique that should be talked about is intercepting network responses as they happen. The web in 2021 still has a whole lot of client-side rendering. For those sites, data is often loaded on the fly with separate network calls (usually with some sort of nonce or contextual key). Much of the hassle in web scraping can be avoided by listening for that specific response instead of parsing an artifact of the JSON->JS->HTML process. I put together a toy site [0] recently that uses this approach for JIT price comparisons of events. When you click on an event, the backend navigates to requested ticket provider pages through a pool of Puppeteer instances and waits for JSON responses with pricing data. [0] https://www.wyzetickets.com https://www.wyzetickets.com