9 ms·
Web Scraping with Python
- f311a 4y agoThat's just a basic introduction. I would not call this article "everything you need to know".
- stingraycharles 4y agoYeah, it’s as if someone posted an article “everything you need to know about cooking” and just explained the concepts of plates, pans, and some of the kitchen appliances. I guess the fact that it’s currently very high on the front page of HN kind of confirms this type of post works, though, which is unfortunate.
- daolf 4y agoHi there, co-author here. Always to improve the content we're writing here. What else would you have expected to read in such an article?
- danmur 4y agoI think it's a pretty good article personally, sounds like the complaint is just about the title :P
- edent 4y agoI think it is good. You have to remember that some of the loudest voices on here take things extremely literally. They have no concept of hyperbole for emphasis. Or, indeed, anything which makes writing interesting to read. Is your guide everything someone needs to know? No. But anyone literate in the ways of modern English understands what you mean. It is an excellent guide and I think you should consider expanding it & perhaps creating a book. Please don't be discouraged by the people on here who don't have the skill or courage to write or submit anything.
- CJefferson 4y agoNowadays I would jump straight to selenium (or similar), as most websites feature AJAX or similar, so need a full browser. Then, you don't actually do anything with selenium, click a button / link, or anything interesting.
- daolf 4y agoAgreed, this is why on the Selenium paragraph we link to this article that go much more in depth https://www.scrapingbee.com/blog/selenium-python/ https://www.scrapingbee.com/blog/selenium-python/
- pfranz 4y agoWhile it's true I often end up needing something like selenium, it's way more heavy handed and I usually reach for it last. It doesn't scale as well, harder to troubleshoot IMHO, and more libraries and dependencies to deal with in a language where that's already not great.
- is_true 4y agoI think it's actually not that bad. Scrapping is a topic that is as broad as the number of sites on the web, a minefield of corner cases.
- daolf 4y agoThank you!
- ohyoutravel 4y agoWhy do these types of low quality, seemingly spam / SEO articles always rise to the top on weekend mornings? Is it lack of competition? Easier to manipulate votes?
- mateuszbuda 4y agoMaybe they use their product to generate upvotes O.O
- daolf 4y ago"low quality". I'm hurt :( PS: we spend tens of hours writing those piece of content and even pay a technical editor to spot the typo and make it more readable since we're not native English. You might not like this post, but I can assure that genuine care was put into writing this!
- degenerate 4y agoScrolling through your article I disagree, it's high quality content. What converts it to "low quality" is the bait-n-switch title. This is not "everything you need to know" -- this is "how to get started from scratch". Metaphor would be "Everything you need to know about fixing cars" and the article shows you how to check the engine light, change oil, rotate tires, and replace spark plugs. There's just no way to make a promise that large and have your article be considered high quality.
- wintercarver 4y agoI thought it was a nice summary, concise, organized, with examples and references. Will revisit it should I need a reminder on scraping. Would not call it low quality at all. Would recommend you ignore passing comments with no constructive criticism. The title is going to be a point of contention as it’s a big claim and probably being misinterpreted as not “everything you need to know [to get started]” but rather “everything you need to know [ever is in this one article and you’ll need not read anything else]”.
- srvmshr 4y agoWe should have some community guidelines to keep out Medium/Towards Data Science and similar low-effort article sources from HN. Genuinely in favor of lesser submission vs. increased noise in submissions. Beginner articles are not taboo, but goes against having high quality insights in general. PS: Flagging is mechanism to filter by community efforts. Guidelines set some general preconditions to the quality of articles for larger dissemination.
- edent 4y agoYou can either hit the "flag" link, or submit something better.
- pedrovhb 4y agoThat's nice, but I don't see much value in learning about sockets for scraping; it's way too low a level. The lowest level I found useful was using a requests/httpx for requests and using regex to parse data when the data you're scraping has a constant enough structure and you're scraping a large number of pages, as regex is a lot faster than parsing html. I'd add that it's often worth spending some time looking at the website for alternate ways than the obvious one of getting the data you're after. sitemap.xml sometimes give useful hints. Another golden trick is to learn reverse engineering mobile app APIs with mitmproxy or something like it. Nowadays it's kind of a pain to do since Android has been locking things down more and more, but it's still quite possible. Apps very often provide endpoints that give you structured data when the web version is server-rendered HTML only, have fewer anti-scraping measures and rate limiting, and even provide data that isn't available at all for the web version.
- Toxygene 4y ago> as regex is a lot faster than parsing html This person would like a word with you -- https://stackoverflow.com/a/1732454 https://stackoverflow.com/a/1732454 :D
- hashmush 4y agoThere's a big difference between parsing HTML and > using regex to parse data when the data you're scraping has a constant enough structure Regex is fine, just don't parse the HTML itself.
- harshreality 4y agoWhat percentage of web scraper routines resort to regex when they should at least start with xpath or some equivalent parser?
- melenaboija 4y agoThe first comment says a lot about it: > I think it's time for me to quit the post of Assistant Don't Parse HTML With Regex Officer. No matter how many times we say it, they won't stop coming every day... every hour even. It is a lost cause, which someone else can fight for a bit. So go on, parse HTML with regex, if you must. It's only broken code, not life and death
- shahidkarimi 4y agoScrapy is there to make all these happening under a single framework.
- bschne 4y agoAside: The last time I had to scrape a lot of data from the web, I additionally used SQLite, which is a breeze to use with Python (basically one import statement and you're set). It might be overkill for some cases, but I found it a huge boon for keeping track of which pages were scraped, which failed, and doing subsequent data processing and parsing "offline". It made it so much easier to recover from the inevitable random error or different markup somewhere deep in your list of pages to scrape etc.
- stall84 4y agoThis is great.. Mainly because the very first thing he does is explain the network requests themselves, focussing on the (somehow often left-out) fact that you are going to have to spoof a browser (or headers associated with it) almost always these days to get around bot-protections.
- kaycebasques 4y agoI think the overall software architecture approach of this post is fundamentally backwards. Given how much of the web is rendered client-side these days you need to start out with a headless option. Headless means that you fire up a true browser and then automate the actions that you need to perform. It's indistinguishable from a real person using a browser. If you try to use urllib3 on a webpage that does heavy client-side rendering then you're going to get incomplete HTML returned from the server (because the website intends to use JavaScript to complete the rendering of the page). On the more rare occasions when you are dealing with static HTML (i.e. there is no rendering on the client; the HTML returned from the server is the complete content) then you can use something like urllib3. Re: which headless library to use this post mentions Selenium which was one of the first headless libs but from what I've heard probably not the best (in terms of developer experience or reliability or robustness) but that's only hearsay... I've never used Selenium myself. Playwright seems like the best option in town if you want to use Python. Built by the former Chrome DevTools team (meaning those people really know how browser internals work). https://playwright.dev/python/docs/intro https://playwright.dev/python/docs/intro
- datalopers 4y agoHeadless scraping quickly becomes a very expensive approach when you try and scale the effort. I only employ it when absolutely necessary. And it’s most definitely distinguishable by any modern (incapsula, perimeterx, cloudflare) WAF.
- apienx 4y agoThe Apify library tries to address most of these issues (I'd say quite successfully). Apify.com provides a platform you can deploy the scrapers on. And yes, at scale. There's also a software marketplace where you can order custom scrapers. 98% of the projects ran thru it have a 5-star rating (Disclaimer: I moderate that marketplace). Pro-tip: submit your project with a Gmail address to skip sales and reach me directly.
- pocket_cheese 4y ago
- dmortin 4y agoHow do scrapers deal with randomized classes in web pages which is more and more common these days? Relying on the page structure only is not a robust alternative.
- edmundsauto 4y agoI've had some success with running a meta-scraper that will search for known value on a page, then back out the page structure from there. It won't help with randomly generated class names, but 95% of tasks I've written aren't this complex. For sites that are hard to scrape (usually bigger sites that get scraped a lot), I pivot towards buying a data feed. Economies of scale incentivize these data companies towards putting someone on maintaining the feed full-time.
- chasd00 4y agoWhat I’ve done is pay very close attention to the network traffic in your browsers dev tools. The data has to get to the browser somehow. Once you’re able to get a session token/cookie then you can figure out what GETs or POSTs you need to get the data you want by watching the requests your browser makes.
- photochemsyn 4y agoOne problematic thing that pops out immediately for a Python-centric approach is that they don't mention that this is all best done in some kind of Python virtual environment, like miniconda or virtualenv. They just suggest 'pip install package', which is not a good approach for anyone (and particularly not beginners) - unless you want to end up with this: https://xkcd.com/1987/ https://xkcd.com/1987/ Looking around a bit with the requirement that the online tutorial mention this rather important fact, I found this alternative option, which helpfully notes: We want to run all our scraping projects in a virtual environment, so we will set that up first. https://python-adv-web-apps.readthedocs.io/en/latest/scraping.html https://python-adv-web-apps.readthedocs.io/en/latest/scrapin... Compare and contrast that discussion with the one presented in this post - the above is far superior. Also, I don't understand why one would suggest PostGreSQL to a beginner when sqlite3 is included already in Python, and is going to be easier to use for small databases. Towardsdatascience seems to have a nice intro-to-sqlite3 tutorial.
- mynameismon 4y agoPerhaps the only issue I would have with this blogpost is using Postgres. By all means, SQLite can do the exact same thing, just easier for a beginner, since they don't have to wade through a mess of networking. Just add the binary to the PATH and one is good to go.
- taosx 4y agoJust don't forget to optimize it for writes (WAL-mode...etc) when having lots of sources.
- ducktective 4y agoConsider pup https://github.com/EricChiang/pup https://github.com/EricChiang/pup
- inshadows 4y agoHow do scrapers deal with being nice to a website these days? I'm talking multiple IPs, request rate, exponential backoff. Is there any body of knowledge for this?
- TBurette 4y agoIs there a good way to combine Scrapy framework (retry, rate limiting,..) with a headless browser such as selenium (to get full js-loaded client-side data)? When I had to do it I ended up duplicating each page request twice. Once for scrapy and once again with selenium.
- ihartley 4y agoYou can use something like scrapy-playwright[0] to run a headless browser framework as your download handler. I think there are versions for some of the other headless systems, if you prefer those. [0] https://github.com/scrapy-plugins/scrapy-playwright https://github.com/scrapy-plugins/scrapy-playwright
- samwillis 4y agoscrapy-playwright is good, and Playwright is awesome. However due to the architecture of Playwright it just keeps accumulating memory until it crashes. You will want to set up your scraper to save its state regularly, cleanly shut down and restart. But once you have that working it does work well.
- holografix 4y agoHow do people get around browser finger printing by “Sign in with Google” these days? All I get is “your browser is not safe” etc which blocks me completely.