12 ms·
Web Scraping in Python – The Complete Guide
- simonw 3y agoI strongly recommend adding Playwright to your set of tools for Python web scraping. It's by far the most powerful and best designed browser automation tool I've ever worked with. I use it for my shot-scraper CLI tool: https://shot-scraper.datasette.io/ https://shot-scraper.datasette.io/ - which lets you scrape web pages directly from the command line by running JavaScript against pages to extract JSON data: https://shot-scraper.datasette.io/en/stable/javascript.html https://shot-scraper.datasette.io/en/stable/javascript.html
- sakisv 3y agoCame here to write about Playwright. I've been using it for the last ~13 months to scrape supermarket prices and it's been a great experience.
- mrtimo 3y agowould love to learn more about what you are doing with supermarket prices
- dommer 3y ago+1 Interested in this area as well.
- sakisv 3y agoMy main drive was to document the crazy price hikes that's been going on in my home country, Greece, so I'm scraping its 3 biggest supermarkets and keep track of the prices of their products* Had a lot of fun building and automating the scraping, especially in order to get around some bot catching rules that they have. For example one of them blocks all the requests originating from non-residential IPs, so I had to use tailscale to route the scraper's traffic through my home connection and take advantage of my ISP's CGNAT. You can take a look here: https://pricewatcher.gr/en/ https://pricewatcher.gr/en/ * I'm not doing any deduplication or price comparisons between the supermarkets, I only show historical prices of the same product, to showcase the changes.
- deanishe 3y agoI can't get prices on my local supermarket's website (in Germany) without selecting a specific branch. Seems kinda suspicious to me.
- sakisv 3y agoYeah in this case between these 3 supermarkets there are 2 options: 1. You choose your general location or 2. You don't choose In the first option you get one less general category to choose from (for example they may not have fresh fish) As far as I can tell, in both cases the supermarket closest to the delivery address is responsible for filling out your order and usually what happens is that they call you to let you know that they don't have something and suggest substitutions.
- samstave 3y agoI too choose this guys supermaket scraper! What Ive long wanted was the the ability to map prices to SCUs by having folks simply take a pic of the UPC + price, just like gasbuddy or what not - in addition to scraping from grocery posting their coupon sheets online for scraping, in addition to people just scanning (non-PII) portions of receipts. Can you share what you've made thus far? * could it be used as an automated "price matching" finder? (for those companies that do a "we price match!" thing
- Oras 3y ago+1 for playwright. The codegen is a brilliant way to simplify scraping.
- BeetleB 3y agoI actually use your shot-scraper tool (coupled with Mozilla's Readability) to extract the main text of a site (to convert to audio and listen via a podcast player). I love it! Some caveats though: - It does fail on some sites. I think the value of scrapy is you get more fine grained control. Although I guess if you can use any JS with shot-scraper you could also get that fine grained control. - It's slow and uses up a lot of CPU (because Playwright is slow and uses up a lot of CPU). I recently used shot-scraper to extract the text of about 90K sites (long story). Ran it on 22 cores, and the room got very hot. I suspect Scrapy would use an order of magnitude less power. On the plus side, of course, is the fact that it actually executes JS, so you can get past a lot of JS walls.
- simonw 3y agoWow, you're really putting it through its paces! When you ran it against 90,000 sites were you running the "shot-scraper" command 90,000 times? If so, my guess is that most of that CPU time is spent starting and stopping the process - shot-scraper wasn't designed for efficient start/stop times. I wonder if that could be fixed? For the moment I'd suggest writing Playwright code for 90,000 site scraping directly in Python or JavaScript, to avoid that startup overhead.
- BeetleB 3y agoYes, indeed I launched shot-scraper command 90K times. Because it's convenient :-) I didn't realize starting/stopping was that expensive. I thought it was mostly the fact that you're practically running a whole browser engine (along with a JS engine). If I do this again, I'll look into writing the playwright code directly (I've never used it).
- ravenstine 3y agoReadability is great, and I use it, but it's odd how half-assed the maintenance for it has been. I've haven't seen any noticeable improvements to it in quite some time, and when I've looked for alternatives, it usually turns out they're using it under the hood in some capacity. Perhaps it's already being made obsolete by LLM technologies? I'd be curious to hear from anyone who's used a locally running LLM to extract written content, especially if it's been built specifically for that task.
- thrdbndndn 3y agoKinda tangent, but Playwright's doc (specifically, the intro https://playwright.dev/python/docs/intro https://playwright.dev/python/docs/intro ) confuses me. It asks you to write a test and then run `pytest`, instead of just letting you to use the library directly (which exists, but is buried in the main text: https://playwright.dev/python/docs/library https://playwright.dev/python/docs/library). I understand that using Playwright in tests is probably the most common use case (it's even in their tagline) but ultimately the introduction section of a lib should be about the lib itself, not certain scenario to use it with a 3rd-party lib B (`pytest`). Especially when it may cause side effect (I wasn't "bitten" by it but surely was confusing: when I was learning it before, I created test_example.py as said in a minefield folder which has batch of other test_xxxx.py files. And running `pytest` causes all of them to run, and gives confusing outputs. And it's not obvious to me at all, since I've never used pytest before and this is not a documentation about pytest, so no additional context was given.) > tagline
- simonw 3y agoHah yeah that's confusing. https://playwright.dev/python/docs/intro https://playwright.dev/python/docs/intro is actually the documentation for pytest-playwright - their pytest plugin. https://playwright.dev/python/docs/library https://playwright.dev/python/docs/library is the documentation for their automation library. I just filed an issue pointing out that this is confusing. https://github.com/microsoft/playwright/issues/29579 https://github.com/microsoft/playwright/issues/29579
- PaulHoule 3y agoBack in the day I used to use HTMLUnit https://htmlunit.sourceforge.io/ https://htmlunit.sourceforge.io/ to crawl Javascript-based sites from Java. I think it was originally intended for integration tests but it sure works well for webcrawlers. I just wrote a Python-based webcrawler this weekend for a small set of sites that is connected to a bookmark manager (you bookmark a page, it crawls related pages, builds database records, copies images, etc.) and had a very easy time picking out relevant links, text and images w/ CSS selectors and beautifulsoup. This time I used a database to manage the frontier because the system is interactive (you add a new link and it ought to get crawled quickly) but for a long time my habit was writing crawlers that read the frontier for pass N from a text file which is one URL per line and then write the frontier for pass N+1 to another text file because this kind of crawler is not only simple to write but it doesn't get stuck in web traps. I have a few of these systems that do very heterogenous processing of mostly scraped content and something think about setting up a celery server to break work up into tasks .
- tnolet 3y ago100%. Playwright (which does have Python support) is completely owning this scene. The robustness is amazing.
- thundergolfer 3y agoWe use shot-scraper internally to automate keeping screenshots in our documentation up-to-date. Thanks for the tool![1] Agree that Playwright is great. It's super easy to run on Modal.[2] 1. https://modal.com/docs/guide/workspaces#dashboard https://modal.com/docs/guide/workspaces#dashboard 2. https://modal.com/docs/examples/web-scraper#a-simple-web-scraper https://modal.com/docs/examples/web-scraper#a-simple-web-scr...
- 3abiton 3y agoHow does it compare to selenium or puppeteer?
- black3r 3y agoPlaywright is a rewrite of puppeteer by people who worked on puppeteer before, but now under Microsoft instead of Github. Not sure if it reached feature parity yet, but all the things we used to do with puppeteer work with playwright, and it seems to be more actively developed.
- sam2426679 3y agoIme playwright is selenium plus some, e.g. you can inspect network activity without having a separately configured proxy.
- sam2426679 3y agoCan anyone recommend a good methodology for writing tests against a Playwright scraping project? I have a relatively sophisticated scraping operation going, but I haven’t found a great way to test methods that are dependent on JavaScript interaction behind a login. I’ve used Playwright’s har recording to great effect for writing tests that don’t require login, but I’ve found that har recording doesn’t get me there for post-login because the har playback keeps serving the content from pre-login (even though it includes the relevant assets from both pre and post login.)
- 65 3y agoI'm not sure why Python web scraping is so popular compared to Node.js web scraping. npm has some very well made packages for DOM parsing, and since it's in Javascript we have more native feeling DOM features (e.g. node-html-parser using querySelector instead of select - it just feels a lot more intuitive). It's super easy to scrape with Puppeteer or just regular html parsers on a Lambda.
- aosaigh 3y agoBecause it’s been around longer. Beautiful Soup was first released in 2004 according to its wiki page and I’m sure there were plenty of libraries before it.
- macintux 3y agoPerhaps more of the people who need to run this kind of data scraping operation are comfortable with Python. Data scientists, operations personnel, etc. I've been using Perl and Python for 30 years, and JS for a few weeks scattered across those same years.
- danpalmer 3y agoHaving done a lot of web scraping, the thing that often matters is string processing. Javascript/Node are fairly poor at this compared to Python, and lack a lot of the standard library ergonomics that Python has developed over many years. Web scraping in Node just doesn't feel productive. I'd imagine Perl is also good for those in that camp. I've also used Ruby and again it was nice and expressive in a way that JS/Node couldn't live up to. Lastly, I've done web scraping in Swift and that felt similar to JS/Node – much more effort to do data extraction and formatting, not without benefits of course. I also suspect that DOM-like APIs are somewhat overrated here with regards to web scraping. JS/Node would only have an emulation of DOM APIs, or you're running a full web browser (which is a much bigger ask in terms of resources, deployment, performance, etc), and to be honest, lxml in Python is nice and fast. I generally found XPath much better for X(HT)ML parsing than CSS selectors, and XPath support is pretty available across a lot of different ecosystems.
- staticautomatic 3y ago
- thomasisaac 3y agoWe've used ScraperAPI for a long time: https://www.scraperapi.com/ https://www.scraperapi.com/ Couldn't recommend them more.
- croemer 3y agoI've used ScrapingBee which has similar pricing and has worked well, can't say which one is better: https://www.scrapingbee.com/ https://www.scrapingbee.com/
- screye 3y agoWe used to be on ScraperAPI, but moved to ScrapingBee after more frequent failures from ScraperAPI. If your scraping needs have realtime requirements, then I'd recommend ScrapingBee.
- thomasisaac 3y agoWeird, we found the exact opposite - what were you scraping? ScrapingBee really struggles on so many domains - ScraperAPI is almost as good as Brightdata when it comes to hard to beat sites.
- daolf 3y agoHi Thomas, really sorry you had a bad experience with ScrapingBee. Would you mind sending me the account you used as I wasn't able to find anything under Thomas Isaac or Tillypa and couldn't see what was going wrong then. I'm sure your comment has nothing to do with the fact that you share the same investor as ScraperAPI but I just wanted be sure.
- deleted 3y ago[deleted]
- DoodahMan 3y agoany HN discount by chance? ;) i'm testing y'all out now for a time-sensitive scrape job that must be done by Mar 1st.
- cnqso 3y agoAny modern web scraping set up is going to require browser agents. You will probably have to build your own tools to get anything from a major social media platform, or even NYT articles.
- mr_00ff00 3y agoMay be misunderstanding what you mean by “browser agents” but I’ve done some web scraping that had dynamic content and it was easy with a simple chrome driver / gecko driver + scraper crate in Rust
- philippta 3y agoShameless plug: Flyscrape[0] eliminates a lot of boilerplate code that is otherwise necessary when building a scraper from scratch, while still giving you the flexibility to extract data that perfectly fit your needs. It comes as a single binary executable and runs small JavaScript files without having to deal with npm or node (or python). You can have a collection of small and isolated scraping scripts, rather than full on node (or python) projects. [0]: https://github.com/philippta/flyscrape https://github.com/philippta/flyscrape
- simonw 3y agoDoes Flyscrape execute JavaScript that is on the page (e.g. by running a headless browser) or is it just parsing HTML and using CSS selectors to extract code from a static DOM?
- philippta 3y agoAs of right now Flyscrape just parses HTML using CSS selectors from static DOM. But as more than enough websites there days are just an empty shell I am working on adding browser rendering support.
- thijsvandien 3y agoI thought scraping is kind of dead given all the CAPTCHAs and auth walls everywhere. The article does mention proxies and rate limiting, but could anyone with (recent) practical experience elaborate on dealing with such challenges?
- caesil 3y agoNot only is scraping not dead but it has won the arms race. There are ways around every defense, and this will only accelerate as AI advances. The CAPTCHAs and walls are more of a desperate, doomed retreat.
- annowiki 3y agoHow do you get around 403/401's from WSJ/Reuters/Axios? Because I've tried user agent manipulation and it seems like I'd have to use selenium and headless to deal with them.
- accidbuddy 3y agoSome months ago, I had problems with captcha. I tried to write an application to access many drugstores and compare the price, but captcha with login system fail the mission. Do you have any piece of advice for me?
- nico 3y agoNot sure about the currently available tools, given the break-neck speed of AI progress, but a couple of years ago I built a scraper that used a captcha-solving service, they sell something like 1000 solutions for $10, it was super cheap. The process was a bit slow because they were using humans to solve the captchas, but it worked really well
- deleted 3y ago[deleted]
- givemeethekeys 3y agoAre scrapers written on a per-website basis? Are there techniques to separate content from menus / ads / filler / additional information, etc? How do people deal with design changes - is it by rewriting the scraper whenever this happens? Thanks!
- staticautomatic 3y agoYeah it’s often gonna be a per site, lots of xpath queries, email me when it breaks kind of endeavor.
- spaniard89277 3y agoYeah. I managed to abstract a bit the structure but in the end websites change.
- hubraumhugo 3y agoI got so annoyed by this kind of tedious web scraping work (maintenance, proxies, etc.) that I'm now trying to fully automate it with LLMs. AI should automate repetitive and un-creative work, and web scraping definitely fits this description. It's a boring but challenging problem. I've started using LLMs to generate web scrapers and data processing steps on the fly that adapt to website changes. Using an LLM for every data extraction, would be expensive and slow, but using LLMs to generate the scraper code and subsequently adapt it to website modifications is highly efficient. The service is using many small AI agents that basically just pick the right strategy for a specific sub-task in our workflows. In our case, an agent is a medium-sized LLM prompt that has a) context and b) a set of functions available to call. Tasks involve automatically deciding how to access a website (proxy, browser), naviage through pages, analyze network calls, and transform the data into the same structure. The main challenge: We quickly realized that doing this for a few data sources with low complexity is one thing, doing it for thousands of websites in a reliable, scalable, and cost-efficient way is a whole different beast. The integration of tightly constrained agents with traditional engineering methods effectively solved this issue. Feel free to give it a try: https://www.kadoa.com/add https://www.kadoa.com/add
- nico 3y agoKadoa looks great. For tool discovery/usage, are you using LangChain or something else? Also, do you support scraping private sites, ie. sites that require a login/password to access the data to scrape? Thank you!
- kalev 3y ago+1 on the question about scraping behind authentication. One huge use case we have as an ecommerce store is to crawl data from our vendors, which do not have (or incomplete) export files
- hubraumhugo 3y agoWe found LangChain and other agentic frameworks to have too much overhead, so we built our own tailored orchestration layer. Authenticated scraping is currently in beta, could you email me your use case (see my profile)?
- 1-6 3y agoHow many complete guides are out there for Python Scraping?
- lagt_t 3y agoHow expensive are the content bundles?
- SinjonSuarez 3y agoCheck out the cloudscraper library if are having speed/cpu issues with sites that require js/have cloudfare defending them. That plus a proxy list plus threading allows me to make 300 requests a minute across 32 different proxies. Recently implemented it for a project: https://github.com/rezaisrad/discogs/tree/main/src/managers https://github.com/rezaisrad/discogs/tree/main/src/managers
- mndgs 3y agoNicely written scraper, btw. Good code.
- SinjonSuarez 3y agoappreciate that! as a few mentioned here, there’s a lot of useful scraping tools/libraries to leverage these days. headless selenium no longer seems to make sense to me for most use cases
- fireant 3y agoI've found myself writing the same session/proxy/rate limiting/header faking management code over and over for my scrapers. I've extracted it into it's own service that runs in docker and acts as a MITM proxy between you and target. It is client language agnostic, so you can write scrapers in python, node or whatever and still have great performance. Highly recommend this approach, it allows you to separate infrastructure code, that gets highly complex as you need more requests, from actual spider/parser code that is usually pretty straightforward and project specific. https://github.com/jkelin/forward-proxy-manager https://github.com/jkelin/forward-proxy-manager
- SinjonSuarez 3y agoThis is great, was totally in the back of my mind as a next step.
- zopper 3y agoThis guide (and most other guides) are missing a massive tip: Separate the crawling (finding urls and fetching the HTML content) from the scraping step (extracting structured data out of the HTML). More than once, I wrote a scraper that did both of these steps together. Only later I realized that I forgot to extract some information that I need and had to do the costly task of re-crawling and scraping everything. If you do this in two steps, you can always go back, change the scraper and quickly rerun it on historical data instead of re-crawling everything from scratch.
- jjice 3y agoI've found this to be a good practice for ETL in general. Separate the steps, and save the raw data from "E" if you can because it makes testing and verifying "T" later much easier.
- iamacyborg 3y agoI realise from working a few places that this isn't entirely common practice, but when we built the data warehouse at a startup I worked at, we engaged with a consultancy who taught us the fundamentals of how to do it properly. One of those fundamentals was separating out the steps of landing the data vs subsequent normalisation and transformation steps.
- ethbr1 3y agoIt's unfortunate that "ETL" stuck in mindshare, as afaik almost all use cases are better with "ELT" I.e. first preserve your raw upstream via a 1:1 copy, then transform/materialize as makes sense for you, before consuming Which makes sense, as ELT models are essentially agile for data... (solution for not knowing what we don't yet know)
- dragonwriter 3y agoI think ETL is right from the perspective where E refers to “from the source of data” and L refers to “to the ultimate store of data”. But the ETL functionality should itself lives in a (sub)system that has its own logical datastore (which may or may not be physically separate from the destination store), and things should be ELT where the L is with respect to that store. So, its E(LTE)L, in a sense.
- naiv 3y agoI wonder what percentage of Google's daily searches are actually coming from scrapers
- aaroninsf 3y agoAs someone who works at a non-profit which is increasingly and regularly crawled, sometimes very aggressively, PLEASE PLEASE PLEASE establish and use a consistent useragent string. This lets us load balance and steer traffic appropriately. Thank you.
- codingminds 3y agoMind sharing which one? I'm curious
- dogman144 3y agoThere was a similar guide on HN titled something like "how to scrape like the big boys" which dug into a setup using mobile IPs, racks of burner phones, and so on. It's been lost to a bad bookmark setup of mine, and if anyone has a lead on that resource, please link, thank you and unlimited e-karma heading your way.
- recursive4 3y agohttps://news.ycombinator.com/item?id=29117022 https://news.ycombinator.com/item?id=29117022
- dogman144 3y agoAmazing, this it! Sincere thanks, been looking around for this for a few years, looks like my HN search abilities needs work.
- deleted 3y ago[deleted]
- justinzollars 3y agoI've tried this for the first time recently in 10 years - it's really become a miserable chore. There are so many countermeasures deployed to web scraping. The best path forward I could imagine is utilizing LLMs, taking screenshots and having the AI tell me what it sees on the page; but even gathering links is difficult. xml site maps for the win.
- moritonal 3y agoLiterally step for step what I spent my weekend putting together. Here's the preview blog I wrote on it. https://blog.bonner.is/using-ai-to-find-fencing-courses-in-london/ https://blog.bonner.is/using-ai-to-find-fencing-courses-in-l... Only step you missed was embeddings to avoid all the privacy pages, and a cookie banner blocker (which arguably the AI could navigate if I cared).
- justinzollars 3y agoAwesome! Things have gotten so bad this is the only alternative. I tried building a hobby search engine then quickly gave up, but did imagine how I would do the scraping!
- evilsaloon 3y agoAlways funny seeing SaaS companies pitch their own product in blog posts. I understand it's just how marketing works, but pitching your own product as a solution to a problem (that you yourself are introducing, perhaps the first time to a novice reader) never fails to amuse me.
- bilater 3y agoI'm convinced there is a gold mine sitting right in front of us ready to be picked by someone who can intelligently combine web scraping knowledge with LLMs e.g. scrape data, feed it into LLMs do get insights in an automated fashion. I don't know exactly what the final manifestation looks like but its there and will be super obvious when someone does it.
- bfeynman 3y agoI feel that the more immediate and impactful opportunity that people are doing is instead of scraping to get/understand content. LLM agents can just interactively navigate websites and perform actions. Parsing/Scraping can be brittle with changes, but an LLM agent to perform an action can just follow steps to search, click on results, and navigate like a human would
- rashkov 3y agoAre you aware of any projects for this? I began to build my own but quickly saw that the context window is not large enough to hold the DOM of many websites. I began to strip unnecessary things from the DOM but it became a bit of a slog. L
- generalizations 3y agoI tried that. Turns out that LLM-generated regex is still better (and a lot faster) than using an LLM directly.
- zffr 3y agoHere are some tips not mentioned: 1. <domain>/robots.txt can sometimes have useful info for scraping a website. It will often include links to sitemaps that let you enumerate all pages on a site. This is a useful library for fetching/parsing a sitemap (https://github.com/mediacloud/ultimate-sitemap-parser https://github.com/mediacloud/ultimate-sitemap-parser) 2. Instead of parsing HTML tags, sometimes you can extract the data you need through structured metadata. This is a useful library for extracting it into JSON (https://github.com/scrapinghub/extruct https://github.com/scrapinghub/extruct)
- deanishe 3y agoThis. A lot of modern sites can be really easy to scrape. Lots of machine-readable data. APIs (for SPAs), OpenGraph/LD+JSON data in <head>, and data- attributes with proper data in them (e.g. a timestamp vs "just now" in the text for the human). Scraping is a lot easier than it used to be.
- zffr 3y agoAdding on to this, if an app uses client-side hydration (ex Next apps) sometimes you can find a big JSON object in the HTML with all the page data. In these cases you can usually write some custom code to extract and parse this JSON object. Sometimes the JSON is embedded in some JavaScript code so you need to use a little regex to extract it.
- korbinschulz 3y ago[dead]
- f311a 3y ago>BeautifulSoup > Features: Excellent HTML/XML parser, easy web scraping interface, flexible navigation and search. It does not feature any parser. It’s basically a wrapper over lxml. >lxml > Features: Very fast XML and HTML parser. It’s fast, but there are alternatives that are literally 5x faster. This article is just another rewrite of a basic introduction. It’s not a guide, since it does mot describe any issues that you face in practice.
- thrdbndndn 3y agoBeautiful Soup comes with a "html.parser", and by default it doesn't not use or even install lxml.
- cmdlineluser 3y agoI'm sorry but BeautifulSoup is not just a wrapper over lxml. lxml even has a module for using beautifulsoup's parser. > lxml can make use of BeautifulSoup as a parser backend https://lxml.de/elementsoup.html https://lxml.de/elementsoup.html > A very nice feature of BeautifulSoup is its excellent support for encoding detection which can provide better results for real-world HTML pages that do not (correctly) declare their encoding.
- antisthenes 3y agoParsing HTML super-fast is very low on the list of priorities when web-scraping things. Yes, in practice. Most of the time it won't even register on the scale, compared to the time spent sending/receiving requests and data.
- labaron 3y agolxml is written in Cython and is very efficient in my tests. Much faster than BeautifulSoup, which is pure Python. What alternatives are 5x faster?
- calf 3y agoI've been writing rudimentary Python scripts to scrape online recipe websites for my hobby cooking purposes, and I wish there was some general software that could do this more simply. One of the websites has started making their images unclickable, so measures like that make me think it might become harder to automatically fetch such content.
- DishyDev 3y agoI've had to do a lot of scraping recently and something that really helps is https://pypi.org/project/requests-cache/ https://pypi.org/project/requests-cache/ . It's a drop in replacement for the requests library but it caches all the responses to a sqlite database. Really helps if you need to tweak your script and you're being rated limited by the sites you're scraping.
- konexis 3y agoGood luck bypassing akamai
- sakisv 3y agoThe way I'm bypassing it is by using tailscale to route the scraper's traffic through my home connection and take advantage my ISP's CGNAT. Works like a charm.
- brianarbuckle 3y agoIt's much simpler to get the links via pandas read_html: import pandas as pd tables = pd.read_html('https://commons.wikimedia.org/wiki/List_of_dog_breeds https://commons.wikimedia.org/wiki/List_of_dog_breeds', extract_links="all") tables[-1]
- throwaway81523 3y agoThis is basically an advertisement for the site's scraping proxy service.
- anotherpaulg 3y agoI recently used Playwright for Python [0] and pypandoc [1] to build a scraper that fetches a webpage and turns the content into sane markdown so that it can be passed into an AI coding chat [2]. They are both powerful yet pragmatic dependencies to add to a project. I really like that both packages contain wheels or scriptable methods to install their underlying platform-specific binary dependencies. This means you don't need to ask end users to figure out some complex, platform-specific package manager to install playwright and pandoc. Playwright let's you scrape pages that rely on js. Pandoc is great at turning HTML into sensible markdown. For example, below is an excerpt of the openai pricing docs [3] that have been scraped to markdown [4] in this manner. [0] https://playwright.dev/python/docs/intro https://playwright.dev/python/docs/intro [1] https://github.com/JessicaTegner/pypandoc https://github.com/JessicaTegner/pypandoc [2] https://github.com/paul-gauthier/aider https://github.com/paul-gauthier/aider [3] https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turbo https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turb... [4] https://gist.githubusercontent.com/paul-gauthier/95a1434a28d9e5c7b62c4e1539db9955/raw/589efa9498203776f9b40797817022c8ed82a409/tmp.pricing.md https://gist.githubusercontent.com/paul-gauthier/95a1434a28d... ## GPT-4 and GPT-4 Turbo GPT-4 is a large multimodal model (accepting text or image inputs and outputting text) that can solve difficult problems with greater accuracy than any of our previous models, thanks to its broader general knowledge and advanced reasoning capabilities. GPT-4 is available in the OpenAI API to [paying customers](https://help.openai.com/en/articles/7102672-how-can-i-access-gpt-4). Like `gpt-3.5-turbo`, GPT-4 is optimized for chat but works well for traditional completions tasks using the [Chat Completions API](/docs/api-reference/chat). Learn how to use GPT-4 in our [text generation guide](/docs/guides/text-generation). +-----------------+-----------------+-----------------+-----------------+ | Model | Description | Context window | Training data | +=================+=================+=================+=================+ | gpt | | 128,000 tokens | Up to Dec 2023 | | -4-0125-preview | | | | | | New | | | | | | | | | | | | | | | | | | | | **GPT-4 | | | | | Turbo**\ | | | | | The latest | | | | | GPT-4 model | | | | | intended to | | | | | reduce cases of | | | | | "laziness" | | | | | where the model | | | | | doesn't | | | | | complete a | | | | | task. Returns a | | | | | maximum of | | | | | 4,096 output | | | | | tokens. [Learn | | | | | more](ht | | | | | tps://openai.co | | | | | m/blog/new-embe | | | | | dding-models-an | | | | | d-api-updates). | | | +-----------------+-----------------+-----------------+-----------------+ | gpt- | Currently | 128,000 tokens | Up to Dec 2023 | | 4-turbo-preview | points to | | | | | `gpt-4 | | | | | -0125-preview`. | | | +-----------------+-----------------+-----------------+-----------------+ ...
- jerzyt 3y agoOf course, Wikipedia is the easiest website to scrape. The HTML is so clean an organized. I'd like to find some code to scrape Airbnb.
- 1vuio0pswjnm7 3y agoI would prefer if people would not submit sites to HN that are using anti-scraping tactics such as DataDome that block ordinary web users making a single HTTP request using a non-popular, smaller, simpler client. One example is www.reuters.com. It makes no sense because the site works fine without Javascript but Javascript is required as a result of the use of DataDome. See below for example demonstration. For anyone who is doing the scraping that causes these websites to use hacks like DataDome: Does your scraping solution get blocked by DataDome. I suspect many will answer no, indicating to me that DataDome is not effective at anything more than blocking non-popular clients. To be more specific, there seems to be a blurring of the line between blocking non-popular clients and preventing "scraping". If scraping can be accomplished with the gigantic, complex popular clients, then why block the smaller, simpler non-popular clients that make a single HTTP request. To browse www.reuters.com text-only in a gigantic, complex, popular browser 1. Clear all cookies 2. Allow Javascript in Settings for the site ct.captcha-delivery.com 3. Block Javascript for the site www.reuters.com 4. Block images for the site www.reuters.com First try browsing www.reuters.com with these settings. Two cookies will be stored. One from reuters.com. This one is the DataDome cookie. And another one from www.reuters.com. This second cookie can be deleted with no effect on browsing. NB. No ad blocker is needed. Then clear the cookies, remove the above settings and try browsing www.reuters.com with Javascript enabled for all sites and again without an ad blocker. This is what DataDome and Reuters ask web users to do: "Please enable JS and disable any ad blocker." Following this instruction from some anonymous web developer totally locks up the computer I am using. The user experience is unbearable. Whereas with the above settings I used for the demonstration, browsing and reading is fast.
- martin82 3y agoGood beginner tutorial and some good stuff in here, but the chances of scraping any site that is behind Cloudflare or AWS WAF (which is almost all interesting sites), are basically zero.