4 ms·
I've actually wrote about this! General tips that I've found from doing more than a few projects [0], and then an overview of Python libraries I use [1]. If yo
by jackschultz 9y ago
I've actually wrote about this! General tips that I've found from doing more than a few projects [0], and then an overview of Python libraries I use [1].
If you don't want to clock on the links, requests and BeautifulSoup / lxml is all you need 90% of the time. Throw gevent in there and you can get a lot of scraping done in not as much time as you think it would take.
And as long as we're talking about web scraping, I'm a huge fan of it. There's so much data out there that's not easily accessible and needs to be cleaned and organized. When running a learning algorithm, for example, a very hard part that isn't talked about a lot is getting the data before throwing it in a learning function or library. Of course, there the legal side of it if companies are not happy with people being able to scrape, but that's a different topic.
I'll keep going. The best way to learn about what are the best tools is to do a project on your own and teat them all out. Then you'll know what suits you. That's absolutely the best way to learn something about programming -- doing it instead of reading about it.
[0] https://bigishdata.com/2017/05/11/general-tips-for-web-scraping-with-python/ https://bigishdata.com/2017/05/11/general-tips-for-web-scrap...
[1] https://bigishdata.com/2017/06/06/web-scraping-with-python-part-two-library-overview-of-requests-urllib2-beautifulsoup-lxml-scrapy-and-more/ https://bigishdata.com/2017/06/06/web-scraping-with-python-p...
- Bromskloss 9y ago> BeautifulSoup / lxml When should one use one or the other, would you say?
- ivan_ah 9y agoYou can use the BeautifulSoup API with the `lxml` parser: https://www.crummy.com/software/BeautifulSoup/bs4/doc/#installing-a-parser https://www.crummy.com/software/BeautifulSoup/bs4/doc/#insta... I've heard that `lxml` can choke on certain badly-formed markup, but it's very fast. Personally has never failed on me.
- gros_roberts 9y agolxml.etree.HTMLParser(recover=True) should work for bad HTML. A few times I had to replace characters before giving the page to lxml, but it was more of an encoding issue.
- toutenrab 9y agoIt may not be related but I also noted that processing HTML with lxml (e.g. update every URL of a HTML document with a different domain for instance) was producing malformed HTML with duplicated tags. So I would recommend to use lxml only as a data extraction tool.
- nerdponx 9y agoLXML also is known to have memory leaks [0][1], so be careful using it in any kind of automated system that will be parsing lots of small documents. I personally encountered this issue, and actually caused to abandon a project until months later when I found the references I linked above. It works nice and fast for one-off tasks, though. Also, a question: how often do you really encounter badly-formed markup in the wild? How hard is it really to get HTML right? It seems pretty simple, just close tags and don't embed too much crazy stuff in CDATA. Yet I often read about how HTML parsers must be "permissive" while XML parsers don't need to be. I've never had a problem parsing bad markup; usually my issues have to do with text encoding (either being mangled directly or being correctly-encoded vestiges of a prior mangling) and the other usual problems associated with text data. [0]: https://benbernardblog.com/tracking-down-a-freaky-python-memory-leak-part-2/ https://benbernardblog.com/tracking-down-a-freaky-python-mem... [1]: https://stackoverflow.com/q/5260261 https://stackoverflow.com/q/5260261
- jackschultz 9y agoBeautifulSoup. The difference is that lxml can run a little faster in certain cases for a huge scrape, but you'll very very very if ever need that. It's interesting and probably worthwhile to try both and know the difference, but bs BeautifulSoup is definitely where to start
- darpa_escapee 9y agoBeautifulSoup has a friendly API, but it is slow. It has a lxml backend, however. If you're familiar with writing XPath queries, lxml is great.
- j_s 9y agoUse https://github.com/kovidgoyal/html5-parser https://github.com/kovidgoyal/html5-parser, which (in my limited understanding) does a better job faster and is backwards-compatible with both. Recommendation by the author (of Calibre fame) on a similar discussion: https://news.ycombinator.com/item?id=15539853 https://news.ycombinator.com/item?id=15539853 Dedicated discussion: https://news.ycombinator.com/item?id=14588333 https://news.ycombinator.com/item?id=14588333