3 ms·
Thanks! I mentioned BeautifulSoup existed very briefly in the lxml section in prerequisites. I just added a short section on CSS Selectors, you can actually use
by Jake232 13y ago
Thanks! I mentioned BeautifulSoup existed very briefly in the lxml section in prerequisites. I just added a short section on CSS Selectors, you can actually use lxml with them. I'll add a little more info with BeautifulSoup and PyQuery mentioned though.
- shuzchen 13y agoPeople who write (python) crawlers use BeautifulSoup because it was created with invalid input in mind. If you're crawling the wild wild internet, chances are you'll find a wide range of invalid markup and lxml barfs on those (it's got recover=True, but that's not supported for all things lxml does).
- Jake232 13y agoI've wrote a lot of python crawlers, and have never really experienced an issue with lxml. It doesn't really ever seem to choke, and I'm sure I must have come across some pretty funky markup in my time. lxml has always handled things fine for me. Maybe I've just gotten lucky I guess?
- aGHz 13y agoThat's not actually lxml's fault, it depends on the libxml2 installed in your environment. Just like Jake232, I haven't found a site that lxml can't handle lately, using libxml 2.9.1. You can find your libxml2 version in Python with: >>> from lxml import etree >>> etree.LIBXML_VERSION If you find a page in the wild that it can't handle, I'd love to know the URL.
- dikei 13y agoEven BeautifulSoup makes use of lxml now (as of version 4)