4 ms·
Would not recommend BeautifulSoup for this type of thing. lxml.html is much better in my experience. If you want to use CSS selectors, there's pyquery.
by halflings 9y ago
Would not recommend BeautifulSoup for this type of thing.
lxml.html is much better in my experience. If you want to use CSS selectors, there's pyquery.
- staticautomatic 9y agoI prefer lxml.etree even for HTML on account of the parser. Either way, I don't understand what's not "for humans" about lxml. It provides a ton of simple and useful abstractions, is easy to learn, and is crazy robust and fast under the hood.
- davidwtbuxton 9y agoYou can also use css selectors with lxml. Works great, in my experience. http://lxml.de/cssselect.html http://lxml.de/cssselect.html
- masklinn 9y ago> lxml.html is much better in my experience. lxml.html has a terrible parser. > If you want to use CSS selectors, there's pyquery. CSS selectors are built into lxml through cssselect[0] which is used to convert CSS3 to XPath 1.0 selectors. [0] If you want to use CSS selectors, there's pyquery.
- EmilStenstrom 9y agoLxml is C so annoying to install in some environments. Also, the API is inconsistent and verbose. I’ve found more luck with html5lib and cssselect2.
- Vindicis 9y agoI remember years ago I needed to parse some html(which was about 2-3 million characters) and after a fair bit of time, I had it up and running with beautifulsoup. Now, my use case was likely quite atypical to most html parsing, but my god was it ever slow! I forget the exact numbers, but I think it was taking about 150 seconds to complete. So, then I wrote it using lxml, which was an improvement, but that was still taking around 100 seconds. Now, I very rarely have any need to scrape and parse html data, and I was scratching my head at how it was taking these parsers so long to parse a 3.5 mib html page. I mean, it should be able to go through that and get what I want in under a second right? So, I said screw it and wrote some regexes. 10-15 seconds was how long it was now taking to parse that html. It actually took 1.5 or so seconds to parse the html; the rest was waiting for it to download the webpage. Ironically, implementing the regexes was actually quicker than figuring out how to use those html parsers and write the code. Of course, that's assuming you know how to craft regexes. Since it was set-up to run every 5 minutes, I wanted something that could do it without spending 1/2 the time parsing the data(amongst other tasks the processors were needed for). YMMV
- jcadam 9y agoYou're using regex to parse html? Have you not read: https://stackoverflow.com/a/1732454/1090568 https://stackoverflow.com/a/1732454/1090568 ?
- wodenokoto 9y agoIf you have an HTML document you want to extract information from, regexes are fast and easy. It's when you don't have guarantees about the structure of the html you are working with, that regex will come up short.
- Giroflex 9y agoI actually had the same experience. I was scraping a large number of pages and upon profiling my script, I found out that bs4 was really slow. Changing the parser from the default to lxml helped things a bit, but I decided I would just try a regex to check quickly whether things could be better. Lo and behold, it was much faster. It's true that it's impossible to parse HTML in its entirety with regex, but if you're looking to extract only a portion of data from a page with a known structure, a bit of regex might be the way to go.