4 ms·
It's actually pretty trivial to speed up, if you have multiple documents to parse, you can use multiprocessing. from multiprocessing import Pool d
by dmayle 4y ago
It's actually pretty trivial to speed up, if you have multiple documents to parse, you can use multiprocessing.
from multiprocessing import Pool
def parse(html):
result = []
soup = BeautifulSoup(html, 'html.parser')
for p in soup.select('div > p'):
result.append(p.text)
return result
with Pool(processes=16) as pool:
for texts in pool.imap_unordered(parse, my_html_texts):
for text in texts:
print(text)
- djtriptych 4y agoit's a pretty common case to be parsing thousands of locally-stored pages. There are only so many cores and the task of scraping an entire site can still easily still be CPU limited.
- driscoll42 4y agoYou might want to try a different parser as well, I tried a basic performance test comparison in another comment (https://gist.github.com/MercuryRising/4061368 https://gist.github.com/MercuryRising/4061368) and the html.parser was very slow compared to the lxml-xml or xml parsers for bs4 ==== Total trials: 100000 ===== bs4 lxml total time: 110.9 bs4 html.parser total time: 87.6 bs4 lxml-xml total time: 0.5 bs4 xml total time: 0.5 bs4 html5lib total time: 103.6 pq total time: 8.7 lxml (cssselect) total time: 8.8 lxml (xpath) total time: 5.6 regex total time: 13.8 (doesn't find all p)
- aumerle 4y agoIf you want fast HTML parsing in python+lxml, use html5-parser https://github.com/kovidgoyal/html5-parser https://github.com/kovidgoyal/html5-parser