3 ms·
I recently learned (by trial and error) some of the headaches associated with running headless browsers at scale that Joel mentions here; wish I'd heard of this
by iamEAP 7y ago
I recently learned (by trial and error) some of the headaches associated with running headless browsers at scale that Joel mentions here; wish I'd heard of this service earlier. I ended up finding other solutions to fill in the gaps: Puppeteer Cluster is one I'd recommend (https://github.com/thomasdondorf/puppeteer-cluster https://github.com/thomasdondorf/puppeteer-cluster)
I especially like the "host it yourself" commercial license model, here; while automating browser _actions_ over a network works well enough, _detailed scraping_ over a network can quickly become inefficient (as many requests for elements or element attributes may incur individual round-trips). In some cases, colocating your browser instance with your scraping logic becomes a necessity.
- mrskitch 7y agoWe hear about puppeteer-cluster _a lot_, and we hear the same thing from folks (that's it's great). browserless.io essentially does "clustering" at an infrastructure level, whereas puppeteer-cluster does it at the application level. Both essentially solve the same problem, just in different ways.