3 ms·
This breaks with using standard web scraping methods (non-headless JS engines). Had to deal with this issue recently due to everything being a damn SPA now. Loo
by rabuse 5y ago
This breaks with using standard web scraping methods (non-headless JS engines). Had to deal with this issue recently due to everything being a damn SPA now. Look into Selenium for running headless browsers if looking to scrape the modern web.
- ricardo81 5y agoI used to use mozrepl back in the day before FF Quantum (or a version nearby) broke it, worked great- run the browser(s) via telnet. Could run as many browsers as memory would allow.
- kordlessagain 5y agoDepending on the use case you might try imaging the page, then send the image to an ML model for full text before indexing. If you need links extracted, Selenium also supports parsing the assembled DOM: https://github.com/kordless/grub-2.0/tree/main/aperture https://github.com/kordless/grub-2.0/tree/main/aperture
- bshipp 5y agoWhile I certainly don't disagree that many sites need selenium these days, I find it's often possible to go straight to the AJAX/xHR source and get a pre-parsed JSON with all the data you're looking for.