22 ms·
I've been using it for quick web scrapping scripts and it's really nice.
by timdeve 4y ago
I've been using it for quick web scrapping scripts and it's really nice.
- nerpderp82 4y agoWhat libraries do you use? I do most of my scraping in Python using beautifulsoup.
- aeonik 4y agoNot OP but I use Reaver with good results. It supports all of JSoup's selectors, and makes it very clean to extract data from HTML. The documentation is a little lacking though, I had to look up other examples on GitHub to figure out how to use all the features. https://github.com/mischov/reaver https://github.com/mischov/reaver
- Borkdude 4y agoBabashka doesn't have a built-in HTML parsing library but it supports it through pods: https://github.com/babashka/pod-registry https://github.com/babashka/pod-registry Pods can be written in any language and they can expose functions to babashka by implementing a protocol. One pod exposing HTML parsing is: https://github.com/retrogradeorbit/bootleg https://github.com/retrogradeorbit/bootleg Here is an example of how to use that: https://github.com/babashka/pod-registry/blob/master/examples/bootleg.clj https://github.com/babashka/pod-registry/blob/master/example...
- noblepayne 4y agoAs mentioned by the one and only Borkdude, bootleg is a nice option for this. It includes the Hickory library: https://github.com/clj-commons/hickory https://github.com/clj-commons/hickory I'm a previous BeautifulSoup user and have found the combination of (1) having the scraped data presented in plain Clojure data structures, and (2) Hickory's built in selectors, to be a very nice experience. Happy scraping!
- nathell 4y agoI plan to port my scraping framework (Skyscraper, https://github.com/nathell/skyscraper https://github.com/nathell/skyscraper) to babashka one day. I’m not sure how easy it will be, though, since it uses core.async (which I believe bb has limited support for) and SQLite via clojure.java.jdbc.
- kot-behemoth 4y agoCore.async is listed as “built-in” in the Babashka Toolbox (https://babashka.org/toolbox https://babashka.org/toolbox). Might be worth checking if the compatibility has improved. And bb supports Honey SQL and SQL pods (https://github.com/babashka/babashka-sql-pods https://github.com/babashka/babashka-sql-pods) - so you might be already compatible!
- timdeve 4y agoAs other people have said Bootleg + Hickory. Here is an, admitedly not very clean, example that grabs stream urls from hltv.org: https://github.com/TimDeve/.dotfiles/blob/master/scripts/general/hltv https://github.com/TimDeve/.dotfiles/blob/master/scripts/gen... Also a basic RSS reader using the clojure XML lib: https://github.com/TimDeve/.dotfiles/blob/master/scripts/general/read-that https://github.com/TimDeve/.dotfiles/blob/master/scripts/gen...
- nerpderp82 4y agoThanks everyone, all great pointers. I'll do my next scrape in Babashka!