3 ms·
Mostly using linux tools. First of all, always respect the site's terms of use and the server's capacity. It is tempting to just do parallel -P 200 but it can
by assafmo 9y ago
Mostly using linux tools.
First of all, always respect the site's terms of use and the server's capacity. It is tempting to just do parallel -P 200 but it can hurt the website and can lead to you being throttled/baned.
cURL. Chrome dev tools -> Network -> Copy as cURL. also very easy to customize requests.
sessions can be done using --cookie and --cookie-jar.
Using tor (sudo apt install tor) with --socks5 localhost:9050 or --socks5-hostname localhost:9050 is great for scraping .onion websites and for anonymity.
The holy grail for me is to find a JSON REST API for the website I am scraping.
lynx --stdin --dump to extract text if there isn't any interesting stuff inside the html/js (man lynx! it has some great options)
awk to extract data. very easy to extract data from tables! to extract data from html I usually use awk -F '[<>]' or awk -F '"'... depends on where the data is in the html. (also sometimes egrep -o to extract with a regex)
seq 1 10 | parallel curl https://banana.papaya.com/?page={} https://banana.papaya.com/?page={} is great for paging.
jq to process JSONs. I usually output to csv with @csv, because then I can use sqlite for quick queries or any other DB that has an easy "import from csv" option (I like using CouchDB with couchimport).
Dynamic pages - You can always get the content using the same request the that site's javascript did. I never find myself needing to execute the page's javascript.
EDIT: typos, grammer
- tmaly 9y agoThanks for the Chrome dev tools tip, I did not even know they had that. lynx brings back memories when all I had was a green screen vt100 terminal in the university library.