4 ms·
POSIX one-liner to download all PDFs: $ curl -s https://www.myharvardclassics.com/categories/20120212 \ | grep "downloads.*download" | sed -e 's/href/\n&
by schizoidboy 8y ago
POSIX one-liner to download all PDFs:
$ curl -s https://www.myharvardclassics.com/categories/20120212 \
| grep "downloads.*download" | sed -e 's/href/\n&/g' \
| sed 's/.*\(http:.*download\)/\1/g' | grep " - " \
| grep -v target | sed 's/"> - / /g' | sed 's/<.*//g' \
| while read link rest; do \
if [ "${i}" = "" ]; then i=1; fi; \
wget -O "Harvard Classics - Volume $(printf '%02d' ${i}) - ${rest}.pdf" "${link}"; \
i=$(( i + 1 )); \
done
- Kovah 8y agoI should have checked the comments before downloading all books manually...
- qop 8y agoExcellent form! I would LOVE to see more of this type of posting in HN. We're supposed to hack everything! There was a really neat post the other day I think you might like about how someone made a shell script that performed a job like 200x quicker than hadoop. You'd probably enjoy it. Anyways, thanks!
- glitch 8y agoFor anyone that wants to download the books, I strongly recommend downloading the PDF archive on Archive.org because the image scans within the PDFs are of much higher quality. https://archive.org/details/Harvard-Classics https://archive.org/details/Harvard-Classics
- schizoidboy 8y agoNice. The `du -ch *pdf` for myharvardclassics.com is 439M, and this higher quality zip is 975MB. $ wget "https://archive.org/compress/Harvard-Classics/formats=TEXT%20PDF&file=/Harvard-Classics.zip"
- deleted 8y ago[deleted]