Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dataslap
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
dataslap
9y ago
21 Gigs in total for everything, ignoring errors (files too big / broken links)
2.
▲
by
dataslap
9y ago
I'm really happy their share it though. The collection is just amazing - it's a good idea for "what books to buy for a personal library when you start making decent money". Definitely would spend money on things like thi
3.
▲
by
dataslap
9y ago
I'm at 300 right now, 9.5 gigs. I think some are skipped cause the download times out (can redownload manually I suppose)
4.
▲
by
dataslap
9y ago
P.S. it will take hours or days =))) <3 <3 <3
5.
▲
by
dataslap
9y ago
Please use responsibly. Recommended to delay by a few days - it seems like there is an initial surge going on right now. Increase download delay to 20 sec - 1 minute. Be respectful. https://github.com/ilyaperepelitsa/me
6.
▲
by
dataslap
9y ago
I think someone else is messing with it - got slower for me too.
7.
▲
by
dataslap
9y ago
It's slow cause things are quite large there (just saw a 2 gig book). I'm using delays, don't worry - hence "in a few hours" <3 <3 <3
8.
▲
by
dataslap
9y ago
<3
9.
▲
by
dataslap
9y ago
Uploading a scrapy crawler that downloads PDF books to github in a few hours, gonna post the link here
10.
▲
by
dataslap
9y ago
some "ethical" measures may do the trick to. scrapy has a setting to integrate delays + you can use fake headers. Some sites are pretty persistent with their cookies (include cookies in requests). It's all case by case basis
11.
▲
by
dataslap
9y ago
depends on the task. For example they have a decent file/image downloading middleware.
12.
▲
by
dataslap
9y ago
scrapy has a pretty decent parser too