2 ms·
Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy. On the other hand, the Internet Archive is a non-profit offering a free
by ronsor 19d ago
Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy.
On the other hand, the Internet Archive is a non-profit offering a free public resource.
- toomuchtodo 19d agoExamples provided as technical examples, strong feelings are out of scope for this thread.
- itintheory 19d agoAs someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid. The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons. [0] https://people.kernel.org/monsieuricon/creepy-crawlies https://people.kernel.org/monsieuricon/creepy-crawlies
- toomuchtodo 19d agoNo strong feelings here is what I meant. Certainly, that energy is best directed into aggressive countermeasures and defense in depth of public goods. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=author%3Adang%20%E2%80%9Cstrong%20feelings%E2%80%9D&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- userbinator 19d agoI say put your data up in torrents, host a few KB of plain HTML linking to them, and let decentralisation do the rest.
- itintheory 18d agoThe data IS available. You can download it all from several sources in one big dump. And yet we're still scraped.
- userbinator 18d agoFurther evidence that they're not actually going after your data, but just DDoS'ing.