3 ms·
>If Reddit merely wanted to restrict the ability to scrape its data, they could have done so without killing off clients – e.g. via licensing deals[1]. They ha
by winddude 3y ago
>If Reddit merely wanted to restrict the ability to scrape its data, they could have done so without killing off clients – e.g. via licensing deals[1].
They haven't taken any steps to stop scrapping. They made access to the data via api extremely expensive. Calling an API is not crawling/scrapping. You can still crawl/scrape, well once/if the mod protest is over. And also stop calling it scrapping, it web crawling.
I also actually wonder about the validity of it as training data. I've done a few experiments with fine tuning models, with a few hundred thousand samples curated from hundreds of millions of threads <https://huggingface.co/winddude/pb_lora_7b_v0.1 https://huggingface.co/winddude/pb_lora_7b_v0.1>. they are interesting, because they end up being so sarcastic. People tend to either be short, or overly opinionated.
At the very least someone would have to do a lot of pre-processing, which would make it a transfomative work anyways.