5 ms·
Why? What is the goal of a scraper, and how does disabling the source of the data benefit them?
by MathMonkeyMan 2y ago
Why? What is the goal of a scraper, and how does disabling the source of the data benefit them?
- TuxMark5 2y agoI guess one could make a point that competition will no longer have the access to the scraped data.
- randmeerkat 2y ago> Why? What is the goal of a scraper, and how does disabling the source of the data benefit them? The next scraper doesn’t get the data. People don’t realize we’re not compute limited for ai, we’re data limited. What we’re watching is the “data war”.
- cyanydeez 2y agoat this point we're _good data_ limited, which has little to do with scraping.
- XorNot 2y agoHonestly it's hard to tell how much more value the LLM people are going to get out of another copy of the internet. It feels a lot like they're stuck for improvements but management doesn't want to hear it.
- Davidzheng 2y agoIt's a bit strange to talk about stuck when the most recent breakthrough is less than a year old.
- LPisGood 2y agoI’m not sure what you mean by breakthrough, but if you’re talking about Deepseek, it’s more of an incremental improvement than a breakthrough.
- DrFalkyn 2y agoWhy kind of data that isn’t public would be so valuable for AI training? Seems like there’s a fuck ton. All of Wikipedia, GitHub for code, etc. I can understand targeting certain sites like Reddit, etc. but not random websites
- timewizard 2y agoIt's to rip off copyrighted content and profit from it instead of the original authors. It's like every other low rent and highly automated scam that finds it's way onto the internet. If you look closely even Google does this. This is probably why many popular sites started getting down ranked in the last 2 years. Now they're below the fold and Google can present their content as their own through the AI box.
- throwaway2037 2y agoPlease remember that Google only needs to be marginally better than the competition. And, of course, their primary biz is ads, not serving great results; that is a distant second priority.
- MathMonkeyMan 2y agoTheir biz is ads, but since search is winner takes all they need only be marginally better than the competition... twenty years ago.
- timewizard 2y ago> Their biz is ads, Yea, but, the FTC doesn't want it to be.
- grotorea 2y agoDiscord I guess would be quite valuable, even the de facto public servers.
- threatofrain 2y agoScraping social media is good data, even without ML. The fact that something is "happening" to people in a social space inherently has importance to people. The specter of law is more threatening to whether companies can get their hands on good data.
- DaSHacka 2y agoNow the only way to obtain that information is through them