4 ms·
Why kind of data that isn’t public would be so valuable for AI training? Seems like there’s a fuck ton. All of Wikipedia, GitHub for code, etc. I can underst
by DrFalkyn 2y ago
Why kind of data that isn’t public would be so valuable for AI training?
Seems like there’s a fuck ton. All of Wikipedia, GitHub for code, etc.
I can understand targeting certain sites like Reddit, etc. but not random websites
- timewizard 2y agoIt's to rip off copyrighted content and profit from it instead of the original authors. It's like every other low rent and highly automated scam that finds it's way onto the internet. If you look closely even Google does this. This is probably why many popular sites started getting down ranked in the last 2 years. Now they're below the fold and Google can present their content as their own through the AI box.
- throwaway2037 2y agoPlease remember that Google only needs to be marginally better than the competition. And, of course, their primary biz is ads, not serving great results; that is a distant second priority.
- MathMonkeyMan 2y agoTheir biz is ads, but since search is winner takes all they need only be marginally better than the competition... twenty years ago.
- timewizard 2y ago> Their biz is ads, Yea, but, the FTC doesn't want it to be.
- grotorea 2y agoDiscord I guess would be quite valuable, even the de facto public servers.