4 ms·
Wait, what? Aren't all of the things that Altman is doing now making all of the data that Reddit is sitting on (and its capacity to generate new ones) EVEN MORE
by ChicagoBoy11 3y ago
Wait, what? Aren't all of the things that Altman is doing now making all of the data that Reddit is sitting on (and its capacity to generate new ones) EVEN MORE valuable? Didn't they just announce some data sharing deal to train an LLM worth many millions not too long ago?
- SilverBirch 3y agoSo this is the story that places like Reddit and Quora would like to tell. They'd like to say that they have this treasure trove of data which can be used to train the models of the future, but that doesn't seem to be what most experts actually beleive. Firstly, Reddit is closing the stable door after the horse has bolted- they were open enough for long enough for most serious players to have most of that data. Second, LLMs are poisoning the well. Reddit's content isn't going to be good training data if it is itself the output of LLMs and that is going to happen. But finally and most importantly, it's not good! The open internet isn't great to train LLMs. A really important basic element of creating good models is training them on good data, and data from reddit as a source is terrible! Reddit as a source is about on par with just scraping the open web. If training data does become valuable you should be looking at places like the New York Times or scientific journals - places that have large volumes of quality training data.
- Workaccount2 3y agoWhile I am sure an LLM only trained on publications would be great, I can also picture it being the most annoyingly square AI to interact with. Just generally being painfully out of touch with how average humans speak and interact. I guess the ideal would be training with some kind of source discriminator or internal source quality parser that allows an LLM to be the chill funny guy who can still school your nerdy ass in distributed network architecture.
- peddling-brink 3y agoTwo LLMs, one pedantic, annoying, and correct. The other a master of speech and creativity. The first explains the answer, the second translates the answer into layman, and the first validates that the translation is still mostly correct.
- ChicagoBoy11 3y agoSincere question: To the extent that "the open internet" IS valuable data, isn't Reddit in the best possible position to capitalize on it? The mechanic of the website allows for people all over the world to contribute free-form content and then has mechanics built in where the content provided is graded by free human reviewers and moderators, and graded accordingly. Not saying it is perfect data, but surely there is not better alternative to it, no? I just don't understand how this data is "terrible" -- other than very limited things like newspapers or other very highly curated (and limited sources), what other large scale data repos are better?
- J_Shelby_J 3y agoI’m in a weird position of both valuing Reddit highly as a source of unique and niche information you literally can’t find any where else, while at the same time feeling that the general quality of discourse in the more mainstream subreddits has taken a nose dive and is now pretty much worthless.
- SilverBirch 3y agoThe problem is the grading is trash. In some sub-reddits the grading is literally inversely proportional to the quality of the content. In others its inversely proportional to traditional grammar rules. In some places it's highly correlated with misogyny. Like how often do you want your LLM to slowly transition it's response to a question into the Fresh Prince of Bel Air meme?[1] Compare it to other forms of content - news sources and journals which are extremely well curated, digitized books - a massive back catalogue of well edited and well structured data, Wikipedia - extremely high quality. I would put it this way, training an LLM on reddit is like teaching a child to play piano entirely by exposing them to freeform jazz. [1]: https://knowyourmeme.com/memes/bel-air-fresh-prince https://knowyourmeme.com/memes/bel-air-fresh-prince
- mooreds 3y agoDeal was with Google: https://www.reuters.com/technology/reddit-ai-content-licensing-deal-with-google-sources-say-2024-02-22/ https://www.reuters.com/technology/reddit-ai-content-licensi...
- reacharavindh 3y agoWait, aren’t there enough scraped dumps of Reddit data already? Sure, it is not as convenient as an API, and you may not be able to put Reddit’s logo next to the results, but why would LLMs care about that as long as they can train on the data in some form?