8 ms·
I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the prov
by JimtheCoder 3y ago
I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving.
I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all...
Even though I am pretty sure it is already included in the training dataset already...
I could be wrong, though...
- coffeebeqn 3y agoReddits community is passive aggressive, thinks it’s really smart, loves memes. I certainly hope OpenAI doesn’t view it at as some kind of a source of truth
- rchaud 3y agoYou just described the C-suite at most SV companies.
- pixl97 3y agoNo, they view it as a language model. You can tell lies with the same language that you can tell truth with. If you want truth, you don't want language, you want references to reviewed work. You also want things like 'show your work' chains of though. These are really different things. If I tell GPT "make up a story" I don't want it coming back and saying, sorry I can only tell the truth.
- moron4hire 3y agoSeriously hope OpenAI is stripping all the canned meme replies that go on for hundreds of sub threads and end up as the top comments in threads before they train their models. How would you even do that reliably?
- jononor 3y agoLabel data and build a meme classifier? Does not have to be perfect to be useful. But yeah, data curation is probably a huge endeavor at the companies making Language model that are fit for production. Like in practically all applications of Machine Learning. But the Reinforcement Learning from Human Feedback (RLHF) is also one of the key tools to getting useful outputs.
- renewiltord 3y agoThey don't. That's why some tokens from /r/counting mess up their models: SolidGoldMagikarp and " davidj12" or something like that
- throwuwu 3y agoYou’re not alone in thinking that. The value in using Reddit as training data would be the question response format of threaded comments. The downside is that the vast majority of comments on Reddit are very low quality and repetitive. You’d have to do a lot of filtering to make it usable and what you’d be left with would be a much smaller pile of training data.
- bzmrgonz 3y agorepetitive is good tho, it presents validation. I'm sure programming a good AI means adding a grading system for repetitive facts, in reddit's case, it may even accommodate likes. The problem would be when we have run-away sarcasm/irony/memes which the llamas can't handle.
- zirgs 3y agoRepetition is not good, because too much of it leads to overfitting.
- dmbche 3y agoI'm sure that the non technical public would be interested in a chatbot fed on reddit data, which is more interesting than how valid the AI models predictions are for the people making money off of it.
- boredumb 3y agoYou are absolutely right minority or not. I actively avoid reddit because the majority of users there. The fact we're seeing government agencies start treading into using GPT models is frightening. We could find ourselves in a tragic comedy where all of the massive institutions and enterprises around us are addressing their serious issues via redditors by proxy.
- deleted 3y ago[deleted]
- zuppy 3y agodepends of where you go on reddit. i've learned a lot of things for my hobbies. for example r/espresso and r/roasting are a source of good information. there are also places like r/askhistorians and many many otheres. reddit is not just r/funny.
- tayo42 3y agoniches and hobbies are dominated by beginners and ideas that can't be challenged (theres a word for this, I cant remember what it is). people aren't just talking about their experience on major subs though. the opinions i run into real life can be very different then with people in the real world. the communities online are made up of the kinds of people who spend their time online, and the content you see on reddit is generally from the people who spend enough time on reddit that they want to browse new. these arent average people. only a minority professionals are actually engaged in reddit. even those ive seen run off because they don't agree with the acceptable opinions
- cannonpalms 3y ago> there's a word for this Hegemony, perhaps.
- pixl97 3y ago>the opinions i run into real life can be very different then with people in the real world. I think you messed that sentence up, but I get what you're trying to say.... And I think it's this. The opinions you get from people 'IRL' are not apt to be as strong as the ones online, and or will run into the regency bias. For example, it's very unlikely you'll actually meet someone that has used 10 different coffee makers because they wanted to see which one was best. Online on some subreddit, you're very likely to meet someone who has done exactly that. Of course those people with strong opinions are the ones that are apt to post most online. So who's option is wrong? Neither. That's why they are opinions.
- dontreact 3y agoOne of the big values of reddit is that it serves as an easy index for non-garbage websites on the internet. This is exactly how it was used in the T5 paper: they threw away all the websites from their crawl that were not linked to from Reddit.
- dontreact 3y agoAnd by garbage I mean literally data/illegible/etc. that would ruin the pretraining of the model.
- zdragnar 3y agoI genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in the thread. Synthesizing multiple comments together requires nuance based on circumstances that I don't think I would trust to an automated process- which heuristics were applied, what was the reputation of the people on either side, etc. Hell, I don't even stop at the first recipe I find if I'm looking for something new to cook for dinner. I look at a couple variations on a dish first.
- ok123456 3y agoThe appeal is that now regular search engines are so bad at giving you useful content that using a LLM is now the equivalent of "google-dorking" to find relevant information.
- rchaud 3y agoOnly because the LLM didn't have any AI generated blogspam to get trained on. That's going to change very quickly.
- ok123456 3y agoAny non-trival LLM that works by scraping the internet is already sufficiently advanced to be able to classify blogspam.
- giantrobot 3y agoThe issue Google has with blogspam is they don't want to filter it because it inevitably uses Google's advertising. So Google gets to self-deal traffic to blogspam which makes them money. They're not incentivized to actually eliminate blogspam. LLMs aren't automatically disincentivized from training on blogspam so they're not going to avoid it either.
- alexghr 3y agoI think what the article tries to say is that OpenAI have already scraped Reddit for training data and with the recent API changes and subreddits going dark, new competitors in the AI space won't have it as easy to get the same training set.
- losteric 3y agoHonestly this sounds like a shower-thought post. With even basic research, Internet Archive and The Eye have Reddit historical data freely available. My desktop PC has all comments and posts from 2007-early 2023, in a convenient jsonl zst. It's only 3TB.
- boh 3y agoI don't think you're the minority. I want to actually read the thread and see what random individuals thought about a thing (with their different views intact), rather than getting an aggregate summary. Often the overlooked, 1 point comment (because he answered the question a month after the question was raised), is the one you're looking for.
- sebzim4500 3y agoThe point isn't really to discover the opinions of redditors, it is to ingest the 'common sense' things that you would never find out from reading scientific papers or even books.
- shaoner 3y agoIn my opinion, it comes back to search engines not giving you a the best answer. If you're looking for the best headset with whatever feature, you'll find mostly sponsored reviews, or at the very least hard to trust reviews. On the other hand, some random people's opinions in a Reddit community with -apparently- no further agenda seem somewhat more honest. Not that the answer is better but it gives you new data points in your search. Basically, it's not one or the other, you can use both tools and that's probably why it makes sense to include Reddit in AI models (which do this job for you automatically)