4 ms·
Story time. I was at a VC conference last year and if I learned nothing else there, I learned how to spell "AI". Every single exhibitor just about had their s
by nyc_data_geek 2y ago
Story time.
I was at a VC conference last year and if I learned nothing else there, I learned how to spell "AI". Every single exhibitor just about had their signage proudly proclaiming their capabilities in this area, but one in particular struck me.
They were touting the API integrations they could offer to train their "Enterprise AI"/LLM, and among those integrations were things like M365, Slack, etc.
It struck me because of the garbage in, garbage out problem. I'd like to think that the amount of shitposting I do on Slack personally will poison that particular well of training data, but this seems to point to a larger problem to me.
LLM's don't have a concept of truth or reality, or awareness of any sort. If the training data they are fed is poorly quality checked/unsanitized by human intelligence, the outputs will be as useless/noisy as the original data set. It feels to me that in the frothy rush to capture market buzz and VC, this is being forgotten.
Am I missing something, here?
- beeboobaa3 2y agoMore ignored than forgotten.
- nyc_data_geek 2y agoWallpapered over?
- leoh 2y agoYes, consider an existing LLM being given “shitpost-y” messages and asking it if there is anything interesting in there. It could probably summarize it well and that could then be used for training another LLM. etc etc
- nyc_data_geek 2y agoThis assumes everything in the training data set is accurate. Sometimes people are wrong, obtuse, sarcastic, etc. LLM's don't have any way of detecting or accounting for this, do they? That output, then being used to train other LLM's, just creates an ouroboros of AI generated dogshit.
- sp332 2y agoLLMs are state-of-the-art at detecting sarcasm. It won't help if the data is just wrong though. Edit: https://arxiv.org/abs/2312.03706 https://arxiv.org/abs/2312.03706 Human performance on this benchmark (detecting sarcasm in Reddit comments) was 0.82, a BERT-based LLM scored 0.79. https://arxiv.org/abd/2106.05752 https://arxiv.org/abd/2106.05752 LSTM, 98% at detecting sarcasm in a Project Gutenburg-based dataset.
- E39M5S62 2y agoI literally can't tell if you're being sarcastic or not.
- rrr_oh_man 2y agoExactly
- comboy 2y ago> LLMs are state-of-the-art at detecting sarcasm. This is such a precious gem.
- brookst 2y agoAnd yet human civilization has survived the fact that many humans are wrong, lying, delusional, etc. There is no assumption that everything in our personal training set is accurate. In fact, things work better when we explicitly reject that idea. LLMs do not rely on 100% factually accurate inputs. Sure, you’d rather have less BS than more, but this is all statistics. Just like most people realize that flat earthers are nutty, LLMs can ingest falsehoods without reducing output quality (again, subject to statistics)
- nyc_data_geek 2y agoOutput quality is low enough that the term "hallucinations" has been coined. Statistics doesn't appear to fix this.
- 2y ago
- chatmasta 2y agoWhy shouldn’t AI be able to shitpost too? At the very least, and much more importantly, AI should be able to recognize shitposting.
- nyc_data_geek 2y agoThis is the crux of it, and where I'm wondering if I'm missing something. Can it, today? My understanding is it cannot discern reality from fiction, thus "hallucinations" (a misnomer because it implies awareness, which these probability models lack).
- williamcotton 2y agoThe sheer scale of data on the long tail. Sure, the head is already a trash pile and has been for decades now, but there is plenty of non-monetized information all over the internet that is barely linked to or otherwise discoverable.
- krainboltgreene 2y agoIt does not matter how hard they try, nothing will rival the CommonCrawl treasure trove except maybe Google's index itself.
- jorisboris 2y agoSame for Reddit or Facebook groups. There's a lot of shitposting there, but absolutely a lot of valuable information if LLMs manage to separate the wheat from the chaff.
- mvkel 2y agoThe best LLMs were trained on data from the open internet, which is full of garbage. They still do a pretty good job (granted it has been fine tuned and RLHF'd, but you can do that with Slack data too)
- antipaul 2y agoWhat do you think chatGPT uses as training data? The whole world’s “sh*tposting”: Reddit, blogs, and the rest of the internet. But also books and Wikipedia and what not. You can “smooth” all the crap out via the training procedure. But even more, Slack can easily filter training data to, say, only posts in high-use channels. Further, slack has other options: eg, use their customer data only for marginal fine-tuning, for example. Or, they don’t even know their use case yet - but want to wrap their arms around your data pronto.
- moneywoes 2y agohow does the training procedure smooth the garbage out?
- thomashop 2y agoThrough regularization techniques, data augmentation, loss functions, and gradient optimization, ensuring the model focuses on meaningful patterns and reduces overfitting to noise.
- bigfudge 2y agoIt’s not obvious how any of those would do anything but better approximate the average of a noisy dataset. RLHF might help, but only if it’s not done by idiots.
- rozap 2y agoWhat makes you think I don't shitpost in the #engineering channel? And heuristics don't even scratch the surface of the bigger problem where it's trained on people who aren't great at their jobs but type a lot of words on slack about circling back on KPIs.
- bee_rider 2y agoI think those types of people are actually shockingly well paid. If slack can make bots to replace them, they’ll print money, right?
- wongarsu 2y agoMost shitposting is probably more straightforward to understand than business communication or press releases where realizing what wasn't said often carries more insight than the things that were said. Of course training an AI model on simple, straightforward and honest data provides good results. That's the essence behind "textbooks is all you need" which lead to the phi LLMs. Those are great small LLMs. But if you want your model to understand the complexity of human communication you have to include it in your training data. If you subscribe to the idea that to be the very best text completion engine possible you would need to have a perfect understanding of reality itself, how different humans perceive reality differently, and how they choose to communicate about this perception and their interaction with reality, themselves and other humans, then it's not unreasonable to expect that back-propagation would eventually find that optimal representation if given enough data, the right architecture and enough processing power. Or at least come somewhat close. In that paradigm there is no "bad data", only insufficient or badly balanced datasets. Just don't try doing that with a 3B parameter LLM.
- swalsh 2y agoAlong the same lines, phi-3 is kind of a sign of what you can do if you focus only on high quality data. It seems like while yes, quantity is very important, quality almot matters just as much.
- bongodongobob 2y agoI think what you're missing is assuming that what an LLM "reads" thinks is a true statement. Shitposting is almost like meta slang. I feel like that's a necessary thing for it to train on to truly understand language. I feel like people underestimate the depth LLMs can pick up on.
- IanCal 2y agoThe more obvious things are that it's not training llms fully on all channels. Some quick ideas: Search and summarize other messages. No new llms training and about mostly linking to existing answers. Fine tune on your messages, but only customer support messages in the public channel, not "eng-shitpost" Natural language requests over your company data.
- deleted 2y ago[deleted]