4 ms·
A problem with AI I don't hear anybody really talking about is how it is creating a massive incentive to start gathering more and more personal information. Mor
by oakpond 2y ago
A problem with AI I don't hear anybody really talking about is how it is creating a massive incentive to start gathering more and more personal information. More and more training data will be needed to make better and better AI products.
- luyu_wu 2y agoIf the data is ananomous (as it should in all telemetry), this is less of a concern. But yeah, it's definitely concerning to see more incentive for data collection beyond ad targeting.
- simonw 2y agoAndrej Karpathy: https://twitter.com/karpathy/status/1797313173449764933 https://twitter.com/karpathy/status/1797313173449764933 > Turns out that LLMs learn a lot better and faster from educational content as well. This is partly because the average Common Crawl article (internet pages) is not of very high value and distracts the training, packing in too much irrelevant information. The average webpage on the internet is so random and terrible it’s not even clear how prior LLMs learn anything at all. People keep assuming that AI companies want to pipe any old junk into their models, but I'm increasingly convinced that this isn't true.
- dylan604 2y agoSetting aside the copyright issues--just like the AI companies already have--it would be interesting to train a model only on text books of all ages. Would it be able to tell the edits/revisions/versions that have been made over time and categorize them from new learning vs changes in political winds?
- amlib 2y agoAnd all the drivel that google's own search AI has been spewing comes from where? Or the incredible breadth of, often wrong, knowledge that ChatGPT can spew out in seconds, where does that comes from? Maybe at some point they will try to create their own, very limited, "knowledge base" but then those AI assistants cease to be a jack of all trades, while still being masters of nothing.
- simonw 2y agoThe Google search stuff recently wasn't training data, it was RAG.
- tentacleuno 2y agoI've noticed you use the acronym "RAG" quite a lot -- I'd like to know what that means. I'm not yet familiar with AI terminology :)
- simonw 2y agoIt stands for Retrieval Augmented Generation. It is the trick where you answer questions that are not in the model's original training data by first running a search For documents relevant to their question, and then invisibly pasting there results into the model as part of the prompt along with their question. It's the trick used by ChatGPT Browsing mode and Bing and Perplexity and Google Gemini and the new Google Search.
- blibble 2y ago> People keep assuming that AI companies want to pipe any old junk into their models, but I'm increasingly convinced that this isn't true. meanwhile in dimension reality: Google's AI tells people to put glue on their pizza, the only source being a single reddit comment from user "fucksmith"
- simonw 2y agoThat's not a model training data thing though, it's a RAG thing. Functionally very different from using that text for training data.
- oakpond 2y agoEvidently they still need to experiment. They don't know if it's junk until they try to use it. I'm convinced they are going to try as much data as they can simply because they need to compete in the market.