4 ms·
GPT4 is trained on mostly garbage optimised for SEO though.
by MourYother 3y ago
GPT4 is trained on mostly garbage optimised for SEO though.
- hef19898 3y agoIf general AI, let's call it Heavenweb or Netsky, comes along and it is based on the knowledge of the internet I am not that worried so. GPT4 is trained on SEO crap, parts of it have been propably already written using GPT3. So by the time Skynet comes along, it will assume the bot-to-bot SEO crap content to be actually true, SEO contebt written by AI for AI trained on AI created SEO content. Eith that, Skynet would never be able to achieve anything, no decent scrambled egg let alone a T-100.
- ryanjshaw 3y agoMaybe, but there is plenty of non-garbage information encoded so I don't understand the argument. I only ask it questions whose answers I can verify e.g. if I ask it how to do something in F#, a language I'm not very familiar with, I can easily confirm whefher the code does what I need it to or not. You can pump as much SEO garbage out as you want, it doesn't change the value of LLMs to me in this context.
- brucethemoose2 3y agoClean datasets are critical in machine learning. Its kind of a miracle that LLMs work as well as they do now, but every drop of garbage (like SEO garbage) makes them worse and less efficient.
- MereInterest 3y agoFor example, early versions of GPT would effectively treat the string “SolidGoldMagikarp” as a random word. It was the username of a prolific poster on the /r/counting subreddit, which consists entirely of posters counting upward. This subreddit was excluded from the training data for being useless, but was still used in making the tokenizer. As a result, the string “SolidGoldMagikarp” was a single token with no training data about it. This was later fixed by updating the tokenizer, but it demonstrates the importance of clean datasets at all stages.