4 ms·
> because they are trained on multilingual data But they were not trained on government-sanctioned homegrown EU data.
by melvinmelih 11mo ago
> because they are trained on multilingual data
But they were not trained on government-sanctioned homegrown EU data.
- saretup 11mo agoThe entirety of the internet vs government-sanctioned homegrown EU data.
- raverbashing 11mo ago> But they were not trained on government-sanctioned homegrown EU data. If none of the LLM makers used the very big corpus of EU multilingual data I have an EU regulation bridge to sell it to you
- tonyhart7 11mo ago"But they were not trained on government-sanctioned homegrown EU data." ok what are you implying on this
- mock-possum 11mo agoSidesteps potential legal issues probably
- sunaookami 11mo agoWho in their right mind would use this?
- tensor 11mo agoI'd use a model trained on a targeted and curated data set over one trained on all the crap on the internet any day.
- sunaookami 11mo agoWe are talking about government-curated data here, the bias should be obvious. Popular LLMs still have huge bias problems but it would be way worse with only government-curated data.
- loandbehold 11mo agoI keep hearing that LLMs are trained on "Internet crap" but is it true? For instance we know from Anthropic copyright case that they scanned millions of books to make a training set. They certainly use Internet content for training but I'm sure it's curated to a large degree. They don't just scrap random pages and feed into LLM.
- nutjob2 11mo ago> I'm sure it's curated to a large degree. They don't just scrap random pages and feed into LLM. How would they curate it on that scale? Does page ranking (popularity) produce interesting pages for this purpose? I'm skeptical.
- airspresso 11mo ago> I keep hearing that LLMs are trained on "Internet crap" but is it true? Karpathy repeated this in a recent interview [0], that if you'd look at random samples in the pretraining set you'd mostly see a lot of garbage text. And that it's very surprising it works at all. The labs have focused a lot more on finetuning (posttraining) and RL lately, and from my understanding that's where all the desirable properties of an LLM are trained into it. Pretraining just teaches the LLM the semantic relations it needs as the foundation for finetuning to work. [0]: https://www.dwarkesh.com/p/andrej-karpathy https://www.dwarkesh.com/p/andrej-karpathy
- ACCount37 11mo ago"Just" is the wrong way to put it. Pretraining teaches LLMs everything. SFT and RL is about putting that "everything" into useful configurations and gluing it together so that it works better.
- ACCount37 11mo agoIt is true. Datasets are somewhat cleaned, but only somewhat. When you have terabytes worth of text, there's only so much cleaning you can do economically.