3 ms·
I am always reminded of Severance, the TV show, when I think of LLMs. Data refinement is what is happening massively right now. Basically, take data harvested o
by EternalFury 2y ago
I am always reminded of Severance, the TV show, when I think of LLMs. Data refinement is what is happening massively right now. Basically, take data harvested on the Internet, which may be of poor quality, and use a teacher-student approach to reformulate it in higher quality. Then use that synthetic data to train new models. It's a bit like the biblical multiplication of the bread loaves. How much farther can it get LLMs? I don't know.
During the next turn, though, a lot of the same synthetic goop is going to turn up on the Internet. That's probably why OpenAI is spending the big bucks to secure access to original content producers. Now, here is to hope "original content producers" are not going to depend on LLMs to generate that original content.
I think training data would be plentiful if autonomous drones roamed the face of the earth, constantly recording audio and video. There is nothing as good as reality to learn what real is. The representation of reality on the Internet has become a matter of opinion for a lot of people.
- Chris_Newton 2y agoThere can be all kinds of strange and possibly perverse incentives at work for the “original content producers”, too. Forums like LinkedIn have had regular listings in recent times looking for programmers to help train new models for answering programming questions and generating code. However, the rates offered are nowhere near enough to compete with real programming jobs for skilled developers, at least not in the major Western economies. One wonders what quality level these new models will actually be trained at. Given that these roles are kinda-sorta asking professional programmers to train themselves out of a job, one also wonders whether 100% of those who do contribute will be doing so in good faith. And as you point out yourself, there is also a risk that the “original” content is anything but. Given the scale of the data involved here, and the sheer amount of skilled human input that would be required to refine that data set manually, it seems unlikely that a Mechanical Turk strategy alone will solve many fundamental problems here.