3 ms·
Likely the case for established model makers, but barring illegal use of outputs from other companies' models, a "first generation" model would still need this
by loufe 1y ago
Likely the case for established model makers, but barring illegal use of outputs from other companies' models, a "first generation" model would still need this as a basis, no?
- eru 1y agoWhy illegal? The more open models (or at least open-weight models) should allow using their outputs. Details depend on license. But yes, 'first generation' models would be trained on human text almost by definition. My comment was only to contradict the claim that 'all LLMs' are trained from stolen text, by noting that some LLMs aren't trained (directly) on human text at all.