5 ms·
Author here -- that's a very good point and as I understand work in progress in different teams. Training autoencoders for language is actually super easy given
by faabian 2y ago
Author here -- that's a very good point and as I understand work in progress in different teams. Training autoencoders for language is actually super easy given the small amount of information contained in text (compared to vision/video), the hard part is making the model focus on the semantic part if all signal we have comes from exact match in token space. Hence Yann LeCun's ideas on joint embedding predictive architectures.
Note also that there is always a trade-off between auxiliary tasks giving more signal but shifting the focus. In our case, we noticed degradation if the number of predicted tokens is too high. So latent prediction methods need to sort out what is useful.
- mike_hearn 2y agoAren't the models already doing this, in a way? We know they can do things like write rhyming poems and song lyrics that do make perfect sense, so at some point the activations must be encoding some sort of overall plan for the upcoming sentences, even if maybe every word isn't predicted yet.
- faabian 2y agoYes. Otherwise next-token models wouldn't be nearly as good as they are. But the question is how to train these capabilities most efficiently! We had some interesting findings on how with increasing model/dataset scale/data quality, capabilities can move from "only learnable with multi-token prediction" to "indifferent" and "multi-token prediction actually hurts". This depends on the capability itself, induction e.g. matures way earlier in this sense than code generation capabilities.
- mike_hearn 2y agoIs it possible that anti-scaling effect occurs because you are removing some middle layers to free up space for the extra output heads? I only scanned the paper quickly but what happens if you treat the technique as strictly additive and don't keep parameter sizes fixed?
- mjburgess 2y ago> so at some point the activations must be encoding some sort of overall plan for the upcoming sentences This isn't obviously the case, compare this "intelligent designer" view with evolution: there was no prior plan for rabbits. it's sufficient to create the appearance of design that sequential steps are simply probabilistically modulated by prior ones. Consider a continuation of "the cat..." merely a distribution over all possible words suffices to create the illusion of a plan, suppose: "the cat sat..." then, "on.., the..." etc. follow from the training data. I think there's a strong argument against trying to model entire sentences exactly because the system isn't modelling semantics: one should expect accuracy to drop off a cliff if there is no actual plan. ie., predicting "sat on the mat" from "cat" shouldnt be a valid prediction, because of the infinite number of possible continuations that as a whole is terrible (eg., what about "chased the mouse" etc.). The space of all possible sentences to continue from "the cat" is infinite, which much of that space actually useful; whereas the number of words is very small, very fininte, and many of them not useful. The only reason that "the cat sat..", "the cat sat on..." is reasonable is because each sequential word can be modulated by the prompt to seem as if planned.
- edmara 2y agoThe modelling is advanced enough that you can't fundamentally distinguish it from (lossy, limited) planning in the way you're describing. If the KQV doesn't encode information about likely future token sequences then a transformer empirically couldn't outperform Markov text generators.
- mjburgess 2y agoNo one is spending $10-50mil building a markov text model of everything ever digitised; if they did so, their performance would approach a basic LLM. Though, more simply, you can just take any LLM and rephrase it as a markov model. All algorithms which model conditional probability are equivalent; you can even unpack a NN as a kNN model or a decision tree. They all model 'planning' in the same way: P(C|A, B) is a 'plan' for C following A, B. There is no model of P("A B C" | "A B"). Literally, at inference time, no computation whatsoever is performed to anticipate any future prediction -- this follows both trivially form the mathematical formalism (which no one seems to want to understand); or you can also see this empirically: inference time is constant regardless of prompt/continuation. The reason 'the cat sat...' is completed by 'on the mat' is that it's maximal that P(on|the cat sat...), P(the|the cat sat on...), P(mat|the cat sat on the...) Why its maximal is not in the model at all, nor in the data. It's in the data generating process, ie., us. It is we who arranged text by these frequencies and we did so because the phrase is a popular one for academic demonstrations (and so on). As ever, people attribute "to the data" or worse, "to the LLM" no properties it has.. rather it replays the data to us and we suppose the LLM must have the property that generates this data originally. Nope. Why did the tape recorder say, "the cat sat on the mat"? What, on the tape or in the recorder made "mat" the right word? Surely, the tape must have planned the word...
- flawsofar 2y agoIn case you’re thinking that rhyming requires planning, that’s just as silly as a rabbit tanning. You can make things up as you go, and the constraints emerge from the flow.
- gbasin 2y agogreat comment