3 ms·
I think you misunderstood how the generation of text works. For each new token it samples probabilities given previous tokens, not averages, then chooses some t
by bitL 4y ago
I think you misunderstood how the generation of text works. For each new token it samples probabilities given previous tokens, not averages, then chooses some token from the top k as the next one with rules that penalize repetition of some order.
Moreover, there is no upper bound for transformers found yet, i.e. the larger the model is and the more data is used for training, the better it performs. It's literally about who is able to throw more money at it at this point, with some closely guarded secrets like warm up steps, training schedules etc. There is also the overfitting effect where one pushes training far beyond overfitting (validation loss growing again) as with transformers at some point the overfitting stops, validation loss starts dropping again and that's when the magic starts happening and money are burnt for scale.
- circuit10 4y agoYou missed the point they were making, which is that the probabilities it’s predicting are based on what it expects the average text in its training set to look like. The loss you’re talking about is how closely its answers match the training set, not how clever the answers sound (though with RLHF it’s different). A model producing better text than what’s in its training set would be penalised for not matching it closely and quickly learn to not do that