12 ms·
You're mixing training loss (cross-entropy/surprisal of next token) and post-training prediction decoding (done e.g. with beam search)
by puttycat 2y ago
You're mixing training loss (cross-entropy/surprisal of next token) and post-training prediction decoding (done e.g. with beam search)
- Xcelerate 2y agoTraining loss considers only the next single token, right? (I’m not up-to-date on the SOTA.) I thought post-training prediction still only directly predicts the next token and beam search is sort of a meta-model applied over that (i.e., it is a model on top of the output of the model that performs next-token prediction—beam search considers at each iteration a subset of the current next-token predictions ranked by their probability to use as multiple starting points for predicting the next token, while keeping track of the joint probabilities to prune the set of candidate sequences at each step). Seems like beam search would fail drastically in cases where the true (unknown) probability distribution over all sequences of tokens of length n has very low conditional probabilities for the first few tokens, each given the computed joint probability of the prior predicted tokens. That is, the true values of p(t2|t1), p(t3|t2,t1), p(t4|t3,t2,t1), ... as derived from the unknown p(t1,t2,...,tn) are very small, but very high when computed via a next-token prediction model. I’m suggesting to modify both. Use cross-entropy of the nth token for training loss. Use cross-entropy of nth token for post-training prediction and then work backward from there to the beginning of your sequence prediction.
- namibj 2y agoThe problem is that a position's probability output is conditioned via attention on all previous positions. If you want to be better you need to switch to DDPMs for example (e.g. an encoder-only transformer to predict diffusion transition probabilities in parallel, then apply steps of denoising). The problem is just that these don't work so well from auto regressive decoder transformers, and encoder-decoder architectures like e.g. Google's T5 have fallen out of favor since about LLAMA dropped.