3 ms·
> The model cannot output a vector and have that same vector fed back in at the next step, it only sees what token the sampler collapsed its vector into. Not c
by wren6991 2mo ago
> The model cannot output a vector and have that same vector fed back in at the next step, it only sees what token the sampler collapsed its vector into.
Not completely true: KV is a projection of the activation at each layer's input, so attention heads see (a representation of) all previous tokens' activations at that layer. The hard decision at the LM head doesn't change that.