2 ms·
there certainly are experiments in keeping the reasoning in latent space. i dont think this is quite the same though, since you arent picking tokens for the ch
by 8note 3mo ago
there certainly are experiments in keeping the reasoning in latent space.
i dont think this is quite the same though, since you arent picking tokens for the chain of thought. inatead, its staying on trying to pick the immediate next token.
as an alternative, maybe you could stack these to produce most likely token lists instead by stacking these?
but i think youd end up with the similar blurriness that llm video generators get where theyre returning an average of all the likely combinations rather than collapsing that wave function