4 ms·
It's still attention and next-token-prediction and nothing else. The only new innovation is MoE, something that's used to optimize local models and not for the
by otabdeveloper4 6mo ago
It's still attention and next-token-prediction and nothing else.
The only new innovation is MoE, something that's used to optimize local models and not for the "SOTA" cloud offerings you're so fond of.
- HumanOstrich 6mo agoYou no listen. Me give up. Go learn on fruit phone.
- otabdeveloper4 6mo agoLLMs are literally next token prediction engines and nothing else. Diffusion for text is not even an academic toy at this point and will likely never be a real thing.