4 ms·
LLMs don't think at all. Forcing it to be concise doesn't work because it wasn't trained on token strings that short.
by otabdeveloper4 6mo ago
LLMs don't think at all.
Forcing it to be concise doesn't work because it wasn't trained on token strings that short.
- HumanOstrich 6mo ago> Forcing it to be concise doesn't work because it wasn't trained on token strings that short. This is a 2023-era comment and is incorrect.
- otabdeveloper4 6mo agoLLMs architectures have not changed at all since 2023. > but mmuh latest SOTA from CloudCorp (c)! You don't know how these things work and all you have to go on is marketing copy.
- HumanOstrich 6mo agoYea you don't know anything about LLM architectures. They often change with each model release. You also aren't aware that there's more to it than "LLM architecture". And you're rather confident despite your lack of knowledge. You're like the old LLMs before ChatGPT was released that were kinda neat, but usually wrong and overconfident about it.
- otabdeveloper4 6mo agoIt's still attention and next-token-prediction and nothing else. The only new innovation is MoE, something that's used to optimize local models and not for the "SOTA" cloud offerings you're so fond of.
- HumanOstrich 6mo agoYou no listen. Me give up. Go learn on fruit phone.
- otabdeveloper4 6mo agoLLMs are literally next token prediction engines and nothing else. Diffusion for text is not even an academic toy at this point and will likely never be a real thing.
- Barbing 6mo agoAnything I can read that would settle the debate?
- rafram 6mo agoThey’re able to solve complex, unstructured problems independently. They can express themselves in every major human language fluently. Sure, they don’t actually have a brain like we do, but they emulate it pretty well. What’s your definition of thinking?
- otabdeveloper4 6mo agoWhen OP wrote about LLMs "thinking" he implied that they have an internal conceptual self-reflecting state. Which they don't, they *are* merely next token predicting statistical machines.