4 ms·
Right. They can do all those things. And none of that will make it smart or able to learn new things. The underlying model is just an llm. But judging from the
by daveguy 3mo ago
Right. They can do all those things. And none of that will make it smart or able to learn new things. The underlying model is just an llm. But judging from the downvotes, it seems AI folks get upset when someone talks honestly about their precious piles of matrix multiplication.
- sureMan6 3mo agoMight bother you to use anthropomorphic terminology like smart and learning but they are capable of producing work that traditionally required human intelligence and the whole point of gpt 3 was the ability to "learn", you can give it an example of an invented brand new coding language and it can write working code in that language
- xylenox 3mo agoYep, people always forget that early LLMs were sold as "Zero Shot Learning".
- kmacdough 3mo agoSold as learning, but that was a marketing term, not a technical one. From a technical perspective, the LLM is not learning. Only reacting based on its original training. You might argue that the systems we've built around them are learning in a way, as they strategically condense and save artifacts from past interactions to pass into the LLMs context. But the LLM itself, which is the source of the intelligence, is not learning. It remains entirely unchanged throughout inference. This difference may seem trite, but it has significant impacts over the long term behavior.
- DiogenesKynikos 3mo agoYou're making a highly arbitrary distinction between learning and ... learning? The LLM can immediately learn and start using new skills. To any lay person, that's what learning is.
- kmacdough 3mo agoContext is not the same as learning. It's easy to conflate because they're tightly coupled in our brains. The underlying structure and tuning of the LLM are entirely unchanged by context. It merely affects the attention and activation of the network. The LLM will not be able to work with this hypothetical new language unless it is in context. This does not fit the computational meaning of learning. Smart is not a well defined term. Nor is it's general idea formally understood. Use it freely, but you won't be saying anything meaningful unless you define your usage.
- DiogenesKynikos 3mo agoThe LLM is the model + context. The output depends on both. You're making an artificial distinction. The LLM sees a new programming language for the first time, and can immediately code in it. That's learning by any reasonable definition. If you go past the context window, it forgets, which is a limitation of current LLMs. But as long as it learned how to code in the new language within its context window, it has gained that new ability.
- daveguy 3mo agoThe context is the input to the LLM model. It seems like you need to study up on how LLMs work, instead of spouting hype.
- DiogenesKynikos 3mo agoIt sounds like you need to be less condescending and engage with my actual comment. First, the "hype" is not externally imposed. It's the genuine reaction of hundreds of millions of people to a technology that would have been considered science fiction just four years ago. The fact that I can't be certain you whether you are an LLM or a human is incredible. Second, the output of the LLM depends on both its weights and the context. If it has seen something in its context window, it now knows it for all intents and purposes. It learns, in other words.
- jatora 3mo agoNo thats probably because you misread what you were replying to and your comment was out of left field. They didnt imply models get better intra-releasally at all.
- DiogenesKynikos 3mo agoI can imagine an AI insulting humans in the same way: "The underlying model is just a biological neutral network. It seems you carbonoids get upset when someone talks honestly about synapses and neuron firing."
- daveguy 3mo agoNeural plasticity is real, and something LLMs are incapable of. So sorry.
- dahart 3mo agoTrue for today’s static models during inference. Not true for self-supervised learning, not true during training or fine-tuning, of course. Ignores that LLMs might start continuous training in the future - there’s no fundamental or technical constraint that prevents LLM ‘plasticity’. And ignores that accumulating context/memories/skills/etc affects performance and might count as a valid analogy to what many people loosely call ‘neural plasticity’, which is sometimes casually mistaking knowledge for network modification.
- daveguy 3mo agoBoth of those things only happen once by the LLM model provider and not every time a prompt is issued.
- dahart 3mo agoToday, depending on which model you use. You’re making unstated assumptions. And that’s not a fundamental property of LLMs, it’s happenstance. LLMs are capable of ‘plasticity’, by design.
- daveguy 3mo agoYou are incorrect. "memory.md" and other context manipulations do not change the underlying model.
- coldtea 3mo agoYou used the word "smart" now, whereas on the comment I replied to, you said "better". Tuning those can definitely make a model respond better or worse. So your claim (quoting 100% as written) that "Their performance depends solely on the model training before release and how well you curate the context you feed it" is wrong. Hence the downvotes. Doesn't matter if LLMs are to be considered intelligent or not for the claim to be wrong. > But judging from the downvotes, it seems AI folks get upset when someone talks honestly about their precious piles of matrix multiplication. Often yes. In this case, it's more like they get upset when someone says something factually wrong, and then defensively changes the goalposts.
- daveguy 3mo ago> Often yes. In this case, it's more like they get upset when someone says something factually wrong, and then defensively changes the goalposts. Oh give me a break. Show me one example of 1) any knob twisting that makes the underlying model better. or 2) any example of the AI providers twisting those knobs to do anything other than degrade performance for their own bottom line or safety. The current post says: "it would be expected for a better model to use different amounts of brevity if it gets better at determining the appropriate amount." When no, the model cannot "get better". It doesn't determine any appropriateness of response realtime except for the weights baked into it from the beginning and whatever context it can muster. If you cram enough guidance that it doesn't decide to ignore maybe you can make it more brief. But it (the model) can do none of those things. LLM models are literally stupid by design.
- blackqueeriroh 3mo ago> If you cram enough guidance that it doesn't decide to ignore maybe you can make it more brief. You are now anthropomorphizing the model yourself.
- coldtea 3mo ago>Oh give me a break. Show me one example of 1) any knob twisting that makes the underlying model better. I mentioned several. You're now once again changing goalpoasts to say you meant the underlying model, not the overall llm performance, even though you explicitly wrote: "Their performance depends solely on the model training before release and how well you curate the context you feed it". So, the context curation was relevant (meaning you didn't constrain your claim to the underlying model), but now somehow all the additional tunables aren't relevant (because suddenly you're just talking about the model). End of discussion.
- Nevermark 3mo agoIntelligence can operate without learning. At a minimum inference and learning don’t need to be co-concurrent. Not disagreeing with your point, but your terminology muddies your point. But your point doesn't acknowledge that even with inference, there is a lot of room to tune the calculations. Multiple models, quantization tradeoffs are just the most obvious examples. Every architecture can be adjusted to increase intelligence/watt or other measure, even without further training.