5 ms·
No OP, but no amount of data will ever enable an LLM to know when it doesn’t know something, for instance. That’s a fundamental limitation with the current arch
by 4death4 3y ago
No OP, but no amount of data will ever enable an LLM to know when it doesn’t know something, for instance. That’s a fundamental limitation with the current architecture.
- bilbo0s 3y agoBut that's different from scaling. You can not know what you don't know, and learn gobs of new stuff at the same time. That's the natural state of most learners.
- deleted 3y ago[deleted]
- datameta 3y agoWhen we have an external source of information that is not self-accountable, it is different from when we are both the learner and functor - when we notice mistakes we learn from them. We're always looking out for errors, and if they happen then (usually) we alone feel responsible for their effects.
- bilbo0s 3y agoAgain, you're talking about something different than scaling. There are many different aspects of emotions, life, or even learning. But the comment was about scaling learning, which can be unaffected by other aspects, such as whether you learned something incorrect. More succinctly: I can still teach Simpson's Rule in a Calculus class to a young man or woman who subscribes to Creationism, or the Flat Earth theory. None of their false knowledge affects their innate ability to take in new knowledge.
- 4death4 3y agoBut what is the purpose of "scaling?" Broadly speaking, it means that there is a marked and obtainable improvement on an objective function in response to additional data. But for certain criteria, like the possession of knowledge, it appears attention-based LLMs aren't capable of that, regardless the amount of data. So they don't "scale" if you use that as an objective function.
- IanCal 3y agoThey can say they don't know things already. How are you defining "know" here?
- nyrikki 3y agoAh the fun of known knowns, known unknowns, unknown unknowns, and weak negation. https://blog.research.google/2024/01/can-large-language-models-identify-and.html https://blog.research.google/2024/01/can-large-language-mode... Negation in the training data, exhaustion, or unknowable unknowns like future events is about an LLM will ever be able to say it "doesn't know" reliably. LLMs have no concept of truth or correctness really, they are just probabilistic appending to output to match learned ideal patterns.
- humansareok1 3y ago>LLMs have no concept of truth or correctness really, they are just probabilistic appending to output to match learned ideal patterns. I think they do have those concepts, like they have many others but they are not bound to them in any way.
- fhd2 3y agoThey're trained (and possibly scripted) to not "know" things or not be able to do things. But at their core, they're going to generate the most plausible continuation, even if that isn't very plausible (like non-existent functions). I think this can be improved (playing with the temperature in self hosted models depending on the nature of the task e.g.), but it doesn't seem fully solvable with current LLMs.