6 ms·
> LLMs as they currently exist cannot master a theory, design, or mental construct because they don't remember beyond their context window. Only humans can can
by paradite 1y ago
> LLMs as they currently exist cannot master a theory, design, or mental construct because they don't remember beyond their context window. Only humans can can gain and retain program theory.
False.
> An LLM is a token predictor. It works only at the level of text. It is not capable of working at a conceptual level: it doesn't reason about ideas, diagrams, or requirements specifications.
False.
Anyone who have spent time in machine learning or reinforcement learning understands that models are projections of higher dimension concepts on to lower dimensions as weights.
- mjburgess 1y agoThe problem with people "who have spent time in machine learning or reinforcement learning" is that they've spent no time, literally none, understanding what a concept is. There is no such thing as a higher dimensional concept, nor can they be projected into a weight space, because they aren't quantities. The concept, say, "Dog" composes with the concept, "Happy" to form "Happy Dog". The extension(Dog) is all possible dogs, the extension(Happy) is all happy objects, the extensions here compose. The intension of "Dog" depends on its context, eg., "get that damned dog!" has a different intension than, "I wish I looked less like a dog!". And the intensions here do not compose like the extensions. Take the act of counterfactual reality-oriented simulation, called "imagination" narrowly, call that `I`. And "curry" this operator with a simulation premise, "(as if) on mars", so we have I' = I(Mars)(...). Now, what is the content of I'(Dog), I(Happy Dog), I(Get that damned dog), I(get that damned happy dog) ? and so on The contents of `I` is nowhere modelled by "projection" because this does not model composition, and is not relevantly discrete and bounded by logical connectives. These are trivial issues which arise the moment you're aware that few writing computer science papers have ever studied the meaning of the words they use with abandon.
- paradite 1y agoYou seem to believe "concept" is a concept in English language. It is not. Concept is an abstraction layer above human languages. Here's a good article that touched on this topic: https://www.neelnanda.io/mechanistic-interpretability/glossary https://www.neelnanda.io/mechanistic-interpretability/glossa...
- mjburgess 1y agoI'm very well-familiar with the literature in this area. I understand that computer scientists, with obscene and wild abandon, will just pick whatever word suits their agenda and define it opportunistically, without regards to the confusion it will cause -- indeed, seemingly with this aim -- to "impress" the reader and make their research seem extraordinary. "Concept" is not a term from computer science, its use here has not only been "narrowed" but flat-out redefined. "Concept" as used in XAI (a field in which i've done research) is an extremely limited attempt to capture the extension of concepts over a given training domain. Concept, as used by the author of this article that you are replying to, and 99.9999999...% of all people familiar with the term, means "concept". It does not mean what it has been asserted to mean in XAI. And one of the most basic features of concepts is their semantic content, that they compose, that they form parts of propositions, and so on.
- paradite 1y agoYou have biased view on the definition of "concept" based on English language and logic. In Chinese language, concept is 概念. In Chinese language, happy dog is 快乐的狗. Notice it has an extra "的" that is missing in English language. This tells you that you can't just treat English grammar and structure as the formal definition of "concept". Some languages do not have words for happiness, or dog. But that doesn't mean the concept of happiness or dog does not exist. The reverse is also true, you can't claim a concept does not exist if it does not exist in English language. Concept is something beyond any particular language or logical construct or notations that you invent.
- mjburgess 1y ago> But that doesn't mean the concept of happiness or dog does not exist. That would be a consequence of your position. The person who wrote the article is english. The claim being evaluated here is from the article. The term "concept" is english. The *meaning* of that term isn't english, any more than the meaning of "two" is english. My analysis of "concept" has nothing to do with the english language. "Happy" here stands in for any property-concept and 'dog' any term which can be a object-concept, or a property-concept, or others. If some other language has terms which are translated into terms that do not function in the same way, then that would be a bad translation for the purpose of discussing the structure of concepts. It is you who are hijacking the meaning of "concept", ignoring the meaning the author intended, substituting one made up 5 minutes ago by self-aggrandising poorly read people in XAI -- and the going off about irrelevant translations into Chinese. The claim the author made has nothing to do with XAI, nor chinese, nor english. It has to do with mental capacities to "bring objects under a concept", partition experience into its conceptual structure ("conceptualise"), simulate scenarios based on compositions of concepts ("the imagination") and so on. These are mental capabilities a wide class of animals possess, who know no language; that LLMs do not possess.
- rhubarbtree 1y agoIf the effectiveness of LLMs has taught me anything, it’s that concepts are quantities, at least in the sense that they are vectors. Good job too, otherwise the brain would have to magically create some entirely new physics to represent concepts.
- viraptor 1y agoI'm not sure why "The contents of `I` is nowhere modelled by "projection" because this does not model composition, and is not relevantly discrete and bounded by logical connectives." In practical terms, what do you think the LLM output cannot contain right now? Because the way I read it now is "LLM can't speculate". But that's trivial to disprove by asking for that happy dog on Mars speculation you have as an example - whether you want the scientific version, or child level fun, it's available and the model will give nontrivial idea connections that I could not find anywhere. (For example childlike speculation from Claude included that maybe dogs would be ok playing in spacesuits since some dogs like wearing little coats) Similarly "And the intensions here do not compose like the extensions." is really high level. What's the actual claim with regards to LLMs?
- mjburgess 1y agoLLMs are just a token->token mapping. They can output any set of tokens for any set of input tokens. So there is no output which isn't in the domain or codomain. The issue is why one (prompt, answer) pair is given. If the answer is given as a "reasoning process" over salient parts of the prompt, that, e.g., involves imagining/simulation as expected, then for {(prompt', answer')} of similar imaginings we will get reliable mappings. If its cheating, then we wont. We can, I think, say for certain that the system is not engaged in counterfactual reasoning. Eg., we can give a series of prompts (p1, p2, p3...) which require increasing complexity of the imagined scenario, and we do not find O(answering) to follow O(p-complexity-increase). Rather the search strategy is always the same, and we can just get "mildly above linear" (pseudo-)reasoning complexity with chain-of-thought.
- viraptor 1y ago> LLMs are just a token->token mapping. They can output any set of tokens for any set of input tokens. So there is no output which isn't in the domain or codomain. This applies the same to humans hearing a question and responding. Tokens in, tokens out (whether words or sound). It's not unique to LLMs, so not useful for explaining differences. > then for {(prompt', answer')} of similar imaginings we will get reliable mappings. If its cheating, then we wont. You're not really showing that this is/isn't the case already. Also this would put people with quirky ideas and wild imagination in the "cheating" group if I understand your claim correctly. There's even a whole game around a similar concept - Dixit - describe an image in a way that as few people as possible will get it. > we can give a series of prompts (p1, p2, p3...) which require increasing complexity of the imagined scenario, and we do not find O(answering) to follow O(p-complexity-increase). Rather the search strategy is always the same You're describing most current implementations, not a property of LLMs. Gemini scales the thinking phase for example. Future models are likely to do the same. Another recent post implemented this too https://news.ycombinator.com/item?id=44112326 https://news.ycombinator.com/item?id=44112326
- TeMPOraL 1y agoYup. That's exactly what language models represent internally; that's what the high-dimensional latent space is exactly about - reifying meaning, defining concepts in terms of relationships to other concepts. LLMs are the idea you describe but made incarnate, in form of a computing artifact we can "hold in our hands", study and play with. IMHO people are still under-appreciating how big a thing this is fundamentally, beyond RAG and chatbots.
- mjburgess 1y agoSure, they're a reification of some aspect of meaning. The question is: which aspect(s), and which not. It is also the case that animals do not "reliably and univerally" implement all aspects of all meanings they are acquainted with, so we aren't looking for 100% of capacities, 100% of the time. Nevertheless, LLMs are only implementing a limited aspect of meaning: mostly association and "some extension". And with this, plus everything ever written, they can narrowly appear to implement much more. Let's be clear though, when we say "implement" we mean that an answer arises from a prompt for a very specific reason: because the answer is meant by the system in the relevant way. In this sense, LLMs can mean any association, perhaps they can mean a few extensions, but they cannot mean anything else. Whenver an LLM appears to partake in more aspects of meaning it is only cheating: it is using familiarity with families of associations to overcome its disabilities. Like the idiot savant who appears to know all hollywood starlets, but is discovered eventually, not to realise they are all film stars. We routinely discover these disabilities in LLMs, when they attempt to engage in reasoning beyond these (formally,) narrow contexts of use. Agentic AI is a very good "on steroids" version of this. Just try to use windsurf, and the brittle edges of this trick appear quickly. It's "reasononing" whenver it seems to work, and "hallucination" when not -- but of course, it just never was reasoning.
- TeMPOraL 1y ago> LLMs are only implementing a limited aspect of meaning: mostly association and "some extension". > Whenver an LLM appears to partake in more aspects of meaning it is only cheating: it is using familiarity with families of associations to overcome its disabilities. I'm not convinced there's anything more to "meaning" - we seem to be defining concepts through relationship to other concepts, and ground that directly or indirectly with experiences. The richer that structure is, the more nuanced it gets. > Like the idiot savant who appears to know all hollywood starlets, but is discovered eventually, not to realise they are all film stars. We routinely discover these disabilities in LLMs, when they attempt to engage in reasoning beyond these (formally,) narrow contexts of use. I see those as limitations of degree, not kind. Less idiot savant, more like someone being hurried to answer questions on the spot. Some associations are stronger and come to mind immediately, some are less "fresh in memory", and then associations can bring false positives and it takes extra time/effort to notice and correct those. It's a common human experience, too. "Yes, those reserved words are 'void', 'var', 'volatile',... wait, 'var' is JS stuff, it's not reserved in C..." etc. Then, of course, humans are learning continuously, and - perhaps more importantly - even if they're not learning, they're reinforcing (or attenuating) existing associations through bringing them up and observing feedback. LLMs can't do that on-line, but that's an engineering limitation, not a theoretical one. I'm not claiming that LLMs are equivalent to humans in general sense. Just that they seem to be implementing the fundamental machinery behind "meaning" and "understanding" in general sense, and the theoretical structure behind it is quite pretty, and looks to me like a solution to a host of philosophical problems around meaning and language.
- TeMPOraL 1y agoProgram theory is the Stochastic Parrot argument of 2025. Suddenly everyone is name-dropping Naur and quoting the same bit of his seminal essay, then pointing at it and saying "this!", without providing any sort of coherent argument why would that point to LLM limitations, or be relevant to the topic in the first place.