12 ms·
Naur's "Programming as Theory Building" and LLMs replacing human programmers
- n4r9 1y agoAlthough I'm sympathetic to the author's argument, I don't think they've found the best way to frame it. I have two main objections i.e. points I guess LLM advocates might dispute. Firstly: > LLMs are capable of appearing to have a theory about a program ... but it’s, charitably, illusion. To make this point stick, you would also have to show why it's not an illusion when humans "appear" to have a theory. Secondly: > Theories are developed by doing the work and LLMs do not do the work Isn't this a little... anthropocentric? That's the way humans develop theories. In principle, could a theory not be developed by transmitting information into someone's brain patterns as if they had done the work?
- IanCal 1y agoSkipping that they say it's fallacious at the start, none of the arguments in the article are valid if you simply have models 1. Run code 2. Communicate with POs 3. Iteratively write code
- n4r9 1y agoI thought the fallacy bit was tongue-in-cheek. They're not actually arguing from authority in the article. The system you describe appears to treat programmers as mere cogs. Programmers do not simply write and iterate code as dictated by POs. That's a terrible system for all but the simplest of products. We could implement that system, then lose the ability to make broad architectural improvements, effectively adapt the model to new circumstances, or fix bugs that the model cannot.
- IanCal 1y ago> The system you describe appears to treat programmers as mere cogs Not at all, it simply addresses key issues raised. That they cannot have a theory of the program because they are reading it and not actually writing it - so have them write code, fix problems and iterate. Have them communicate with others to get more understanding of the "why". > . Programmers do not simply write and iterate code as dictated by POs. Communicating with POs is not the same as writing code directed by POs.
- n4r9 1y agoOh, I think I see. You're imagining LLMs that learn from PO feedback as they go?
- IanCal 1y agoThis can be as simple as giving them search over communication with a PO, and giving them a place to store information that's searchable. How good they are at this is a different matter but the article claims it is impossible because they don't work on the code and build an understanding like people do and cannot gain that by just reading code.
- Jensson 1y ago> To make this point stick, you would also have to show why it's not an illusion when humans "appear" to have a theory. Human theory building works, we have demonstrated this, our science letting us build things on top of things proves it. LLM theory building so far doesn't, they always veer in a wrong direction after a few steps, you will need to prove that LLM can build theories just like we proved that humans can.
- jerf 1y agoYou can't prove LLMs can build theories like humans can, because we can effectively prove they can't. Most code bases do not fit in a context window. And any "theory" an LLM might build about a code base, analogously to the recent reasoning models, itself has to carve a chunk out of the context window, at what would have to be a fairly non-trivial percentage expansion of tokens versus the underlying code base, and there's already not enough tokens. There's no way that is big enough to build a theory of a code base. "Building a theory" is something I expect the next generation of AIs to do, something that has some sort of memory that isn't just a bigger and bigger context window. As I often observe, LLMs != AI. The fact that an LLM by its nature can't build a model of a program doesn't mean that some future AI can't.
- imtringued 1y agoThis is correct. The model context is a form of short term memory. It turns out LLMs have an incredible short term memory, but simultaneously that is all they have. What I personally find perplexing is that we are still stuck at having a single context window. Everyone knows that turing machines with two tapes require significantly fewer operations than a single tape turning machine that needs to simulate multiple tapes. The reasoning stuff should be thrown into a separate context window that is not subject to training loss (only the final answer).
- fouc 1y agoOr have at least 2 models. Each with their own dedicated context.
- 1y ago
- ryandv 1y ago> To make this point stick, you would also have to show why it's not an illusion when humans "appear" to have a theory. This idea has already been explored by thought experiments such as John Searle's so-called "Chinese room" [0]; an LLM cannot have a theory about a program, any more than the computer in Searle's "Chinese room" understands "Chinese" by using lookup tables to generate canned responses to an input prompt. One says the computer lacks "intentionality" regarding the topics that the LLM ostensibly appears to be discussing. Their words aren't "about" anything, they don't represent concepts or ideas or physical phenomena the same way the words and thoughts of a human do. The computer doesn't actually "understand Chinese" the way a human can. [0] https://en.wikipedia.org/wiki/Chinese_room https://en.wikipedia.org/wiki/Chinese_room
- CamperBob2 1y agoYou're seriously still going to invoke the Chinese Room argument after what we've seen lately? Wow. The computer understands Chinese better than Searle (or anyone else) understood the nature and functionality of language.
- ryandv 1y agoYou're seriously going to invoke this braindead reddit-tier of "argumentation," or rather lack thereof, by claiming bewilderment and offering zero substantive points? Wow.
- CamperBob2 1y agoYes, because the Chinese Room was a weak test the day it was proposed, and it's a heap of smoldering rhetorical wreckage now. It's Searle who failed to offer any substantive points. How do you know you're not arguing with an LLM at the moment? You don't... any more than I do.
- ryandv 1y ago> How do you know you're not arguing with an LLM at the moment? You don't. I wish I was right now. It would probably provide at least the semblance of greater insight into these topics. > the Chinese Room was a weak test the day it was proposed Why?
- psychoslave 1y ago> To make this point stick, you would also have to show why it's not an illusion when humans "appear" to have a theory. That burden of proof is on you, since you are presumably human and you are challenging the need of humans to have more than a mere appearance of having a theory when they claim to have one. Note that even when the only theoretical assumption we go with is that we will have a good laugh watching other people going crazy after random bullshits thrown at them, we still have a theory.
- dcre 1y agoI agree. Of course you can learn and use a theory without having developed it yourself!
- jimbokun 1y agoHe doesn't prove the claim. But he does make a strong argument for why it's very unlikely that an LLM would have a theory of a program similar to what a human author of a program would have: > Theories are developed by doing the work and LLMs do not do the work. They ingest the output of work. And this is certainly a true statement about how LLMs are constructed. Maybe this latently induces in the LLM something very similar to what humans do when writing programs. But another possibility is that it's similar to the Brain Teasers that were popular for a long time in programming interviews. The idea was that if the interviewee could use logic to solve riddles, they were probably also likely to be good at writing programs. In reality, it was mostly a test of whether the interviewee had reviewed all the popular riddles commonly asked in these interviews. If they had, they could also produce a realistic chain of logic to simulate the process of solving the riddle from first principles. But if that same interviewee was given a riddle not similar to one they had previously reviewed, they probably wouldn't do nearly as well in solving it. It's very likely that LLMs are like those interviewees who crammed a lot of examples, again due to how LLMs are trained. They can reproduce programs similar to ones in their training set. They can even produce explanations for their "reasoning" based on examples they've seen of explanations of why a program was written in one way instead of another. But that is a very different kind of model than the one a person builds up writing a program from scratch over a long period of time. Having said all this, I'm not sure what experiments you would run to determine if the LLM is using one approach vs another.
- latexr 1y agoFull title is: > Go read Peter Naur's "Programming as Theory Building" and then come back and tell me that LLMs can replace human programmers Which to me gives a very different understanding of what the article is going to be about than the current HN title. This is not a criticism of the submitter, I know HN has a character limit and sometimes it’s hard to condense titles without unintentionally losing meaning.
- deleted 1y ago[deleted]
- IanCal 1y agoWhat's the purpose of this? > In this essay, I will perform the logical fallacy of argument from authority (wikipedia.org) to attack the notion that large language model (LLM)-based generative "AI" systems are capable of doing the work of human programmers. Is any part of this intended to be valid? It's a very weak argument - is that the purpose?
- karmakaze 1y agoI stopped thinking that humans were smarter than machines when AlphaGo won game 3. Of course we still are in many ways, but I wouldn't make the unfounded claims that this article does--it sounds plausible but never explains how humans can be trained on bodies of work and then synthesize new ideas either. Current AI models have already made discoveries that have eluded humans for decades or longer. The difference is that we (falsely) believe we understand how the machine works and thus doesn't seem magical like our own processes. I don't know that anyone who's played Go and appreciates the depth of the game would bet against AI--all they need is a feedback mechanism and a way to try things to get feedback. Now the only great unknown is when it can apply this loop on its own underlying software.
- skydhash 1y ago> The difference is that we (falsely) believe we understand how the machine works and thus doesn't seem magical like our own processes. We do understand how the machine works and how it came to be. What most companies are seeking for is a way to make that useful.
- codr7 1y agoTo make that _seem_ useful enough for people to part with their money.
- karmakaze 1y agoThat's like saying we know how we think because we understand how neurons fire and accumulate electric charges with chemical changes. We have little idea now the information is encoded and what is represented. We're still trying to get models to explain themselves because we don't know how it arrived at a response.
- thanatropism 1y agoAlphaGo passes Hubert Dreyfus's test -- it has a world -- in a way that LLMs don't.
- philipswood 1y ago> Theories are developed by doing the work and LLMs do not do the work. They ingest the output of work. It isn't certain that this framing is true. As part of learning to predict the outcome of the work token by token, LLMs very well might be "doing the work" as an intermediate step via some kind of reverse engineering.
- skydhash 1y ago> As part of learning to predict the outcome of the work token by token They're already have the full work available. When you're reading the source code of a program to learn how it works, your objective is not to learn what keyword are close to each other or extract the common patterns. You're extracting a model which is an abstraction about some real world concept (or some other abstractions) and rules of manipulation of that abstraction. After internalizing that abstraction, you can replicate it with whatever you want, extends it further,... It's an internal model that you can shape as you please in your mind, then create a concrete realization once you're happy with the shape.
- philipswood 1y agoAs Naur describes this, the full code and documentation, and the resulting model you can build up from it is merely "walking the path" (as the blogpost put it), and does not encode "building the path". I.e. the theory of the program as it exist in the minds of the development team might not be fully available for reconstruction from just the final code and docs since it includes a lot of activity that does not end up in the code.
- skydhash 1y agoIt could be, if you were trying to only understand how the code does something. But more often, you're actively trying to understand how it was built by comparing assumptions with the code in front of you. It is not merely walking the path, if you've created a similar path and are comparing techniques.
- 1y ago
- philipswood 1y agoThe paper he quotes is a favorite of mine and I think is has strong implications for the use of LLMs, but I don't think that this implies that LLMs can't form theories or write code effectively. I suspect that the question to his final answer is: > To replace human programmers, LLMs would need to be able to build theories by Ryle’s definition
- skydhash 1y agoHaving a theory of the program, means you can argue about its current state or its transition in a new state, not merely describing what it is doing. If you see "a = b + 1" it's obvious that the variable a is taking the value of variable b incremented by one. What LLMs can't do is explaining why we have this and why it needs to change to "a = b - 1" in the new iteration. Writing code is orthogonal to this capability.
- philipswood 1y ago> What LLMs can't do is explaining why we have this and why it needs to change to "a = b - 1" in the new iteration. I did a search on Github for code containing `a=b+1` and found this: https://github.com/haoxizhong/problem/blob/a2b934ee7bb33bbe947cf2c66aaacfa3db1672bf/p82/source/std/c.cpp#L124 https://github.com/haoxizhong/problem/blob/a2b934ee7bb33bbe9... It looks to me that ChatGPT specifically does a more than OK job at explaining why we have this. https://chatgpt.com/share/680f877d-b588-8003-bed5-b425e14a5396 https://chatgpt.com/share/680f877d-b588-8003-bed5-b425e14a53... While your use of 'theory' is reasonable Naur uses a specific and more elaborate definition of theory. Example from the paper: >Case 1 concerns a compiler. It has been developed by a group A for a Language L and worked very well on computer X. Now another group B has the task to write a compiler for a language L + M, a modest extension of L, for computer Y. Group B decides that the compiler for L developed by group A will be a good starting point for their design, and get a contract with group A that they will get support in the form of full documentation, including annotated program texts and much additional written design discussion, and also personal advice. The arrangement was effective and group B managed to develop the compiler they wanted. In the present context the significant issue is the importance of the personal advice from group A in the matters that concerned how to implement the extensions M to the language. During the design phase group B made suggestions for the manner in which the extensions should be accommodated and submitted them to group A for review. In several major cases it turned out that the solutions suggested by group B were found by group A to make no use of the facilities that were not only inherent in the structure of the existing compiler but were discussed at length in its documentation, and to be based instead on additions to that structure in the form of patches that effectively destroyed its power and simplicity. The members of group A were able to spot these cases instantly and could propose simple and effective solutions, framed entirely within the existing structure. This is an example of how the full program text and additional documentation is insufficient in conveying to even the highly motivated group B the deeper insight into the design, that theory which is immediately present to the members of group A.
- voidhorse 1y agoRyle's definition of theory is actually quite reductionist and doesn't lend itself to the argument well because it is too thin to really make the kind of meaningful distinction you'd want. There are alternative views on theorizing that reject flat positivistic reductions and attempt to show that theories are metaphysical and force us to make varying degrees of ontological and normative claims, see the work of Marx Wartofsky, for example. This view is far more humanistic and ties in directly to sociological bases in praxis. This view will support the author's claims much better. Furthermore, Wartofsky differentiates between different types of cognitive representations (e.g. there is a difference between full blown theories and simple analogies). A lot of people use the term "theory" way more loosely than a proper analysis and rigorous epistemic examination would necessitate. (I'm not going to make the argument here but fwiw, it's clear under these notions that LLMs do not form theories, however, they are playing an increasingly important part in our epistemic activity of theory development)
- BiraIgnacio 1y agoGreat post and Naur's paper is really great. What I can't help stop thinking is of the many other cases where something should-not-be because being is less than ideal, and yet, they insist on being. In other words, LLMs should not be able to largely replace programmers and yet, they might.
- codr7 1y agoMight, potentially; it's all wishful thinking. I might one day wake up and find my dog to be more intelligent than me, not very likely but I can't prove it to be impossible. It's still useless.
- lo_zamoyski 1y agoIn some respects, perhaps in principle they could. But what is the point of handing off the entire process to a machine, even if you could? If programming is a tool for thinking and modeling, with execution by a machine as a secondary benefit, then outsourcing these things to LLMs contributes nothing to our understanding. By analogy, we do math because we wish to understand the mathematical universe, so to speak, not because we just want some practical result. To understand, to know, are some of the highest powers of the human person. Machines are useful for helping us enable certain work or alleviate tedium to focus on the important stuff, but handing off understanding and knowledge to a machine (if it were possible, which it isn't) would be one of the most inhuman things you could do.
- BiraIgnacio 1y agoAs a software engineer, I really hope that will be the case :) Thanks for the reply!
- falcor84 1y ago> First, you cannot obtain the "theory" of a large program without actually working with that program... > Second, you cannot effectively work on a large program without a working "theory" of that program... I find the whole argument and particularly the above to be a senseless rejection of bootstrapping. Obviously there was a point in time (for any program, individual programmer and humanity as a whole) that we didn't have a "theory" and didn't do the work, but now we have both, so a program and its theory can appear "de novo". So with that in mind, how can we reject the possibility that as an AI Agent (e.g. Aider) works on a program over time, it bootstraps a theory?
- Jensson 1y ago> So with that in mind, how can we reject the possibility that as an AI Agent (e.g. Aider) works on a program over time, it bootstraps a theory? Lack of effective memory, that might have worked if you constantly retrained the LLM incorporating the new wisdom iteratively like a human does, but current LLM architecture doesn't enable that. The context provided is neither large enough nor can it use it effectively enough for complex problems. And this isn't easy to solve, you very quickly collapse the LLM if you try to do this in the naive ways. We need some special insight that lets us update LLM continuously as it works in a positive direction the way humans can.
- falcor84 1y agoYeah, that's a good point. I absolutely agree that it needs access to effective long-term memory, but it's unclear to me that we need some "special insight". Research is relatively early on this, but we already see significant sparks of theory-building using basic memory retention, when Claude and Gemini are asked to play Pokemon [0][1]. It's clearly not at the level of a human player yet, but it (particularly Gemini) is doing significantly better than I expected at this stage. [0] https://www.twitch.tv/claudeplayspokemon https://www.twitch.tv/claudeplayspokemon [1] https://www.twitch.tv/gemini_plays_pokemon https://www.twitch.tv/gemini_plays_pokemon
- Jensson 1y ago
- ebiester 1y agoFirst, I think it's fair to say that today, an LLM cannot replace a programmer fully. However, I have two counters: - First, the rational argument right now is that one person and money spent toward LLMs can replace three - or more - programmers total. This is the argument with a three year bound. The current technology will improve and developers will learn how to use it to its potential. - Second, the optimistic argument is that a combination of the LLM model with larger context windows and other supporting technology around it will allow it to emulate a theory of mind that is similar to the average programmer. Consider Go or Chess - we didn't think computers had the theory of mind to be better than a human, but it found other ways. For humans, Naur's advice stands. We cannot assume that this is true if there are tools with different strengths and weaknesses than humans.
- ActionHank 1y agoI think that everyone is misjudging what will improve. There is no doubt it will improve, but if you look at a car, it is still the same fundamental "shape" of a model T. There are niceties and conveniences, efficiency went way up, but we don't have flying cars. I think we are going to have something, somewhere in the middle, AI features will eventually find their niche, people will continue to leverage whatever tools and products are available to build the best thing they can. I believe that a future of self-writing code pooping out products, AI doing all the other white collar jobs, and robots doing the rest cannot work. Fundamentally there is no "business" without customers and no customers if no one is earning.
- ebiester 1y agoYou cannot build a tractor unit (the engine-cab half of the tractor-trailer) with Model T Technology even if they are close. And the changes will be in the auxiliary features. We will figure out ways to have LLMs understand APIs better without training them. We will figure out ways to better focus its context. We will chain LLM requests and contexts in a way that help solve problems better. We will figure out ways to pass context from session to session that an LLM can effectively have a learning memory. And we will figure out our own best practices to emphasize their strengths and minimize their weaknesses. (We will build better roads.) And as much as you want to say that - a Model T was uncomfortable, had a range of about 150 miles between fill-ups, and maxed out at 40-45 mph. It also broke frequently and required significant maintenance. It might take 13-14 days to get a Model T from new york to los angeles today notwithstanding maintenance issues, and a modern car could make it reliably in 4-5 days if you are driving legally and not pushing more than 10 hours a day. I too think that self-writing code is not going to happen, but I do think there is a lot of efficiency to be made.
- BenoitEssiambre 1y agoSolomonoff induction says that the shortest program that can simulate something is its best explanatory theory. OpenAI researchers very much seem to be trying to do theory building ( https://x.com/bessiambre/status/1910424632248934495 https://x.com/bessiambre/status/1910424632248934495 ).
- ninetyninenine 1y agoCan he prove what he says? The foundation of his argument rests on hand waves and vague definitions on what a mind is and what a theory is that it ultimately goes nowhere. Then he makes a claim and doesn’t back it up. I have a new concept for the author to understand: proof. He doesn’t have any. Let me tell you something about LLMs. We don’t understand what’s going on internally. LLMs say things that are true and untrue just like humans do and we don’t know if what it says is a general lack of theory building ability or if it’s lying or if it has flickers of theory building and becomes delusional at other times. We literally do not know. The whole thing is a black box that we can only poke at. What ticks me off is all these geniuses who write these blog posts with the authority of a know it all when clearly we have no fucking clue about what’s going on. Even more genius is when he uses concepts like “mind” and “theory” building the most hand wavy disagreed upon words in existence and rest his foundations on these words when no people ever really agree on what these fucking things are. You can muse philosophically all you want and in any direction but it’s all bs without definitive proof. It’s like religion. How people made up shit about nature because they didn’t truly understand nature. This is the idiocy with this article. It’s building a religious following and making wild claims without proof.
- turtlethink 1y agoTo add on that - the human mind is a black box that we don't understand, we don't know what's going in internally, and most descriptions like the one here are arbitrary, made up, and don't reflect reality
- ninetyninenine 1y agoyep. Same principle. Totally agreed. The irony is that we feel we can make more claims about LLMs because we created LLMs. We created something that we don't understand.
- stevenhuang 1y agoWell said. I wouldn't be surprised if at the root of it for these people is motivated reasoning from belief there is something special about the mind that machines cannot possess. Akin to the wayward belief animals can't feel pain, only humans do. Which we now realize is wrong, actually some animals understand pain and suffer just as much as humans can. Would not be surprised if we come to a similar realization for LLMs and our understanding what it means to reason.
- fedeb95 1y agothere's an additional difficulty. Who told the man to build a road? This is the main stuff that LLMs or any other technology currently seem to lack, the "why", a reason to do stuff a certain way and not another. A problem as old as human itself.
- lo_zamoyski 1y agoYes, but it's more than that. As I've written before, LLMs (and all AI) lack intentionality. They do not possess concepts. They only possess, at best, conventional physical elements of signs whose meaning, and in fact identity as signs, are entirely subjective and observer relative, belonging only to the human user who interprets these signs. It's a bit like a book: the streaks of pigmentation on cellulose have no intrinsic meaning apart from being streaks of pigmentation on cellulose. They possess none of the conceptual content we associate with books. All of the meaning comes from the reader who must first treat these marks on paper as signs, and then interpret these signs accordingly. That's what the meaning of "reading" entails: the interpretation of symbols, which is to say, the assignment of meanings to symbols. Formal languages are the same, and all physical machines typically contain are some kind of physical state that can be changed in ways established by convention that align with interpretation. LLMs, from a computational perspective, are just a particular application. They do not introduce a new phenomenon into the world. So in that sense, of course LLMs cannot build theories strictly speaking, but they can perhaps rearrange symbols in a manner consistent with their training that might aid human users. To make it more explicit: can LLMs/AI be powerful practically? Sure. But practicality is not identity. And even if an LLM can produce desired effects, the aim of theory in its strictest sense is understanding on the part of the person practicing it. Even if LLMs could understand and practice theory, unless they were used to aid us in our understanding of the world, who cares? I want to understand reality!
- fedeb95 1y agoI get your point and I agree to a certain extent. However, it's arguable that everyone shares the same aim, that is, to understand reality. Some want to go down a road, no matter how it got built, or, some think, or better don't think, no matter where it leads. In that world, an artificial entity that can 1) create an aim and 2) build enough understanding to execute and 3) execute could be valuable. Right now we're at the 3) in the specific context of byte arrays. Now, an artificial system that could also understand, i.e. possess some kind of structure of concepts, and from there also produce the "need" to create something, that would be a huge leap forward. Forward toward what? I don't know.
- xpe 1y ago> Theories are developed by doing the work and LLMs do not do the work. They ingest the output of work. This is often the case but does not _have_ to be so. LLMs can use chain of thought to “talk out loud” and “do the work”. It can use supplementary documents and iterate on its work. The quality of course varies, but it is getting better. When I read Gemini 2.5’s “thinking” notes, it indeed can build up text that is not directly present in its training data. Putting aside anthropocentric definitions of “reasoning” and “consciousness” are key to how I think about the issues here. I’m intentionally steering completely clear of consciousness. Modern SOTA LLMs are indeed getting better at what people call “reasoning”. We don’t need to quibble over defining some quality bar; that is probably context-dependent and maybe even arbitrary. It is clear LLMs are doing better at “reasoning” — I’m using quotes to emphasize that (to me) it doesn’t matter if their inner mechanisms for doing reasoning don’t look like human mechanisms. Instead, run experiments and look at the results. We’re not talking about the hard problem of consciousness, we’re talking about something that can indeed be measured: roughly speaking, the ability to derive new truths from existing ones. (Because this topic is charged and easily misunderstood, let me clarify some questions that I’m not commenting on here: How far can the transformer-based model take us? Are data and power hungry AI models cost-effective? What viable business plans exist? How much short-term risk, to say, employment and cybersecurity? How much long-term risk to human values, security, thriving, and self-determination?) Even if you disagree with parts of my characterization above, hear this: We should at least be honest to ourselves when we move the goal posts. Don’t mistake my tone for zealotry. I’m open to careful criticism. If you do, please don’t try to lump me into one “side” on the topic of AI — whether it be market conditions, commercialization, safety, or research priorities — you probably don’t know me well enough to do that (yet). Apologies for the pre-defensive posture; but the convos here are often … fraught, so I’m trying to head off some of the usual styles of reply.
- geraneum 1y ago> it indeed can build up text that is not directly present in its training data. I’m curious how you know that.
- triclops200 1y ago
- analyte123 1y agoThe "theory" of a program is supposed to be majority embedded in its identifiers, tests, and type definitions. The same line of reasoning in this article could be used to argue that you should just name all your variables random 1 or 2 letter combinations since the theory is supposed to be all in your head anyway. Indeed, it's quickly obvious where an LLM is lacking context because the type of a variable is not well-specified (or specified at all), the schema of a JSON blob is not specified, or there is some other secret constraint that maybe someone had in their head X years ago.
- woah 1y agoThese long winded philosophical arguments about what LLMs can't do which are invariably proven wrong within months are about as misguided as the gloom and doom pieces about how corporations will be staffed by "teams" of "AI agents". Maybe it's best just to let them cancel each other out. Both types of article seem to be written by people with little experience actually using AI.
- turtlethink 1y agoMost people in the tech space talking about AI, not only misunderstand AI - but they usually have a far greater misconception of the human mind/brain. The basic argument in the article above (and in most of this comment thread) is that LLMs could never reason because they can't do what humans are doing when we reason. This whole thread is amusingly a rebuttal of itself. I would argue it's humans that can't reason, because of what we do when we "reason", the proof being this article which is a silly output of human reasoning. In other words, the above argument for why LLMs can't reason are so obviously fallacious in a multiple ways, the first of which is that human reasoning is a golden standard of reasoning, (and are a good example of how bad humans are at reasonin. LLMs use naive statistical models to find the probability of a certain output, like "what's the most likely next word". Humans use equally rationally-irrelevant models that are something along the lines of "what's the most likely next word that would have the best internal/external consequence in terms of dopamine or more indirectly social standing, survival, etc." We have very weak rational and logic circuits that arrive at wrong conclusions far more often than right conclusions, as long as it's beneficial to whatever goal our mind thinks is subconsciously helpful to survival. Often that is simple nonsense output that just sounds good to the listener (e.g. most human conversation) Think how much nonsense you have seen output by the very "smartest" of humans. That is human reasoning. We are woefully ignorant of the actual mechanics of our own reasoning. The brain is a marvelous machine, but it's not what you think it is.
- edanm 1y agoI really don't understand how people can write entire articles which can be disproven with an hour of work. This is like writing a long polemic on how AIs will never be able to play Chess because of... reasons... four years after Deep Blue beat Kasparov.
- andai 1y agoIf I understand correctly, the critique here is that is that LLMs cannot generate new knowledge, and/or that they cannot remember it. The former is false, and the latter is kind of true -- the network does not update itself yet, unfortunately, but we work around it with careful manipulation of the context. Part of the discussion here is that when an LLM is working with a system that it designed, it understands it better than one it didn't. Because the system matches its own "expectations", its own "habits" (overall design, naming conventions, etc.) I often notice complicated systems created by humans (e.g. 20 page long prompts), adding more and more to the prompt, to compensate for the fact that the model is fundamentally struggling to work in the way asked of it, instead of letting the model design a workflow that comes naturally to it.
- drbig 1y ago> If I understand correctly, the critique here is that is that LLMs cannot generate new knowledge, and/or that they cannot remember it. > The former is false, and the latter is kind of true -- the network does not update itself yet, unfortunately, but we work around it with careful manipulation of the context. Any and all examples of where an LLM generated "new knowledge" will be greatly appreciated. And the quotes are because I'm willing to start with the lowest bar of what "new" and "knowledge" mean when combined.
- andai 1y agoThey are fundamentally mathematical models which extrapolate from data points, and occasionally they will extrapolate in a way that is consistent with reality, i.e. they will approximate uncharted territory with reasonable accuracy. Of course, being able to tell the difference (both for the human and the machine) is the real trick! Reasoning seems to be a case where the model uncovers what, to some degree, it already "knows". Conversely, some experimental models (e.g. Meta's work with Concepts) shift that compute to train time, i.e. spend more compute per training token. Either way, they're mining "more meaning" out of the data by "working harder". This is one area where I see that synthetic data could have a big advantage. Training the next gen of LLMs on the results of the previous generation's thinking would mean that you "cache" that thinking -- it doesn't need to start from scratch every time, so it could solve problems more efficiently, and (given the same resources) it would be able to go further. Of course, the problem here is that most reasoning is dogshit, and you'd need to first build a system smart enough to pick out the good stuff... --- It occurs to me now that you rather hoped for a concrete example. The ones that come to mind involve drawing parallels between seemingly unrelated things. On some level, things are the same shape. I argue that noticing such a connection, such a pattern, and naming it, constitutes new and useful knowledge. This is something I spend a lot of time doing (mostly for my own amusement!), and I've found that LLMs are surprisingly good at it. They can use known patterns to coherently describe previously unnamed ones. In other words, they map concepts onto other concepts in ways that hasn't been done before. What I'm referring to here is, I will prompt the LLM with some such query, and it will "get it", in ways I wasn't expecting. The real trick would be to get it to do that on its own, i.e. without me prompting it (or, with current tech, find a way to get it to prompt itself that produces similar results... and then feed that into some kind of Novelty+Coherence filtering system, i.e. the "real trick" again... :). A specific example eludes me now, but it's usually a matter of "X is actually a special case of Y", or "how does X map onto Y". It's pretty good at mapping the territory. It's not "creating new territory" by doing that, it's just pointing out things that "have always been there, but nobody has looked at before", if that makes sense.
- andriesm 1y agoMany like to say that LLM's cannot do ANY reasoning or "theory building". However, is it really true that LLM's cannot reason AT ALL or cannot do theory construction AT ALL? Maybe they are just pretty bad at it. Say 2 out 10. But almost certainly not 0 out of 10. They used to be at 0, and now they're at 2. Systematically breaking down problems and Systematically reasononing through parts, as we can see with chain-of-thought hints that further improvements may come. What most people however now agree is that LLMs can learn and apply existing theories. So if you teach an LLM enough theories iy can still be VERY useful and solve many coding problems, because an LLM can memorise more theories than any human can. Big chunks of computer software still keeps reinventing wheels. The other objection from the article, that without theory building an AI cannot make additions or changes to a large code base very effectively - this suggests an idea to try - before promting the AI for a change on a large code base, prepend it with a big description of the entire program, the main ideas and how they map yo certain files, classes, modules etc, and see if this doesn't improve your results? And in case you are concerned that. documenting and typing out entire system theories for every new prompt, keep in mind that this is something you can write once and keep reusing (and adding to over time incrementally). Of course context limits may still be a constraint. Of course I am not saying "definitely AI will make all human programmers jobless". I'm merely saying, these things are already a massive productivity boost, if used correctly. I've been programming for 30 years, started using cursor last year, and you would need to fight me to take it away from me. I'm happy to press ESC to cancel all the bad code suggestions, to still have all thr good tab-completes, prompts, better than stack-overflow question answering etc.
- hnaccountme 1y agoI completely agree with this. But most programmers i've encountered are just converting English to <programming language>. If a bug is reported then convert English to <programming language> AI is the new Crypto