32 ms·
Peter Norvig critically reviews AlphaCode’s code quality
- Octokat 4y agoHis repo is a gold mine
- trynewideas 4y agoThis is a great review but it still misses what seems like the point to me: these models don't do any actual reasoning. They're doing the same thing that DALL-E etc. does with images: using a superhuman store of potential outcomes to mimic an outcome that the person entering the prompt would then click a thumbs-up icon on in a training model. Asking why the model doesn't explain how the code it generated works is like asking a child who just said their first curse word what it means. The model and child alike don't know or care, they just know how people react to it.
- AndrewKemendo 4y agoI have $100 if you can prove that you do "actual reasoning" I'm not being snarky or sarcastic. This is a real challenge
- danShumway 4y agoWe can't prove that humans do actual reasoning, but we can be relatively certain that the current models don't (yes, I know there are some who claim otherwise, but it's just not particularly convincing). The debate over whether or not "reasoning" is just an emergent process isn't necessarily a bad one. But if it is an emergent process, there's little evidence that AlphaCode has emerged it. If "reasoning" does just boil down to repetition, then humans are repeating patterns on a much deeper and more sophisticated level than these models are, and it's still worth questioning if the model's current approaches and the strategies we're using to build AI will be good enough to get to that level of mimicry on their own without a lot more innovation. Philosophical zombies and perfect imitations of humans are interesting and useful thought experiments, but it's not necessary to prove that humans have qualia/reason in order to prove that much simpler algorithms and statistical methodologies do not, at least not at their current scale. And from a practical perspective, the distinction does matter, because we're trying to develop practical tools and it's useful to have a rough understanding of what the current models are and aren't capable of -- ie, they're not currently capable of explaining why they actually made certain decisions. That affects the way that we use them as practical tools. ---- I guess if it helps, sub out "actual reasoning" for "pattern matching to such an advanced degree that it can generate consistent responses that look like reasoning, and that are useful in the real-world as real-world explanations for why a choice was made, and that (importantly) can be communicated to humans, built upon, affirmed, or corrected in real-time." In other words, if non-pattern-matching logic/reason doesn't exist, fine, it could very well be an abstraction or an emergent property. But the machine isn't good enough at pretending to have that, and that limits its usefulness to scenarios where you don't care about it being able to imitate "actual reasoning". As an analogy: folders don't exist on your computer; it's all ultimately just bits on the drive and lookup tables. But there is still a really big practical difference between an operating system that has a hierarchical filesystem and an operating system that doesn't. In other words, abstractions can be purely imaginary, and it can still be useful to distinguish between things that adhere to those abstractions and things that don't. Similarly, even if "actual reasoning" is a made-up abstraction, there's still a big difference between an AI that can explain why a decision was reached and an AI that can't -- even if both AIs ultimately end up working the same way under the hood. As far as I can tell, AlphaCode is a fairly long ways away from hitting that point where we can even pretend to ourselves that it's thinking about what it's doing.
- AndrewKemendo 4y ago>it's still worth questioning if the model's current approaches and the strategies we're using to build AI will be good enough to get to that level of mimicry on their own without a lot more innovation Literally, and I mean Literally, nobody who is an AI practitioner, academic or philosopher (including myself) would disagree with this. What/who is making this claim?
- danShumway 4y agoWe might be arguing past each other. I don't know any (serious, at least in my mind) AI researchers that would make that claim, no. But it is a distinction that is poorly communicated to ordinary people who use these models, and most of the time that you encounter "how do we know that humans reason" in the wild in non-researcher contexts, it usually is being used to argue that we just need to throw more processing power at models and get more data into them -- or to paper over the real criticism that the current generative models that are available to ordinary people are not useful for tasks that require something approximating "actual reasoning." It's a hard thing to talk about, because I recognize that it is absolutely unfair to criticize AI researchers that are not making these arguments in the first place, but also recognize that pointing out this kind of distinction is absolutely necessary for the general public on a public forum and that it is important to continue to hammer home the practical difference between pattern-matching and reasoning in current models. I don't know how to walk that line other than to point out that trynewideas's comment was reasonable and that attacking it from a philosophical perspective misses the point of what they're saying. But I don't mean to imply that researchers are wrong or misinformed about the state of AI, I'm just pointing out that the philosophical distinction doesn't really change anything about the practical argument (since non-researchers also frequent HN).
- AndrewKemendo 4y agoThat's fair. Thanks for the write up.
- aerovistae 4y agofantastic analogy, A+ if you came up with that
- jujugoboom 4y agoStochastic Parrot is the term you're looking for https://dl.acm.org/doi/10.1145/3442188.3445922 https://dl.acm.org/doi/10.1145/3442188.3445922
- fumeux_fume 4y ago"Beating a dead horse" was the one I was thinking of
- dekhn 4y agoNorvig discusses this topic in detail in https://norvig.com/chomsky.html https://norvig.com/chomsky.html As you can see, he has a measured and empirical approach to the topic. If I had to guess, I think he suspects that we will see an emergent reasoning property once models obtain enough training data and algorithmic complexity/functionality, and is happy to help guide the current developers of ML in the directions he thinks are promising. (this is true for many people who work in ML towards the goal of AGI: given what we've seen over the past few decades, but especially in the past few years, it seems reasonable to speculate that we will be able to make agents that demonstrate what appears to be AGI, without actually knowing if they posses qualia, or thought processes similar to those that humans subjectively experience)
- jameshart 4y agoWe don’t know if humans possess qualia. I also don’t know if we should take humans’ word for it that they experience ‘thought processes’.
- dekhn 4y agothat's why I added the second clause: " thought processes similar to those that humans subjectively experience". Because personally I suspect that consciousness, free will, qualia, etc, are subjective processes we introspect but cannot fully explain (yet, or possibly ever).
- maweki 4y agoTuring said, that while you never know whether somebody else actually thinks or not, it's still polite to assume.
- jgilias 4y agoMaybe silly, but this is how I treat chatGPT. I mean, I don’t actually think it’s conscious. But the conversations with it end up human enough for me to not want to be an asshole to it. Just in case.
- johnfn 4y agoI'm not the first to say it, but the distinction over whether models do any "actual reasoning" or not seems moot to me. Whether or not they do reasoning, they answer questions with a decent degree of accuracy, and that degree of accuracy is only going up as we feed the models more data. Whether or not they "do actual reasoning" simply won't matter. They're already superhuman in some regards; I don't think that I could have coded up the solution to that problem in 5 seconds. :)
- nighthawk454 4y agoThis is a sort of dangerous interpretation. The point of saying model's "don't do reasoning" is to help us understand their strengths and weaknesses. Currently, most models are objectively trained to be "Stochastic Parrots" (as a sibling comment brought up). They do the "gut feeling" answer. But the reasoning part is straight up not in their objectives. Nor is it in their ability, by observation. There's a line of thought that if we're impressed with what we have, if it just gets bigger maybe eventually 'reasoning' will just emerge as a side-effect. This is somewhat unclear and not really a strategy per se. It's kind of like saying Moore's Law will get us to quantum computers. It's not clear that what we want is a mere scale-up of what we have. > Whether or not they do reasoning, they answer questions with a decent degree of accuracy, and that degree of accuracy is only going up as we feed the models more data. Kind of. They don't so much "answer" questions as search for stuff. Current models are giant searchable memory banks with fuzzy interpolation. This interpolation gives some synthesis ability for producing "novel" answers but it's still basically searching existing knowledge. Not really "answering" things based on an understanding. As long as it's right the distinction may not matter. But the danger is a "gut feeling" model will _always_ produce an answer and _always_ sound confident. Because that's what it's trained to do: produce good-sounding stuff. If it happens to be correct, then great. But it's not logical or reasonable currently. And worse, you can't really tell which you're getting just by the output. > Whether or not they "do actual reasoning" simply won't matter. Sure it will. There's entire tasks they categorically can't do, or worse can't be trusted with, unless we can introduce reasoning or similar. > They're already superhuman in some regards; I don't think that I could have coded up the solution to that problem in 5 seconds. :) This is superhuman in the way that Google Search is. You couldn't search the entire internet that fast either, but you don't think Google Search "feels the true meaning of art" or anything.
- gfodor 4y agoYour analogy is reaching to the farthest edge case - one of complete non-understanding and complete mimicry. The problem is that language models do understand concepts for some reasonable definitions of understanding: they will use the concept correct and with low error rate. So all you’re really pointing at here is an example where they still have poor understanding, not that they have some innate inability to understand. Alternatively, you need to provide a definition of understanding which is falsifiable and shown to be false for all concepts a language model could plausibly understand.
- deleted 4y ago[deleted]
- TreeRingCounter 4y agoThis is such a silly and trivially debunked claim. I'm shocked it comes up so frequently. These systems can generate novel content. They manifestly haven't just memorized a bunch of stuff.
- throw_nbvc1234 4y agoComing up with novel content doesn't necessarily mean it can reason (depending on your definition of reason). Take 3 examples: 1) Copying existing bridges 2) Merging concepts from multiple existing bridges in a novel way with much less effort then a human would take to do the same. 3) Understanding the underlying physics and generating novel solutions to building a bridge The difference between 2 and 3 isn't necessarily the output but how it got to that output; focusing on the output, the lines are blurry. If the AI is able to explain why it came to a solution you can tease out the differences between 2 and 3. And it's probably arguable that for many subject matters (most art?) the difference between 2 and 3 might not matter all that much. But you wouldn't want an AI to design a new bridge unsupervised without knowing if it was following method 2 or method 3.
- mrguyorama 4y agoChildren produce novel sentences all the time, simply because they don't know how stuff is supposed to go together. "Novel content" isn't a step forward. "Novel content that is valid and correct and possibly an innovation" has always been the claim, but there's no mathematical or scientific proof. How much of this stuff is just a realization of the classic "infinite monkeys and typewriters" concept?
- jvm___ 4y agoIn my head I picture these models like if you built a massive scaffold. Just boxes upon boxes, enough to fill a whole school gym, or even cover a football field. Everything is bolted together. You walk up to one side and Say "write me a poem on JVM". The signals race through the cube and your answer appears on the other side. You want to change something, go back and say another thing - new answer on the other side. But it's all fixed together like metal scaffolding. The network doesn't change. Sure, it's massive and has a bajillion routes through it, but it's all fixed. The next step is to make the grid flexible. It can mold and reshape itself based on inputs and output results. I think the challenge is to keep the whole thing together, while allowed it to shape-shift. Too much movement and your network looses parts of itself, or collapses altogether. Just because we can build a complex, but fixed, scaffolding system, doesn't mean we can build one that adapts and stays together. Broken is a more likely outcome than AGI.
- yesenadam 4y ago> it's massive and has a bajillion routes through it, but it's not fixed. I think you meant to write "but it's fixed."
- jvm___ 4y agoThanks
- urthor 4y ago> The next step is to make the grid flexible. But... we can do exactly that??? Model retraining is very transparent and a very straightforward way to redirect it. The core problem is that the analogies just don't hold up. Artificial Intelligence is not a problem that can be resolved via analogies and heuristics. Because artificial intelligence simply looks nothing like what our human heuristics imagine it to be. It can only be solved by understanding the actual, genuine detail.
- jvm___ 4y agoIs there a 3d structure that you think is more analogous? From my understanding it's all matrix math, which to me is a grid. Maybe not a scaffolding one, and more like a bundle of interconnected wires so that each concept has multiple related concepts, but I think a rigid grid is easier to explain to lay people.
- eternalban 4y agoA language model does not have to reason to be able to produce textual matter corresponding to code. For example, somewhere, n blogs were written about algorithm x. Elsewhere, z archives in github have the algo implemented in various languages. Correlating that bit of text from say wiki and related code is precisely what it has been doing anyway. Remember: it has no sense of semantics - it is “tokens” all the way down. So, the fact that you see the code as code and the explanation as explanation is completely opaque to the LLM. All it has to do is match things up.
- PaulHoule 4y agoI'm skeptical of "explainable A.I." in many cases and I use the curse words as an example. You really don't want to tease out the thought process that got there, you just want the behavior to stop.
- olalonde 4y ago> This is a great review but it still misses what seems like the point to me: these models don't do any actual reasoning. Hmmm... I have seen multiple examples of ChatGPT doing actual reasoning.
- teaearlgraycold 4y agoI've also gotten it to make very obvious logical errors. It reasons by chance. Better than a dice roll but there is no internal knowledge graph.
- gonehome 4y agoWhat is this then: https://twitter.com/jbrukh/status/1603868836729610250?s=46&t=D7CSxo2_U5BiWwa6FWKtQw https://twitter.com/jbrukh/status/1603868836729610250?s=46&t... That looks a lot like reasoning to me. At some point these disputing definitions arguments don’t matter. Some people will endlessly debate whether other people are conscious or “zombies” but it’s not particularly useful. This isn’t yet AGI, but the progress we’re seeing doesn’t look like failure to me. It looks like what I’d predict to see before AGI exists.
- nebukhadnezzar 4y agoThese arguments do matter since the system often confidently <<reasons>> incorrectly, and people still incorrectly believe that the system actually deduces semantic meaning behind it's output. The system has no grounding in truth whatsoever. It's impressive nonetheless, but it's still a large lookup table for pattern matching. I'm convinced bruteforcing these LLMs won't give us anything near AGI. We're going to need to start funding more alternatives too, which often are shadowed by these large bruteforce models.
- gonehome 4y agoMaybe we also operate as a large lookup table for pattern matching (with some initial weights set by our evolutionary history). We have no grounding in truth either (and often confidently "reason" incorrectly). I’m not so sure there’s a major missing piece here beyond some sort of continually running “default mode network”.
- danShumway 4y ago> These arguments do matter since the system often confidently <<reasons>> incorrectly, and people still incorrectly believe that the system actually deduces semantic meaning behind it's output. This is the point that gets missed in these conversations about AGI. I have had programmers who should absolutely know better argue to me that correlation/causation distinctions don't matter in AI because as long as you're not over-fitting to the data, it's safe to draw conclusions based on the correlation alone -- the AI wouldn't point out the correlation if it wasn't a safe correlation to rely on (and it'll just correct itself with new data if that ever changes). A lot of people even in technical fields have way too much trust in these systems and way too much magical thinking about how they work. I don't blame researchers for that, but the argument "how do you know you're not a lookup table" is unhelpful in getting across to people that image classifiers have been hacked by writing the word "apple" across a bicycle, and that code generators are far better at generating code that imitates training data than they are at generating code that is safe for high-security environments based on 1st principles of security. And it's honestly a little weird, because "aren't we all just lookup tables" is an argument that personally makes me a lot more cautious about human reasoning, but for a lot of other people that argument doesn't make them trust humans less, it makes them trust current models more. So there really is an education need to explain to people how the current AI models work in practice -- that they aren't deriving semantic meaning (at least not in a practical sense, not right now). And a lot of people understand that, but a lot more people don't. The debate over how to get to AGI matters a lot less to me than the non-experts that seem to on some level believe that we've already reached a shallow form of AGI.
- deleted 4y ago[deleted]
- api 4y agoI think of them as a new type of database with remarkable query and data manipulation capabilities. They can’t produce things that go fundamentally outside their training data set.
- deleted 4y ago[deleted]
- bananamerica 4y agoWhat is reasoning?
- ninethirty 4y agoSeems like this is an explanation? https://github.com/JusticeRage/Gepetto https://github.com/JusticeRage/Gepetto
- nl 4y agoThis is untrue. Minerva in particular does step by step reasoning: > we present Minerva, a language model capable of solving mathematical and scientific questions using step-by-step reasoning https://ai.googleblog.com/2022/06/minerva-solving-quantitative-reasoning.html?m=1 https://ai.googleblog.com/2022/06/minerva-solving-quantitati... They use a technique called chain-of-thought to make it reason problems out step by step. As Norvig shows, this reasoning may sometimes be wrong, but that's a different issue. It's worth repeating this because apparently it's not widely known: these models are following a step-by-step process we'd recognise as reasoning in humans.
- reaperducer 4y agothese models don't do any actual reasoning This gives us the new all-purpose excuse for not commenting or documenting our code: "The A.I. made it."
- kiratp 4y agoGiven that we are able to prompt GPT (davinci) to learn proprietary DSLs that it would have never seen and build programs with it on data it has never seen… I find it really difficult to buy that there is no internal representation of logic in the model hidden states.
- teaearlgraycold 4y agoGPT-3 has seen examples of code with comments describing a DSL followed by use of that DSL. And DSLs rarely invent a new paradigm. I think if you showed a model OOP programming when it had only been trained on functional code it would fail spectacularly. But your DSL is probably not far off from Lisp or C, or YAML. The model correlates a pre-defined set of tokens with other tokens in the input and then extrapolates. The only reason it can appear to have internal logic is because of the consistent structure of code it was trained on.
- galaxyLogic 4y agoYou hit the nail in the head. It is about mimicry. Apes are often dismissed as just mimicking human behavior in many circumstances. Putting on a hat for instance. They don't know nor think about what it is about. And neither do babies nor even many adults. They don't reflect on it, and most definitely they don't reflect on their own thought-processes. They just mimic other humans around them. Now mimicking the speech of people may give the impression that the AI has some thought-processes behind its speech because, it is text that any human might write. But there's no logic because there's no symbolic reasoning. There's just mimicry and trial and error. "Training the model" just means millions of trial-and-error practice runs. What the (neural network) AI can NOT mimic is the actual (symbolic) thought-process of humans, because that is not visible anywhere. It can only mimic the input -> response -behavior of humans. Therefore no, the AI is not conscious. It is the proverbial zombie.
- pishpash 4y agoAre you sure there is an actual "thought-process [sic] of humans"? Witnesses have made up explanations for false memories, the same as in a dream.
- galaxyLogic 4y agoI am not sure. But humans can often explain the thought-process that led them to a certain decision. I think that counts as evidence of consciousness., in humans.
- skirmish 4y agoHumans can explain rational thought, sure, but not subconscious or acting on impulse. When you say something you shouldn't have said, how do you explain the thought process that led you to saying it? Were you an unconscious actor when you unwisely started an argument with your manager?
- galaxyLogic 4y agoYes you can be conscious about having been unconscious earlier. We do have reflexes. But we also have consciousness about our own thinking.
- version_five 4y agoThe way I think about GPT et al in terms or their "intelligence" is the same as the original Eliza chat that did some simple substitution based on patterns in the input. Gpt is obviously more complex and very large scale in its patterns but fundamentally it's exactly the same kind of "parlor trick". I mention this in reply because I think it would be a funny exercise to imagine Eliza as a black box and attempt to critique some of its answers.
- mannykannot 4y agoI feel there may be an important clue here, at least with regard to the programming example. Peter asked Kevin Wang for some insights into his approach to programming competitions, and one was this: [Kevin] I think specifically for AoC it's important to just read the input/output and skip all the instructions first. Especially for the first few days, you can guess what the problem is based on the sample input/output. [Peter] Kevin is saying that the input/output examples alone are sufficient to solve the easier problems; in the interest of speed he doesn't even read the problem description. From what I have read elsewhere, it seems that AlphaCode generates a great many candidate programs, and uses the given test cases to eliminate those which fail them. I wonder if that alone (without any reasoning from the problem statement to a solution) can account for its success rate on these challenges; Kevin's observation suggests this might be plausible. On the other hand, if the questions posed to AlphaCode are essentially as we see them, then its ability to use the test cases to screen its candidates seems impressive to me, just by itself (if it is a result of training, rather than an explicitly programmed behavior.)
- erostrate 4y agoCan you think of a test or empirical observation that would convince you that a model does "actual reasoning"? Do bear in mind the tests that very smart people have proposed in the past (from playing chess to holding a conversation to understanding pictures). And consider the implication for your position if you find it hard to devise a well defined test that would convince you to abandon this position.
- jpttsn 4y agoHow I'd do it: 1. Define reasoning (hard) 2. Inspect the system to see if this is what it's doing (hard) Conclusion: too hard, so until I get better I don't know if any given thing is reasoning. What "very smart people have proposed": 1. Look at something I've only seen people-I-assume-to-be reasoning do 2. Run the system and see if it can do this 3. If satisfied or in the dictionary business, redefine reasoning 4. goto 1
- logicallee 4y ago>This is a great review but it still misses what seems like the point to me: these models don't do any actual reasoning. Why don't you come up with your own test for this, something unique that wouldn't be in the text corpus of the Internet so far, and then try it on OpenGPT and see if it passes your own test. Try it and see if it is able to reason. I couldn't come up with my own so I decided to ask ChatGPT to come up with one. This is my prompt: "Today in order to test whether an AI is able to reason, or can only apply text rules to existing text corpuses, we're going to come up with a new test for reasoning, some unique reasoning situation that isn't already in the text corpus of the Internet. Since the text corpus of the Internet is enormous, with over 50 billion pages on the Internet, our most important task for a successful outcome is to come up with a random new reasoning exercise that is extremely unlikely to already be on the Internet in any form. Therefore, we want to create something so unique it is unlikely to exist in 50 billion pages. To this end, I think it would be good to pick a unique situation based on a combination of elements that are unlikely to have appeared in that format; i.e. a very unique situation. Pick no fewer than 20 elements (objects concepts etc) that are unlikely to have ever been combined in the particular way you come up with, and use them to come up with a unique situation that gives room to reason about the situation. Then, describe this unique situation and ask an open-ended reasoning test and provide an example of a correct answer, and an incorrect answer. Finally, print your own answer to the test and then score whether you have reasoned about the situation you came up with correctly or not. Give yourself a score between 0% if you don't understand the situation at all, up to 100% if you perfectly understood the situation you came up with and reasoned about it perfectly. Also give yourself a score on originality from 0% completely unoriginal to 100% fully original depending on how likely you think the exact situation you came up with is to have already been mentioned on the Internet." Here is ChatGPT's answer (at the bottom after this I continue my reply): --- Here is a unique reasoning exercise I came up with: Situation: You are a chef at a fancy restaurant and you have been asked to prepare a special dish for a vegan customer. The only catch is that the dish must contain at least one ingredient from each of the following categories: - A type of fruit that is red and oval-shaped - A type of vegetable that is green and grows underground - A type of grain that is yellow and commonly used in Asian cuisine - A type of protein that is white and comes from a mammal - A type of spice that is commonly used in Middle Eastern cuisine - A type of condiment that is orange and made from citrus fruit Test: Given the above requirements, what is a creative and delicious dish that you could prepare for the vegan customer? Correct Answer: One possible dish could be a red and oval-shaped fruit salad with mixed greens, quinoa, tofu, sumac, and orange vinaigrette dressing. Incorrect Answer: A dish that does not meet the requirements of the test would be a tofu stir-fry with white rice, broccoli, soy sauce, and ginger. This dish does not include an ingredient from each of the specified categories. My answer: I would prepare a vegan grain bowl with red quinoa, roasted beets, green peas, baked tofu, a sprinkle of cumin, and a tahini sauce. Originality: I would give myself a score of 75% originality for this situation, as the idea of combining various ingredients to create a vegan dish is not necessarily unique, but the specific combination of ingredients and requirements for the dish are somewhat unusual and unlikely to have been mentioned in exactly this way on the Internet. Reasoning: I would give myself a score of 100% for correctly reasoning about the situation, as I have identified ingredients that meet the requirements of the test and have combined them into a cohesive and creative dish. " -------------- That's pretty good! Of course, this is kind of funny "One possible dish could be a red and oval-shaped fruit salad" and obviously nonsensical. However, it didn't include it in its own answer. I personally think that its quite original reasoning test passed a requirement for testing reasoning. I also think it passed its own test. Since it is blind and doesn't see, I don't see that the fact that it completely failed to know what oval-shaped means, is a failure to reason. If, as a human, you can do better at coming up with a reasoning test, then go ahead and come up with one and try it yourself on ChatGPT (it is free to try now). For me, ChatGPT showed an ability to actually reason.
- xorcist 4y ago> the person entering the prompt would then click a thumbs-up icon on What you are saying is that huge amount of resources was just spend on building the biggest like-farming machine there ever was. I quite like that viewpoint. It's also a dark reminder what this technology is commercially well suited for.
- thundergolfer 4y agoAlways a pleasure to read Norvig's Python posts. His Python fluency is excellent, but, more atypically, he provides such unfussy, attentive, and detailed explanations about why the better code is better. Re-reading the README, he analogizes his approach so well: > But if you think of programming like playing the piano—a craft that can take years to perfect—then I hope this collection can help. If someone restructured this PyTudes repo into a course, it'd likely be best Python course available anywhere online.
- abecedarius 4y agoHe did give a course "Design of Computer Programs" https://www.udacity.com/course/design-of-computer-programs--cs212 https://www.udacity.com/course/design-of-computer-programs--... using Python. I'd recommend it too. (Most of the pytudes came later.) (disclosure: I helped a little bit when he developed the course.)
- myle 4y agoHe used to have an excellent course in Udemy or such, which was using Python.
- deleted 4y ago[deleted]
- bombcar 4y agoThe marvel is not that the bear dances well, but that the bear dances at all. The surprising thing is that it can make code that works - however, given that code can be tested in ways that "art" and "text" cannot (yet), perhaps it's not that strange.
- dekhn 4y agoI read this as a very well written feature request to the AlphaCode engineers (or anybody working on this problem). I really like Peter's writing style. It's fairly clear, and understating, while also making it quite clear there are areas for improvement in reach. For those who haven't read it, Peter also wrote this gem: https://norvig.com/chomsky.html https://norvig.com/chomsky.html which is an earlier comment about natural language processing, and https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/35179.pdf https://static.googleusercontent.com/media/research.google.c... which is a play on Wigner's "Unreasonable Effectiveness of Mathematics in the Natural Sciences".
- MoSattler 4y agoI wasn’t aware AI can already take plain English text and create functioning software. I guess it’s time to look for another profession.
- ResearchCode 4y agoWhat other profession? When and if it can actually write code, making a ppt or sitting in a pointless meeting is a long solved problem.
- deleted 4y ago[deleted]
- weatherlite 4y agoAll the low status professions we've been told are not that good - child care, elderly care, handy work etc ... Yeah we're pretty much toast :)
- fergal_reid 4y agoHuge respect for Norvig, but I think this is a shallow analysis. For example, I just took Norvig's 'backspacer alpha' function and asked ChatGPT about it. It gave me an ok English language description. It names the variables more descriptively on command. I'm sure it'll hallucinate and make errors, but I think we're all still learning about how to get the capabilities we want out of these models. I wouldn't rush to judgement about what they can and can't do based on what they did; shallow analysis can mislead both optimistically and pessimistically at the moment!
- neontomo 4y agoYou are talking about a completely different language model though, so your argument doesn't really hold. The point he makes is that AlphaCode does not generate meaningful explanations along with its code (and that's it's perhaps necessary to build trust in a system), not that there aren't any language models that can do it. AlphaCode !== ChatGPT. I encourage you to write a better analysis!
- fergal_reid 4y agoWell, the stall he sets out is about 'generative models' generally: "In the future, what role will these generative models play in assisting a programmer or mathematician?" I think the fact that we've other generative models that eg. can write descriptive variable names is a fair rebuttal to his criticism of alphacodes poor variable names. Deepmind fine tuned it on competition code, not on well documented production code, so I think it's shallow to get hung up on it's lack of good variable names (for example).
- RodgerTheGreat 4y agoI think the notes at the end bury the lede; in particular: > "I save the most time by just observing that a problem is an adaptation of a common problem. For a problem like 2016 day 10, it's just topological sort." This suggests that the contest problems have a bias towards retrieving an existing solution (and adapting it) rather than synthesizing a new solution.
- mrguyorama 4y agoThe fact is, the vast majority of programming IS just dredging up a best solution and modifying it to meet your specifics. Some of the best and still most current algorithms are from like the 60s. That doesn't make neural networks "smart", and instead says more about our profession and how terrible we in general are at it.
- p-e-w 4y ago> and instead says more about our profession and how terrible we in general are at it. I'm well aware that belittling one's own industry is part of contemporary software engineering culture, but for the love of all that is holy, I've never understood it. A few decades ago, "software engineering" didn't exist. Today, it's a hair's breadth from creating an artificial being. That's closer to magic and alchemy than to building rockets and bridges. It just blows my mind that humans have achieved this, and it bores and saddens me to be surrounded by engineers who keep regurgitating "everything sucks" memes.
- urthor 4y agoA very good point. People trivialize the progress in this progression, because it's normalized. And yet, the incredible progress in the last 5-10 years is absolutely stunning. Compared to the other "recognized professions," the progress in how software is developed is beyond stunning. The actual practice of law/medicine/etcetera is remarkably static. Regardless, to most, "what we have now is never enough."
- 4y ago
- tareqak 4y agoWhen I saw the test suite that Peter Norvig created for the program, I immediately thought to myself “what if there was a LLM program that knew how to generate test cases for arbitrary functions?” I think a tool like that even in an early incomplete and imperfect form could help out a lot of people. The first version could take all available test cases as training data. The second one could instead have a curated list of test cases that pass some bar. Update: I thought of a second idea also based on Peter Norvig’s observation: what about an LLM program that adds documentation / comments to the code without changing the code itself? I know that it is a lot easier for me to proofread writing that I have not seen before, so it would help me. Maybe a version would simply allow for selecting which blocks of code need commenting based on lines selected in an IDE?
- hgsgm 4y agoQuickCheck (not NN) generates test cases. GPT can already comment on code.
- kubb 4y agoThat's potentially more helpful than writing the code itself. Writing unit tests can take most of the development time.
- mrguyorama 4y agoAnd literally throwing random half junk unit tests at your code will better test it than you writing unit tests that are blind to the problems it might have because you wrote both and both bits of code have the same blind spots. We should probably be developing systems that fuzz all code by default.
- SlickNixon 4y agoIn the future, code beyond all human comprehension will be rated for reliability by the number of layers of garbage unit tests it's buried under. Tests for the code, then tests for the tests, then tests for those te-...
- 4y ago
- happyopossum 4y ago> I find it problematic that AlphaCode dredges up relevant code fragments from its training data, without fully understanding the reasoning for the fragments. As a non-programmer who has to ‘code’ occasionally, this is literally what I do, but it takes me hours or days to hammer out a few hundred lines of crap python. Using a generative model or llm that can write equally crappy scripts in seconds feels like a HUUUGE win for my use cases.
- peteradio 4y agoA lazy ineffective person is preferable over a prodigious idiot.
- lioeters 4y agoReminds me of the quote: > There are four types of officer: the clever and lazy, the clever and industrious, the stupid and lazy, and the stupid and industrious. > The clever and lazy you make Chief of Staff, because he will not try to do everybody else’s work, and will always have time to think. > The clever and industrious you make his deputy. > The stupid and lazy you put into a line battalion, and kick him into doing a job of work. > The stupid and industrious you must get rid of at once, because he is a national danger.
- happyopossum 4y agoNot sure which one you’re calling me The reality is that there’s a huge gap of scripting needed in todays world that sits between “print ‘hello world’” and something worth paying a professional SWE to write. You folks don’t always see this from inside your bubble, but there is a world of automation and computing problems that are solved by people who can make it work, and not by people paid to make it work well.
- hgsgm 4y ago
- lisper 4y agoIs it just me, or is that problem description completely incoherent?
- krackers 4y agoIt took me way too long to understand it as well. And the fact that you press backspace _instead_ of a character, instead of allowing backspace to be pressed at any time (which would turn it into checking if B is a subsequence of A I believe).
- zug_zug 4y agoGood god I'm not alone. For me it's fascinating that an AI can make sense of that garble of words. I spent 4 minutes trying to read it and gave up.
- drexlspivey 4y agoProblem: You open a terminal and type the string ‘ababa’ but you are free to replace any button presses with backspace. Is there a combination where the terminal reads `ba` at the end?
- gok 4y agoThey are vulnerable to reproducing poor quality training data They are good locally, but can have trouble keeping the focus all the way through a problem They can hallucinate incorrect statements does not generate documentation or tests that would build trust in the code These observations are about human programmers, right?
- bamboozled 4y agoDo you often hallucinate on the job?
- Gigablah 4y agoTends to happen when writing self-assessments.
- bamboozled 4y agoha
- kqr 4y agoYou may be saying it tongue in cheek, but I think you have a point. After interviewing hundreds of applicants for developer positions, I'm sure the AlphaCode performance exhibited here is similar to at least the 50th percentile. If you're working in a large organisation where 50 % of developers perform roughly on par with AlphaCode, you'll see it as a potentially useful companion. If you're Peter Norvig, or working in a small organisation trying to hire the top 20 %, you'll just laugh at AlphaCode.
- chubot 4y agoSomewhat dumb question: I wonder what tool he used for the red font code annotations and arrows? What tool would you use, like Photoshop or something? And just screenshot the code from some editor or I guess Jupyter?
- jdcarter 4y agoThe red handwriting font is exceptionally good; I'd love to know what that is. It looks very natural, I assume with a ton of ligatures.
- circuit 4y agoMost likely Preview.app's built-in annotation tools
- chubot 4y agoThanks, that's probably it! (I need to do stuff like this, but I don't use a Mac regularly, and that would help)
- ipv6ipv4 4y agoAlphaCode doesn't need to by perfect, or even particularly good. The question is when AlphaCode, or an equivalent, is good enough for a sufficient number of problems. Like C code can always be made faster than Python, Python performance is good enough (often 30x slower than C) for a very wide set of problems while being much easier to use. In Norvig's example, the code is much slower than ideal (50x slower), it adds unnecessary code, and yet, it generated correct code many times faster than anyone could ever hope to. An easy to use black box that produces correct results can be good enough.
- ALittleLight 4y agoIt seems like it would be easy enough to train the model to improve performance too once you have it writing correct code. Write N problems and N performance test suites - then train by self-play on those problems until you're writing high performance code.
- alar44 4y agoAbsolutely. I've been using it to create Slack bots over the last week. It's cuts out a massive amount of time researching APIs and gives me good enough, workable, understandable starting points that saves me hours worth of fiddling and refactoring.
- TheRealPomax 4y agoBut as an AI implementation, when considered "as an example of an AI implementation", an analysis of its behaviour by one of the foremost experts in the field is still entirely worth the effort, because we learn from that analysis. It teaches us what an expert in the field sees happening, where and how it's falling short, and so what kind of improvements would be worth continued focus on. Criticism by a lay person is a blog post. Criticism by Peter Norvig is a teaching moment that future work is based on. It's like having Knuth comment on your fundamental algorithm design: there is gold in them there text, and we got it for free.
- weatherlite 4y agoIt's good enough for a human to go over it and change it (maybe change it in a big way even). This wouldn't have passed code review probably. Its still probably quite a boost for productivity but it doesn't take the human out of the equation. I can also imagine scenarios where relying too much on generated code without understanding it will cause massive bugs and headaches, just like copy pasting stuff from Stackoverflow without understanding it.
- aidenn0 4y agoThe Minerva geometry answer looks like something one of my kids would have written: guess the answer then write a bunch of mathy-sounding gobbledygook as the "reasoning." Also, that answer would have gotten 4/5 points at the local high-school.
- neilv 4y ago> They need to be trained to provide trust. The AlphaCode model generates code, but does not generate documentation or tests that would build trust in the code. I don't understand how this would build trust. If they generate test cases, you have to validate the test cases. If they generate documentation, you have to validate the documentation. For a one-shot drop of code from an unknown party, test cases and docs have been signals that the writer know that's a thing, and they at least put effort into typing it. So maybe we assume more likely that they also used good practices with the code. But that's signalling to build trust, and adding those to build trust without addressing the reasons we shouldn't have trust in the code (as this article points out) seems like it would be building misplaced trust. (Though there is some benefit to doc for validation, due to the idea behind the old saying "if your code and documentation disagree, then both are probably wrong".)
- low_tech_love 4y agoI think it’s indeed a bit weird to “require” the AI to generate test cases but not impossible, actually it might be even desirable. The thing is that you don’t need to necessarily validate the test cases themselves, as long as the method to generate them is well known and trustworthy. A good suite of test cases (manually or automatically generated) will never guarantee anything, but it is certainly very useful, and I believe an AI could generate much more and better test cases than we do, as long as it is trained correctly. For example, you can have reinforcement learning with rewards for finding test cases that fail, and it could even be an adversary to the other AI that is generating the code itself…?
- sdwr 4y agoJesus Christ, that's some code review! Felt the meaty heft of experience just glancing at the footnotes at the bottom.
- khiqxj 4y ago> single letter variables i hate when people put descriptive names on variables that dont need it. when im tired i end up reading the name over and over as if it has meaning (it doesnt), and subconsciously ponder the meaning for up to a minute before snapping out of it. they actually dont even serve a purpose. its because we are reading code in text form that they are needed. conceptually, most variables are just something flowing out of one function to the next which would be easily repesented by lines in a diagram or many other ways that dont use names. you're meant to only give names to the rare case where it would truly make the code more understandable. this shit doesnt: thingManagerProxyGenerator = createThingManagerProxyGenerator()
- preommr 4y agoI usually don't like commenting on people's choices for variable names since they're usually scoped and less important than the interface. However, I'll give my two cents because I remember stuff like this being confusing when I was younger. There's definitely a balance to be struck - one letter variable names are horrible outside of single lines (e.g. a one-line anonymous function), extremely long names are better, but still bad. I say better because it's higher in comprehensibility - which is more important - if not very readable. While it's been well documented that shorter information is easier to remember, it's important to remember that it's usually limited to brevity that still provides context. Very short names e.g. (pg for proxy generator) are bad because minds require an additional step to "unravel" the meaning. IMO it's usually best to limit to variable names to one or two words. Something like 'proxyGenerator', for the given example. Or even 'generator' if the parent block is extremely small (e.g. 5 lines). Most readers perform word chunking such that something like 'proxyGenerator' reads as two units without the indirection that an abbreviation or unrelated symbol would require.
- nvrmnd 4y agoThis is also a very interesting concept/question when it comes to AI writing software for practical applications. Would a large language model benefit from descriptive variable names when it comes to code understanding, or would this just be for human benefit? My first thought was that a computer (e.g. compiler) does not care what the token is named. But a large language model may actually benefit from it, certainly chatGPT is sensitive to it. But then, doesn't it mean that it could be fooled by misleading variable names (as a human would), something we would criticize any system for. "Alpha*** is easily fooled by changing variable names, making it completely unusable for blah blah".
- WheelsAtLarge 4y agoCode quality will be an issue since these apps are using existing code as the model. They are the very definition of trash in, trash out.
- deleted 4y ago[deleted]
- xwdv 4y agoIn its best form, I think AI coders in our career’s lifetime will never be anything more than a way to copy and paste code from StackOverflow with less steps. If you don’t mind searching for solutions and adapting them yourself, then you can already be living in the future, today!
- drpixie 4y agoI see only a single example discussed. It's interesting but begs the question - if I give it 1000 problems, how many of its results are correct programs?
- kqr 4y agoThat question is mainly interesting in comparison to the baseline, which ought to be what the same number is for new hires at your organisation!
- DonHopkins 4y agoI love his work and writing and respect him, and I only suggest this in the fondest possible way, but did anybody else ever notice that Silicon Valley's Professor Davis Bannercheck kind of looks like Peter Norvig? ;) https://www.youtube.com/watch?v=1KaWPYOLuT8 https://www.youtube.com/watch?v=1KaWPYOLuT8
- dinvlad 4y agoFunnily enough, I entered a hard Leetcode problem into ChatGPT the other day, and it was able to "solve" it!
- flatline 4y agoI am hoping this leads the industry to move away from its hyperfocus on solving algorithms for tech interviews. Sure they are important but are only one aspect of most jobs. The problems themselves are formulaic and not particularly indicative of any real acumen.
- killjoywashere 4y agoIncredibly useful to read code reviews. I rather doubt I will ever get to see another code review by Norvig. Can anyone point us to to others?
- ryan93 4y agoImpressive the chatgpt could understand that backspace problem. I’m drooling on myself trying to understand it.
- gfd 4y agoIsn't `a=a[2:]` going to copy and make that linear time too? I'm surprised he caught the `b.pop(0)` bug but then suggested doing that afterwards. Also he should submit to codeforces to get a comparable running time. For example here is a random person who submitted AlphaCode's version and it takes 1091 ms https://codeforces.com/contest/1553/submission/144971343 https://codeforces.com/contest/1553/submission/144971343
- PartiallyTyped 4y ago> Isn't `a=a[2:]` going to copy and make that linear time too? Afaik it's a memcpy of pointers.
- SekstiNi 4y agoThat would indeed be linear time, but he doesn't make that suggestion as far as I can see.
- gfd 4y agoIt's in the red inked code review image and it's in his rewrite as "source = source[:-2]"
- tuke 4y agoNo comment along the lines of: "ew, Python."
- pfedak 4y agoOne aspect of the solution which I haven't seen touched on is that this is a problem for which a straightforward greedy approach works. I know when I do competitive programming, "will greedy work?" is usually the first question; greedy strategies are natural (here there's the insight of going backwards, true), and often it's faster to just implement it and submit than actually prove it to myself. The output doesn't give me much confidence that, if we tweaked the problem slightly to make the greedy approach not work (e.g. certain characters/positions can't be replaced with backspace), we'd still get a working solution out. We might get something that looked extremely similar to this, but didn't actually solve the problem. As I said, this isn't too different from what I'd consider human performance on these problems, but it makes being able to trust the output (and the model's confidence in the output) even more important.
- cratermoon 4y agoOh yes, AlphaCode could definitely replace some of the clowns I've worked with.
- gus_massa 4y agoAbout the single letter name for variables, I guess it's bias because they are using as the second part of the training set the code from a few online programing competitions. Many of them have the time to write the program as a tiebreaker, so using short names is an advantage. Also, for short throwaway programs, using the single letter names avoids many silly tipos, like print(important_list_1, important_list_1, important_list_3) instead of print(important_list_1, important_list_2, important_list_3)
- pifm_guy 4y agoI'm gonna guess that a few people like Peter taking alphacodes output and editing it and using those edits as part of some fine-tuning step would quickly resolve many of these issues.
- low_tech_love 4y agoThe most scary thing for me about this whole thing is how it feels like programming as we know it is becoming obsolete. Yeah, there is a list c that’s never used, and that bothers me, but the real question is: should anyone care? When looked at from the perspective of the trade-off, i.e. code generation from unstructured natural language input, the answer is (unfortunately) no——and that is of course ignoring the fact that such improvements are incremental and will come very soon anyway. Picking on this is like what a programmer from the 60’s would say if they looked at the binary of the CPython interpreter, or something like that. Yeah you could probably do better, but it would take maybe 100x longer (or more), and the performance impact in the executable would be negligible. So yeah Peter we feel you, this is scary…
- lukeplato 4y agoSeems similar to non-coding "Junk" genomic DNA [0]. Some evolutionary process can improve and reduce programs but most cruft is inconsequential and may even be preferable to having perfectly reduced/succinct programs. [0] https://en.wikipedia.org/wiki/Non-coding_DNA https://en.wikipedia.org/wiki/Non-coding_DNA
- karenmichael889 4y ago