16 ms·
GPT-3 can run code
- ivegotnoaccount 5y ago> For example, it seems to understand how to find a sum, mean, median, and mode. > Input: 1, 4, 5, 6, 2, 1, 1 > Output: 2.28571428571 Well, even with those small numbers, it's wrong. The first "2" after the dot should not be there. The result it gives is 16/7, not 20/7.
- loganmhb 5y agoI wonder how much of this is an illusion of precision that comes from pattern matching on content from filler sites like https://www.free-hosting.biz/division/16-divided-7.html https://www.free-hosting.biz/division/16-divided-7.html (I do not recommend clicking the link, but the result appears there).
- ivegotnoaccount 5y agoI was thinking the same thing, especially as we are talking about division, and the result is "correct" for 16/7 to a great number of digits. See also the "x = x + x three times", for which the result is not random but the result for the same thing... Two times instead of three (so result/2). That heavily smells like it has read sites that had nearly the same code on them.
- kaetemi 5y agoIt has a ton of programming books in its training data. It only "runs" anything that's close enough to any samples it has seen that included output. Anything complex, and it fails, because it does not reason about it logically. It's bad at the same things humans are bad at.
- mr_toad 5y agoHuman programmers rely on intuition and experience much more than some people give them credit for. An experienced programmer can find common errors quickly, simply because they’ve seen (and made) so many. Being able to intuit what a block of code does is actually a core skill; having to actually step through code in your head is slow and difficult.
- kevincox 5y agoI find it quite interesting that in the JSON to YAML example it reordered the list. If this was an access control list that could be a serious security issue that could have easily been missed in review. (Especially if dozens of files like this were changed at once). Of course a malicious user could have done this as well and likely got by code review but the fact that it was accidental is scarier in a way.
- daenz 5y agoNit, but YAML is a superset of JSON, so no conversion required :)
- jefftk 5y agoThis sort of "do what I mean" situation, where doing the thing the user intended is different from doing something technically correct, is a place GPT-3 excels. Even though returning the input would be easiest, it has the pragmatic judgement to predict that's not what the user wants.
- kcorbitt 5y agoFor folks wanting to play around with the GPT-3 code-editing capabilities referenced in the article within your own codebase, I wrote a simple open source VS Code plugin that lets you run commands against your currently-open file and get GPT-3's suggested edits back in a diff: https://marketplace.visualstudio.com/items?itemName=clippy-ai.clippy-ai https://marketplace.visualstudio.com/items?itemName=clippy-a...
- thrtythreeforty 5y ago> GPT-3 struggles with large numbers, decimal numbers, and negative numbers. When used it returns answers that are close but often incorrect. Regarding GPT-3's "guesstimates," intuitively it feels like the network has to guess because it hasn't been given a way to do exact computation--a neural network is built out of nonlinear functions--even if it "understands" the prompt (for whatever value you want to give to "understand"). Are there any techniques that involve giving the model access to an oracle and allowing it to control it? To continue the analogy, this would be the equivalent of giving GPT-3 a desk calculator. If this is a thing, I have other questions. How do you train against it? Would the oracle have to be differentiable? (There are multiple ways to operate a desk calculator to evaluate the same expression.) Also, what control interface would the model need so that it can learn to use the oracle? (Would GPT-3 emit a sequence of 1-hot vectors that represent functions to do, and would the calculator have "registers" that can be fed directly from the input text? Some way of indirectly referring to operands so the model doesn't have to lossily handle them.)
- ravi-delia 5y agoI believe the dominant thinking is that GPT-3 has trouble with math because it doesn't see individual digits. It obviously has no trouble working on words, which are much more discreet than numbers. I wouldn't be surprised if it had trouble carrying a long equation though. When writing it can reconsider the whole context with each new word, externalizing that memory, but with most computations it would have to carry out the whole thing in one go. That's a lot of dedicated parameters for a single subtask.
- thrtythreeforty 5y ago> with most computations it would have to carry out the whole thing in one go Is there a way to allow models to say "let me think about this some more"? With language models like GPT-3 you emit one token per inference iteration, with its previous output fed back in as input/state. Can models opt out of providing a token, but still update state? That would allow it to break up the computation into discrete steps.
- Veedrac 5y ago> GPT-3 seems to have issues with large numbers. Moyix’s gist covers this in detail. GPT-3 tends to guesstimate an algebraic function instead of evaluating the numbers, so the answer is only correct to a certain approximation. There are two issues here. One is the lack of working memory, which means that there is very little scratch space for calculating things with a meaningful sequential depth. GPT-3 is very unlike traditional evaluation methods in this regard, in that it is easier for it to interpret the meaning of a program you give it and then intuit the result given the context than it is to mechanically execute its steps. The other issue is the text encoding, which makes it much harder for GPT-3 to do digit-by-digit operations. Many arbitrary numbers are just their own token. A fixed length number to us looks like a fixed number of characters, but for GPT-3 they can be and almost arbitrary number of tokens divided into almost arbitrary chunks. Using thousands separators is very helpful for it. If you account for these and design a prompt that mitigates them you can get much stronger results. Here is an example: https://news.ycombinator.com/item?id=30299360#30309302 https://news.ycombinator.com/item?id=30299360#30309302. I managed an accuracy of 42% for 3-by-3 digit multiplication.
- YeGoblynQueenne 5y ago>> There are two issues here. One is the lack of working memory, which means that there is very little scratch space for calculating things with a meaningful sequential depth. It's a language model. It can generate text, not "calculate things". If you give it the right prompt, it will generate the right text, but if there's any computation going on, that's you computing the right prompt. See Clever Hans: https://en.wikipedia.org/wiki/Clever_Hans https://en.wikipedia.org/wiki/Clever_Hans
- canjobear 5y agoIf this were true, then engineered prompts would fail for held-out problem instances. But they don’t.
- deleted 5y ago[deleted]
- 7373737373 5y agoHas anyone tried using it for SAT problems yet?
- timdellinger 5y agomy recollection is that the original journal article announcing GPT-3 included some data on how it performed against SAT-style questions
- 7373737373 5y agoApparently these were college SAT questions, I'm wondering about https://en.wikipedia.org/wiki/SAT_solver https://en.wikipedia.org/wiki/SAT_solver
- Avalaxy 5y agoJust because you can, doesn't mean that you should. For some things it's just better to use a rules-based engine that is always correct, rather than a heuristics based algorithm that gives answers that are merely close.
- tasty_freeze 5y agoI don't think the author of the piece (or anyone for that matter) thinks GPT-3 should be used for running programs or evaluating functions. It is being discussed because it is surprising that GPT-3 can do it at all. It is worth investigating what types of emergent knowledge and behavior are encoded in the trained network, as the boundaries of its capabilities may help illuminate future neural network architecture design.
- algon33 5y agoIf I remember rightly, the AlphaCode paper includes a list of benchmarks, including the results of a finetuned GPT-3 for coding. I think they did it because Codex wasn't available to them when were doing their tests, but I might be wrong there.
- curling_grad 5y agoTheir paper has results from Codex. (see p. 21) https://arxiv.org/pdf/2203.07814v1.pdf https://arxiv.org/pdf/2203.07814v1.pdf
- imranq 5y agoAn interesting research direction would be to see how much the GPT3 deviates as we get more precise on various computational tasks. Possibly this would give some measure of some of the concepts the model has learned
- sho_hn 5y agoDo we today have any test suites/benchmarks for models along those lines?
- ogogmad 5y agoTheorem: There are infinitely many prime numbers. Proof: [GPT-3's "proof"] Let p be a prime number. Then by definition, p is a positive integer that has no positive divisors other than 1 and p. Now consider the integer q = p + 1. We will show that q is not a prime number, which will prove that there are infinitely many prime numbers. Since p is a positive divisor of q, q cannot be prime. [...another attempt with the same prompt] We know that there are infinitely many natural numbers. Let's suppose that there are only finitely many prime numbers. Then we can list them all: p_1, p_2, p_3, ..., p_n Now, consider the number N = p_1 p_2 p_3 ... p_n + 1 This number is greater than any of the prime numbers in our list, so it cannot be divisible by any of them. But it is also a natural number, so it must be prime. This contradicts our assumption that there are only finitely many prime numbers, so we must conclude that there are infinitely many prime numbers. ~~Personal remark: I had to give it the same prompt like 5 times to get a correct-ish proof. The first attempts came out like the first example.~~ The proof is still wrong.
- ogogmad 5y agoI keep asking GPT-3 to prove that the LR algorithm (for finding eigenvalues and eigenvectors) converges for PSD matrices. It keeps insisting that it's a form of gradient descent. Is that true?
- deleted 5y ago[deleted]
- gnulinux 5y agoBoth proofs are wrong, second one is closest. Second one should not claim that N is a prime (it likely isn't). It should say N is not divisible by any of p_i, and since due to Fun. Theo. of Arith. it is such that N = Sum {c_i q_i} where q_i are prime, and none of q_i in {p_i} which shows a finite list of primes is not possible construct.
- brian_cloutier 5y agoInterestingly, these attempts are about the same as what pops up when I try to remember the proof: - It's a proof by contradiction - The key step is in taking the finite list of primes, multiplying them together, and adding 1 I then try to flesh out the details, it might take a second to realize that this new number is also prime, and then a few moments more to remember the exact rationale why. Along the way the proof lives in a kind of superposition where I'm not clear on the exact details. The "proofs" you gave here seem to be serializations of a similar superposition! GPT-3 seems to remember the proof about as well as I do, but it's missing the final sanity check which tweaks the proof until all the pieces correctly fit together. In this case, you seem to be performing a version of this sanity check by running the prompt multiple times until a correct answer comes out. I wonder if it's possible to prove something more obscure using a similar process: GPT-3 comes up with ideas and the human sanity checks.
- csmeder 5y agoWhat year will GTP be able to take an app written in Swift/SwiftUI and output a spectacular Android translation? 3-years? 5-years? 10-years? This is an interesting benchmark because it is a very difficult problem, however: GTP has both everything it needs to do this without needing a fundamental improvement to the core of GTP (this process is more of a science than art) and using automated UI testing GTP can check if its solution worked. Thus this challenge is in the realm of what GTP already is, however, once it can do this it will have massive implications for how software is built.
- anyfoo 5y agoA terrible prospect. It's hard enough for people to faithfully port an application. People who participate and live in the world that makes up our reality. Leaving this up to an AI will at best flood us with low quality junk. At worst it's actively harmful.
- learndeeply 5y agoAnyone have any ideas on how they're doing text insertion using an auto-regressive model?
- lucidrains 5y agoyes, they are most likely finetuning with this type of pretraining https://arxiv.org/abs/2103.10360 https://arxiv.org/abs/2103.10360 quite easy to build
- spupe 5y agoThis is fascinating. I feel that we are still in the infancy of the field, however. These observations are analogous to naturalists of the past describing an animal's behavior, but we need to get to the point where more accurate estimates are made (ie, how often does it do each thing, how accurate it is after 100+ tries, etc). Every day we see a new observation showing wha GPTs can do, we also need a good way to make these observations systematic.
- lopatin 5y agoCan someone explain for a dummy how this is possible? How does it know that range() is zero indexed? Was it specifically trained on Python input/function/output data? Or did it just "learn" it? Do the researchers know how it learned it? Does it actually "run" the code? Like, if it was looping over 1 billion iterations would it take 1B times longer than if it was just one iteration? I have so many questions.
- etskinner 5y agoIt seems more likely that it learned it. If you knew nothing about Python, but understood the word "for" a little, and understood code a little, you're likely to figure out that range() is zero-indexed after you see something like this a few times >>> for i in range(3): print(i) 0 1 2
- lopatin 5y agoMy mind is just blown that it learned a language runtime based on examples. What would happen if you gave it an infinitely recusrive function? It can't stack overflow, there's no stack! Wait, is there?
- stevenhuang 5y agoMy guess is it would respond with the standard stack overflow error, from examples of similar output posted in its training set.
- deleted 5y ago[deleted]
- YeGoblynQueenne 5y agoIt hasn't. It's memorised the examples and can generate them with some variation and according to what output is more likely given the input. But that's generation, not computation. I don't know if this example helps, but a computer can generate (pseudo-)random numbers by executing an algorithm. A pair of dice can also generate random numbers because they're thrown so that they land at random and someone has marked pips on their faces that a human can read as numbers. The result may be similar, but one is generated by a computation and the other by a random process. The random process is not a computation. It's a random process. (Or just unpredictable).
- zora_goron 5y agoA quick question for anyone familiar with the architecture of these Transformer-based models -- I've heard that one reason why they don't work well with numbers is how the inputs are tokenized (i.e. as "chunks" rather than individual words/numbers). Is there anything architecturally preventing an exception in this form of tokenizing in the data preprocessing step, and passing numbers into the model in the format of 1 digit == 1 token? It seems like such a change could possibly result in a better semantic "understanding" of digits by the model.
- Veedrac 5y agoNothing prevents it, no. Transformers are certainly capable of learning mathematical tasks; consider [1] as an example, which uses big but regular token lengths. Alternatively you could just scale 'till the problem solves itself. [1] https://arxiv.org/abs/2201.04600 https://arxiv.org/abs/2201.04600
- deleted 5y ago[deleted]
- berryg 5y agoI struggle to understand how GPT-3 executes code. Is it simply running a python (or any other language) interpreter? Or is GPT-3 itself interpreting and executing python code? If the latter question is true that would be amazing.
- bidirectional 5y agoIt is the latter.
- deleted 5y ago[deleted]
- PeterisP 5y agoIt does not execute code, it "guesses" what the output of the code should be, given all the data it has seen during training - and, surprisingly, for many types of problems these guesses are accurate or close to that.
- aplanas 5y agoSeems that it can convert from Python to Perl: https://beta.openai.com/playground/p/o4qZWSXVz8JMmVaI9j9NMIKM?model=text-davinci-edit-001 https://beta.openai.com/playground/p/o4qZWSXVz8JMmVaI9j9NMIK...
- happycube 5y agoOddly, "convert to C" failed completely for me (wrapping the Python in main() { }) but C++ worked.
- charcircuit 5y ago>Is GPT-3 Turing complete? Maybe. It's obviously not. To handle infinite loops it needs to solve the halting problem. Which is not possible.
- anyfoo 5y agoI don't quite understand your answer. You don't need to solve the halting problem to be Turing complete, quite obviously. Why would GPT-3 need to in order to be?
- charcircuit 5y agoGPT-3 always gives an output after a certain amount of time. What should GPT-3 return when running: print("A"); potentiallyHalts(); print("B");
- anyfoo 5y agoPotentially time out? I don’t see the difference to, say, a python interpreter with a timeout. What would a human do, are we not Turing complete? I mean, in the strictest sense that isn’t Turing complete either, because when you have a timeout you cannot run every program a theoretical Turing machine could. But then no practical computer is, because resources are always constrained (e.g. finite memory instead of an infinite tape). So when we talk about something being Turing complete, we usually disregard the resource limitations and effectively substitute something like “we mean Turing complete in the sense that it would be if we also had infinite memory and time”. So, I still don’t understand why GPT-3 would have to (impossibly) solve the halting problem to be Turing complete[1], but everything else including a python interpreter or lambda calculus doesn’t. [1] Note that I don’t assert that GPT-3 could or could not be Turing complete, I just don’t know why the halting problem predicates that.
- _Nat_ 5y agoProbably easier to just observe that, if GPT-3 isn't reliably correct, then it's not consistent enough to simulate a Turing-machine and therefore isn't Turing-complete. As for loops: a Turing-machine could do infinitely many loops, so Turing-completeness implies that a system can do the same. If GPT-3 can't do infinitely many loops, it's not strictly Turing-complete; and if it can't do many loops, then it wouldn't seem like a meaningful approximation of a Turing-complete system.
- PaulHoule 5y agoIt would be remarkable if it got the right answers. But it can't because it doesn't have the right structure (e.g. GPT-3 finishes in a finite time, a program in a real programming doesn't necessarily!) GPT-3's greatest accomplishment is that it has "neurotypical privilege", that is if it gets an answer that is 25% or 95% correct people give it credit for the whole thing. People see a spark of intelligence in it the way that people see faces in leaf axels or in martian rock formations or how G.W. Bush looked in Vladimir Putin's eyes and said he got a sense of Putin's soul. (That was about the only thing in his presidency that he later said he regretted!) As an awkward person I am envious because sometimes it seems I get an answer 98% correct or 99.8% correct and get no credit at all.
- Micoloth 5y agoGPT3 does not think like a human, but it definitely executes code in a way that is more similar to a human than a computer.. Proof is, that indeed humans do get the wrong answer in quizzes like these sometimes! So i cannot understand this point of view of diminishing it as "spark of intelligence". It is exactly what advertised: a very big step forward towards real AI, even if definitely not the last one?
- PaulHoule 5y agoIt is the Emperor's New Clothes incarnate. It has the special talent of hijacking your own intelligence to make you think it is intelligent. People understood this about the 1966 ELIZA program but intellectual standards have dropped greatly since then.
- YeGoblynQueenne 5y ago>> Proof is, that indeed humans do get the wrong answer in quizzes like these sometimes! GPT-3 gets the wrong answer because it has memorised answers and it generates variations of what it has memorised. It generates variations by sampling at random from a probability distribution over what it's memorised. If it has the correct answer memorised, sometimes it will generate the correct answer, sometimes it will generate a slight variation of it, sometimes it will generate a large variation of it, sometimes it will generate something completely irrelevant (i.e. with a very small probability). Failure is not an exclusive characteristic of humans. In particular, any mechanical device will fail, eventually. For example, a flashlight will stop functioning when it runs out of battery. But not because it is somehow like a human and it just got it wrong that one time.
- mountainriver 5y agoThis is such an interesting field but I think there needs to be more focus on determinism and correctness. The stuff that’s happening with retrieval transformers is likely where this is heading
- unixhero 5y agoGreat so how do I run GPT-3 on my own hardware at home?
- mr_toad 5y agoIt’s not available to the public or open source so you can’t. Only the smallest models might run on a single GPU, the largest would need a large grid.
- unixhero 5y agoI have many computers and hundreds of threads, and not afraid to acquire 3090 cards. However GPT-3 seems elusive, and I can't find out what it takes to run it myself.
- bitwize 5y agoGPT-3 is starting to remind me of SCP-914. Give it an input, and its millions of tiny wheels churn and it produces something like what you want, but otherwise quite unexpected. Let's hope it doesn't turn into something like SCP-079...
- DC-3 5y agoVery far from an expert on ML, but isn't GPT-3 trivially not Turing Complete since it halts deterministically?
- deleted 5y ago[deleted]
- a-dub 5y agois there a search engine for the training data so that one can verify that it is actually performing novel operations and not just quoting back stuff from its incredibly large training set?
- timdellinger 5y agoI assume that GPT-3 is just exhibiting rote memory. For small numbers, it has accurate answers memorized from the training set, but for larger numbers, it just "remembers" whatever is close... hence the ability to estimate. My take is not that GPT-3 can run code, but rather that GPT-3 has memorized what code looks like and what the output looks like.
- mbowcut2 5y agoSo, for people unfamiliar with deep language models like GPT, it's essentially a program that takes in a prompt and predicts the next set of words based on a training corpus -- which in GPT-3's case is a large portion of the internet. In these examples GPT is not executing any python code, it has just been trained on enough Python code/output to successfully predict what kinds of outputs these functions would produce.
- FartyMcFarter 5y ago> GPT is not executing any python code, it has just been trained on enough Python code/output to successfully predict what kinds of outputs these functions would produce. This distinction is not that clear though. If you can predict well the output of a function, that's equivalent to executing the code.
- toomanydoubts 5y agoFor that to be somewhat true, I can see at least two prerequisites: 1) the function must be pure(no side-effects); 2) predicting well is not enough, GPT must predict perfectly a hundred percent of the time. Still, technically you're not executing the code.
- wyattpeak 5y agoComputers don't execute code perfectly 100% of the time. I agree that it's fundamentally different, but I'm not exactly sure how, and I think it's subtler than you're suggesting.
- tluyben2 5y agoIt is fundamentally different in mathematical foundations; some functions are proven formally verified and therefor will execute 100% perfectly (I guess you are talking about actual bugs like hardware issues?); what gpt3 does is not even close to that; if you put the same input to gpt3 multiple times it comes up with different answers. That is nowhere close to a computer executing an algorithm.
- graiz 5y agoGPT3 is a really impressive auto-complete. It takes inputs and predicts what text should be output. It's super impressive and it looks like it's smart but it is not running code, it's not Turing complete and if you understand how it works it's very easy to cause it to produce significant errors.
- luxurytent 5y agoSimilar to how my four year old can read books, because he’s memorized the words I’ve read to him through repeated story times.
- deleted 5y ago[deleted]
- inopinatus 5y agoEven a stopped clock tells the right time, twice a day.