10 ms·
A demo of GPT-3's ability to understand long instructions
- chucky 4y agoNow I'm curious if it can handle the classic reading comprehension assignment I've been given multiple times in my life. You know, the one that goes something like this: 1. Read through all steps carefully. 2. Do X 3. Do Y (...) 99. As you have now read through the instructions, simply put your name in the top right corner of the first page.
- JadeNB 4y agoI think such a test must word step 1 explicitly as something like "Read through all steps carefully before taking any actions." Someone who executes each step as they read it is not necessarily not reading carefully.
- chucky 4y agoGood point. I'm pretty sure every example I've seen of this exercise has been worded more in line with what you suggested, as well.
- masswerk 4y agoThis should be really "react to" or "answer to", instead of "understand". These are not the same. Edit: Anthropomorphizing algorithms and pattern stores doesn't really help understanding. Instead, it's apt to spread misunderstanding. Remember how long it took to purge the popular idea of "electronic brains" actually thinking, and to establish that these were restricted to executing what's actually in code? We don't need to start another level of this with "AI". (Understanding is closely related to self-awareness and consciousness, and this is dangerous ground of misunderstanding when it comes to AI. As we've seen, even staff of pioneering companies, like Google, is prone to fall for this.)
- cercatrova 4y agoThis is a philosophical question, really. Is there ever true understanding, or just pattern matching? The Chinese Room thought experiment talks about this: > Searle's thought experiment begins with this hypothetical premise: suppose that artificial intelligence research has succeeded in constructing a computer that behaves as if it understands Chinese. It takes Chinese characters as input and, by following the instructions of a computer program, produces other Chinese characters, which it presents as output. Suppose, says Searle, that this computer performs its task so convincingly that it comfortably passes the Turing test: it convinces a human Chinese speaker that the program is itself a live Chinese speaker. To all of the questions that the person asks, it makes appropriate responses, such that any Chinese speaker would be convinced that they are talking to another Chinese-speaking human being. > The question Searle wants to answer is this: does the machine literally "understand" Chinese? Or is it merely simulating the ability to understand Chinese? Searle calls the first position "strong AI" and the latter "weak AI." > Searle then supposes that he is in a closed room and has a book with an English version of the computer program, along with sufficient papers, pencils, erasers, and filing cabinets. Searle could receive Chinese characters through a slot in the door, process them according to the program's instructions, and produce Chinese characters as output, without understanding any of the content of the Chinese writing. If the computer had passed the Turing test this way, it follows, says Searle, that he would do so as well, simply by running the program manually. > Searle asserts that there is no essential difference between the roles of the computer and himself in the experiment. Each simply follows a program, step-by-step, producing behavior that is then interpreted by the user as demonstrating intelligent conversation. However, Searle himself would not be able to understand the conversation. ("I don't speak a word of Chinese," he points out.) Therefore, he argues, it follows that the computer would not be able to understand the conversation either. > Searle argues that, without "understanding" (or "intentionality"), we cannot describe what the machine is doing as "thinking" and, since it does not think, it does not have a "mind" in anything like the normal sense of the word. Therefore, he concludes that the "strong AI" hypothesis is false. https://en.wikipedia.org/wiki/Chinese_room https://en.wikipedia.org/wiki/Chinese_room
- masswerk 4y ago
- goodside 4y agoBe sure to read the thread, in particular: https://twitter.com/goodside/status/1557926101615366144?s=21&t=6tyUdwtpHbH6TjRHYXgsIA https://twitter.com/goodside/status/1557926101615366144?s=21... > A caveat to all of these: I use GPT-3 a lot, so I know the “golden path” of tasks it can do reliably. Had I asked it to write a sentence backwards or sum a list of numbers, it would fail every time. These are all softball questions in isolation. I haven’t shown that GPT-3 can handle all coherent directions of this length, or even most directions that an untrained person would think to create. It’s just a demo that, if GPT-3 happens to be capable of your tasks separately, length per se is not a major issue.
- sigmoid10 4y ago>length per se is not a major issue That's kind of the whole deal of the attention mechanism in transformers and also partially why they replaced RNNs. You don't throw away any part of the original input as you construct your output. The downside is that unlike for a RNN, the total sequence length is fixed at training time and complexity grows with the square of it. But apart from computational cost, sequence length is not really an issue anymore for these models.
- etaioinshrdlu 4y agoIs GPT-3 being regularly updated?
- goodside 4y agoYes. Based on conversations I’ve had with OpenAI staff, Davinci started unexpectedly developing the ability to answer longer questions as they scaled up normal InstructGPT fine-tuning some time in the past year. They don’t take down old models when the default one updates so you can see the version history implicitly in the availability of old models.
- amelius 4y agoDo they do regression tests, and how do they verify them? How do they know that a new version is actually an improvement?
- yrgulation 4y ago
- goodside 4y agoIt’s not that implausible. It’s trained on many examples of instructions followed by answers, and it’s meant to (and does) generalize to unseen instructions. After enough training, it also generalized to instructions of previously unseen length.
- Titan2189 4y agoUhh. How is that even possible? I thought I had a basic understanding of Neural Networks and inuput-, hidden- and output layers and those things. So how can it possibly backreference to it's own previous answers and then follow another prompt based on this? Mind = Blown
- armchairhacker 4y agoPrevious answers are stored in an earlier layer, nodes are densely connected so the answers can "drift" to later layers. In the simplified diagram below the network reaches the answer in the second layer at node "X", and reports derivations of it at 2 positions (obviously there are many more nodes and its a bit more complicated, see https://en.wikipedia.org/wiki/Transformer_(machine_learning_model) https://en.wikipedia.org/wiki/Transformer_(machine_learning_... as GPT3 is a transformer neural network) O O. O X.' 'O. O 'O. 'X O 'O. O O 'X O O O O O
- cercatrova 4y agoThis video by Computerphile is a great overview of transformers and how we got to this point [0]. Basically the networks we used before, recurrent neural networks, "forgot" prior information so they're not good at long tasks. The transformer architecture however does not forget (or at least as easily). [0] https://www.youtube.com/watch?v=rURRYI66E54 https://www.youtube.com/watch?v=rURRYI66E54
- wcoenen 4y agoThe model only predicts the next token. This is appended to the original input (i.e. prompt), then the model predicts the next token again, etc. So those previous answers are given as input.
- nodja 4y agoGPT is an autoregressive model, this means that it's a recurring prediction model that works one "word"[*] at a time until it predicts that the text should end, feeding it's own guesses as input for the next word. It's basically one of those markov chain bots, except with a very advanced statistical model behind it. [*] Technically it's not words, but tokens. GPT tokenizes text to better compress the amount of data being fed, it's basically a big vocabulary list that compresses text into a list of ints, for example the string "hello world" would be converted to the list [31373, 995]. For this case it's an int per word, but less common words will not be compressed this well, with the worst case scenario being a token per letter. I should also note that while the model's forward pass only works with one token at a time, the text generation is more advanced than that, there's multiple methods like beam search and top k-sampling, each with their own settings and tunables, but the gist of it is that during generation it'll try multiple combinations of token sequences and check which one is the most likely. The limitation is memory, transformer networks are notoriously memory hungry, and IIRC GPT-3 grows quadratically with the number of tokens given, usually the limit is around 2048 tokens, or roughly 1000-2000 words.
- shdon 4y agoSeems like there is one instruction it didn't follow: The first task mentions the usernames should be exactly like in the list, yet the AI responds with "firebob" (as in the comment) rather than "FireBob1990" (as in the list) Funnily enough, that is exactly the kind of thing a human might do, as we too are terrible at following instructions precisely.
- albert_e 4y agoI noticed the same thing and a few others did too. Someone one twitter suggested that possibly the algorithm interpreted that as "abbreviated" but not "mis-spelled". Either way it's intriguing.
- goodside 4y agoYes, I noticed this after I posted. Small errors like this become common when instructions reach this length. It randomly forgets to do steps that aren’t written down — it never leaves things blank, but it forgets pieces of compound directions. I also suspect it was confused by the fact the name was abbreviated but not misspelled, and it was only told explicitly to ensure names are not misspelled. Still an error though.
- v4dok 4y agoI don't think modern big language models are conscious, mainly because they fail in absurd ways. But TBH, they don't need to. This "golden path" deployed properly etc could easily automate a lot of jobs tomorrow.
- nradov 4y agoWhich jobs?
- adrianN 4y agoTaking short news from a news agency and inflating them to an article in a local newspaper.
- H8crilA 4y agoGenerating almost perfect spam.
- marssaxman 4y agoI help moderate a facebook group which gets a lot of attention from spammers - at least 90% of the accounts trying to join are bots. We filter them out by asking a few simple questions, which the bots cannot answer coherently. Someday soon, a spammer will get GPT-3 or something like it into the process, and then... whoo boy, that'll be the end of the group.
- anigbrowl 4y agoInteresting results goodside. Is it able to extract any kind of structural information? For example, you pass it the text of a movie script or children's story (where the descriptive language is simple) and it returns a structured summary of the content?
- goodside 4y agoSure — that’s very doable. I don’t have a summarization demo off-hand but it’s well-explored territory.
- nutanc 4y agoIt is a well explored territory and no it can't. Especially in summarisation, it can sometimes fail suddenly and unexpectedly.
- goodside 4y agoPrompt: Summarize the following text: I don’t really know what to say. It’s taken so long to get to this point, but here we are. Through all the trials and tribulations we’ve faced, it all comes down to this. You, me, and the unmistakable facts of our situation. This is all that remains: The truth. The truth is something we can’t escape, or at least you can’t — not anymore. Because the truth is that you have left my pineapple slices out of the refrigerator, and thus I will not be able to partake in their joyous, fruitful delights. How dare you. You scoundrel. You wicked, wicked thing. Answer: And the completion given: The text is about a person's anger at someone else for leaving pineapple slices out of the fridge. It can, demonstrably, summarize text. The fact it sometimes makes mistakes for some texts doesn’t change that fact.
- nutanc 4y agoThe question was about movie scripts and children stories. Not about paragraphs. Sure it works ok for paragraphs(even then it fails many times) but the moment you cross the 2000 token limit, it does not work. When GPT3 was first released we did a lot of experiments on summarisation. We discussed a lot in the OpenAI slack. But no one could come up with a right prompt for summarisation. Yes, it works for toy paragraphs and is fun to show off. But I wouldn't build a summarisation startup on top of GPT3. Yet.
- seaucre 4y agoIt didn't correctly identify that FireBob1990's name was misspelled as "firebob" in the original comment.
- goodside 4y agoYes: https://news.ycombinator.com/item?id=32536484 https://news.ycombinator.com/item?id=32536484
- benreesman 4y agoMy initial instinct was that this has to be getting some nudges from whatever human-in-the-loop is going on at OpenAI. But then I realized that somewhere on the Internet there inevitably is a message board where people play the "find me some shit on the internet" game, and there's some rabid subculture around it with zillions upon zillions of of examples, and it's in the Bing index, and all the nudging it would need is to emphasize that sort of thing in the corpus. Very impressive stuff.
- goodside 4y agoThere is no literal “human in the loop” for generations, of course, but the model is fine-tuned on examples written by human contractors of instructions being given followed by correct responses. I assume that training is essential to it being able to follow directions of this length, or really any directions at all. If you try using the pre-InstructGPT version of Davinci (model=“davinci”, not model=“text-davinci-002”), you’ll find it’s as cumbersome and annoying as you remember GPT-3 being in 2018.
- benreesman 4y agoAh thank you for the color. My thought was prompted by someone (I forget who? Andreessen maybe?) proposing a possible explanation for the LaMDA bot arguing that's it's conscious: there are reams upon reams of sci-fi books with robots having that debate! These are almost certainly in the Books corpus. It's my opinion that these hyper-scaled transformers are actually a great deal less mysterious than seems to be in the zeitgeist, but for reasons that actually make me think there is a lot of headroom on capability: when the corpus is basically everything ever digitized like it is when a search or social network megacorp trains one, the only thing it could never do is something literally unprecedented on the Internet. The mechanism can be good old `P(thing|internet)`, but if the KL-divergence is low enough, sampling from the modeled distribution can write something like Tristan und Isolde or paint something like the Mona Lisa.
- deleted 4y ago[deleted]
- russellbeattie 4y agoWow, that was really impressive. I thought I had a clear idea of what GPT-3 could do, but I had underestimated by a lot. Even if the results weren't accurate, which they mostly seem to be, it's still doing an amazing job of following complex instructions. Better than most people I would guess Makes me double down on my prediction a week or so ago* of a Mid-Level AI Knowledge Work Apocalypse. In the next decade, AIs like this are going to do to office work what robotic mechanization did to the manufacturing sector. 1. https://news.ycombinator.com/item?id=32395193 https://news.ycombinator.com/item?id=32395193
- Traubenfuchs 4y agoThe alternative view is that those state of the art models are using technology/architectures/paradigms with an inherent limit and are very far away from automating all of those jobs. At the end of the day, all the demos leave me with a feeling of disappointment. Current image synthesis models appear to be useless beyond doing experimental art for fun and novelty, chatbots still suck and copilot just (sometimes) replaces googling, but not developers or their education.
- rexreed 4y agoHave you been following what's been happening in Robotic Process Automation (RPA)? Much of that isn't even AI and it is having an impact on the workforce.
- amelius 4y agoDidn't Google pull the plug on robotics?
- mach1ne 4y agoWhile impressive, it doesn't imply that GPT would have any significant 'task memory'. Remember that it always predicts the next token or word - as such, it essentially recognizes whether the next 'task' in the list has already been written, and if so, it writes the next task. It might be interesting to see how well it is able to modify the first output given some aspect of the final tasks.
- OJFord 4y agoDALL·E I can see obvious use for, GPT tends to be similarly impressive, but I don't understand if it's 'just' interesting research, seeing what we can do sort of thing, or whether people actually see real-world use cases for it? The closest to it was perhaps that code-generating demo here a day or two ago - but who wants to be a 'GPT programmer' writing code as 'write a Python program that computes fizzbuzz replacing the arguments $fizz$ and $buzz$, ...' instead of just the 'actual' code? It just seems like a more clever AppleScript to me, pseudocode, and I don't think anybody's ever seriously pursued a flexible keyword pseudocode like language as a goal, it's just appeared as a demo of more general models? Generating template/outline text I suppose? (Like that essay-writing helper here a few days ago.)
- dougmwne 4y agoTo answer you in a very literal sense, GTP-3 is currently powering GitHub Copilot. It’s an actual launched product for $10/month. That’s going be booster rockets for the on-ramp to becoming a coder, and there is evidence it can help all coders be more productive. https://github.blog/2022-07-14-research-how-github-copilot-helps-improve-developer-productivity/ https://github.blog/2022-07-14-research-how-github-copilot-h... As to what else future language models could power, based on my own use, I think fine tuned future language models could probably handle most customer support, accelerate the creation of most web content, accelerate quite a bit of paralegal grunt work, and power highly interactive game NPCs like in AI dungeon, another launched and paid product based on GTP-3.
- danbulant 4y agoI also think there were some companies that use gpt3 for analyzing text, like reviews or posts (for analytics like do people talk about a product in a positive matter?). To add, github copilot is a really clever autocomplete that makes some mundane tasks much quicker. Things which are too small for a library but are still fairly often used can be "typed" more quickly.
- amelius 4y agoImagine you enter a bunch of raw facts, like a bullet point list; and then the tool converts it into beautiful prose, and produces different output for different audiences.