7 ms·
I am totally confused by people not being impressed with gtp3. If you asked 100 people in 2015 tech industry if these results would be possible in 2020, 95 woul
by Bx6667 6y ago
I am totally confused by people not being impressed with gtp3. If you asked 100 people in 2015 tech industry if these results would be possible in 2020, 95 would say no, not a chance in hell. Nobody saw this coming. And yet nobody cares because it isn’t full blown AGI. That’s not the point. The point is that we are getting unintuitive and unexpected results. And further, the point is that the substrate from which AGI could spring may already exist. We are digging deeper and deeper into “algorithm space” and we keep hitting stuff that we thought was impossible and it’s going to keep happening and it’s going to lead very quickly to things that are too important and dangerous to dismiss. People who say AGI is a hundred years away also said GO was 50 years away and they certainly didn’t predict anything even close to what we are seeing now so why is everyone believing them?
- abernard1 6y ago> And yet nobody cares because it isn’t full blown AGI. That’s not the point. The point is that we are getting unintuitive and unexpected results. I don't think these are unintuitive or unexpected results. They seem exactly like what you'd get when you throw huge amounts of compute power at model generation and memorize gigantic amounts of stuff that humans have already come up with. A very basic Markov model can come up with content that seem surprisingly like a human would say. If anything, what all of the OpenAI hype should confirm is just how predictable and regular human language is.
- yters 6y agoExactly. Even gpt3 is not creating new content. It is just permuting ecisting content while retaining some level of coherence. I don't reason by repeating various tidbits I've read in books in random permutations. I reason by thinking abstractly and logically, with a creative insight here and there. Nothing at all like a Markov model trained on a massive corpus. Gpt3 may give the appearance of intelligent thought, but appearance is not reality.
- CrazyStat 6y agoGPT-3 is nothing like a Markov model.
- yters 6y agoSame sort of generative probabilistic model idea.
- deleted 6y ago[deleted]
- mietek 6y ago> I don't reason by repeating various tidbits I've read in books in random permutations. Are you sure?
- yters 6y agoYes, I would fail any sort of math exam if I used the GPT-3 model.
- catalogia 6y agoAll creative work is derivative.
- yters 6y agoNot all derivative work is creative.
- Aqueous 6y agoRight...but at the end of the day that's what intelligence is. You are just an interconnected model of billions of neurons that has been trained on millions of facts created by other humans. Except for this model can vastly exceed the amount of factual knowledge that you could possibly absorb over your entire lifetime.
- abernard1 6y ago> You are just an interconnected model of billions of neurons that has been trained on millions of facts created by other humans. ...but I didn't pop out of the womb that way, and as you said, over my lifetime I will read less than 1 millionth of the data that GPT-3 was trained on. GPT-2 had a better compression ratio than GPT-3, and I'm sure a GPT-4 will have a worse compression ratio than GPT-3 on the road we're on. Rote memorization is hardly what I'd call intelligence. But that's what we're doing. If these things were becoming more intelligent over time, they'd need less training data per unit insight. This isn't a dismissal of the impressiveness of the algorithms, and I'm not suggesting the classic AI effect "changing the goalposts over time." I fundamentally believe we're kicking a goal in our own team's net. This is backwards.
- Voloskaya 6y ago> They seem exactly like what you'd get when you throw huge amounts of compute power I disagree with that. The one/few shot ability of the model is much much better than what I would have imagined, and I know very few people in the field that saw GPT-3 and were like "yep, exactly what I thought".
- canjobear 6y ago> A very basic Markov model can come up with content that seem surprisingly like a human would say. This is false. Natural language involves long-term dependencies that are beyond the ability of any Markov model to handle. GPT-2 and -3 can reproduce those dependencies reliably. > If anything, what all of the OpenAI hype should confirm is just how predictable and regular human language is. Linguists have been trying to write down formal grammars for natural languages since the 1950s. Some of the brightest people around have essentially devoted their lives to this task. And yet no one has ever produced a complete grammar of any human language. So no, human language is not predictable and regular, at least not in any way that we know how to describe formally.
- abernard1 6y agoW.r.t. the Markov model, I just mean that something even that trivial can sound lifelike. It's not surprising that throwing billions of times more data at the problem with more structure can make the parroting better. > So no, human language is not predictable and regular, at least not in any way that we know how to describe formally. I don't know what to say about this other than perhaps the NLP community has been a little too "academic" here and I disagree. Grade schoolers routinely are forced to make those boring diagrams for their particular language, and that has tremendous structure. When you add that structure (function) with the data of billions of real-world people talking, it's not surprising that the curve fit looks like the real thing. Given how powerful things like word2vec have been that do very, very simple things like distance diffs between words, it's not surprising to me that the state of the art is doing this.
- FeepingCreature 6y agoIt is surprising! You could throw all the data of the entire human race at a Markov model and it would not sound a tenth as good as even GPT-2. Transformers are simply in a new class.
- Bx6667 6y agoWere you alive in 2010?
- quonn 6y agoFirst, I remember the demos for GTP-2. Later, when it was available and I could try it myself, I was kind of disappointed in comparison. Second, while impressive we are also finding out at the same time just how much more is needed to make something of value. It‘s like speech recognition in 1995. Mostly there, but in the end it took another 20 years to actually work. But still, it‘s exciting.
- m0zg 6y agoThe problem is that not only is this "not full blown AGI". The problem is that, if you understand how this works, it's not "intelligence" at all (using the layperson meaning of the word, not the marketing term), and it's not even on the way to get us there.
- liuliu 6y agoI disagree. Yes, it is just a decoder of transformer. But it looks like we are really close, with some tweaks on the network structure, reward function design and inputs / outputs. On the same time, GPT-3 also points how far away we are at hardware level. Let me put it this way: I don't know how challenging the rest is going to be, but it surely looks like we are on the right path finally.
- m0zg 6y agoIt fundamentally has no _reasoning_. There is no AGI without reasoning.
- deleted 6y ago[deleted]
- hackinthebochs 6y agoWhat makes you think this? The fact that it can produce working code from a prompt in some cases shows rudimentary non-trivial reasoning. Hell, GPT-2 demonstrated rudimentary reasoning of the trivial sort.
- abernard1 6y ago> The fact that it can produce working code from a prompt in some cases shows rudimentary non-trivial reasoning. It doesn't at all. It indicates that it read stackoverflow at some point, and that on a particular user run, it replayed that encoded knowledge. (I'd also argue it shows the banality of most React tutorials, but that's perhaps a separate issue.) Quite a lot of these impressive achievements boil down to: "Isn't it super cool that people are smart and put things on the internet that can be found later?!" I don't want to trivialize this stuff because the people who made it are smarter than I will ever be and worked very hard. That said, I think it's valid for mere mortals like myself to question whether or not this OpenAI search engine is really an advancement. It also grates on me a bit when everybody who has a criticism of the field is treated like a know-nothing Luddite. The first AI winter was caused by disillusionment with industry claims vs reality of what could be accomplished. 2020 is looking very similar to me personally. We've thrown oodles of cash and billions of times more hardware at this than we did the first time around, and the most use we've gotten out of "AI" is really ML: classifiers. They're super useful little machines, but they're sensors when you get right down to it. AI reality should match its hype, or it should have less hype (e.g. not implying GPT-3 understands how to write general software).
- riotman 6y ago"People who say AGI is a hundred years away also said GO was 50 years away" this is not true. The major skeptics never said this. The point skeptics were making was that benchmarks for chess (IBM), Jeopardy!(IBM), GO (Google), Dota 2 (OpenAI) and all the rest are poor benchmarks for AI. IBM Watson beat the best human at Jeopardy! a decade ago, yet NLP is trash, and Watson failed to provide commercial value (probably because it sucks). I'm unimpressed by GLT-3, to me nothing fundamentally new was accomplished, they just brute forced on a bigger computer. I expect this go to the same way as IBM Watson.
- justinpombrio 6y agoWhat would be a good benchmark? In particular, is there an accomplishment that would be: (i) impressive, and clearly a major leap beyond what we have now in a way that GPT-3 isn't, but (ii) not yet full-blown AGI?
- rtpg 6y agoMaybe nothing? “Search engines through training data” are already the state of the art, and have well documented and mocked failure cases. Unless someone comes along with a more clever mechanism to pretend it’s learning like humans, you’re not looking at a path towards AGI in my opinion.
- justinpombrio 6y ago> you’re not looking at a path towards AGI in my opinion What I'm trying (and apparently failing?) to ask is, what would a step on the path towards AGI look like? What could an AI accomplish that would make you say "GPT-3 and such were merely search engines through training data, but this is clearly a step in the right direction"?
- abernard1 6y ago> What I'm trying (and apparently failing?) to ask is, what would a step on the path towards AGI look like? That's an honest and great question. My personal answer would be to have a program do something it was never trained to do and could never exist in the corpus. And then have it do another thing it was never trained to do, and so on. If GPT-3 could say 1) never receive any more input data or training, and then 2) read an instruction manual for a novel game that shows up a few years from now (so it can't be replicated from the corpus), and 3) plays that game, and 4) improves at that game, that would be "general" imo. It would mean there's something fundamental with its understanding of knowledge, because it can do new things that would have been impossible for it to mimic. The more things such a model could do, even crummily, would go towards it being a "general" intelligence. If it could get better at games, trade stocks and make money, fly a drone, etc. in a mediocre way, that would be far more impressive to me than a program that could do any of those things individually well.
- screye 6y agoFrom the POV of an AI practitioner, there is one and only one reason I remain unimpressed with GPT3. It is nothing more than one big transformer. At a technical level, it does nothing impressive, apart from throw money at a problem. So in that sense, having already been impressed at Transformers and then ELMO/BERT/GPT-1 (for making massive pretraining popular). There is nothing in GPT3 that is particularly impressive outside of Transformers and massive pre-training, both of which are well known in the community. So, yeah, I am very impressed by how well transformers scale. But, idk if I'd give OpenAI any credit for that.
- neurologic 6y agoThe novelty of GPT3 is its few shot learning capabilities. GPT3 shows a new, previously-unknown, and, most importantly, extremely useful property of very large transformers trained on text -- that they can learn to do new things quickly. There isn't any ML researcher on record who predicted it.
- ricksharp 6y agoYes, the emergent ability to understand commands mixed in with examples is pretty crazy.
- MiroF 6y ago> There isn't any ML researcher on record who predicted it. That's just absurd - this was an obvious end-result for LM. NLP researchers knew that something like this was absolutely possible, my professor predicted it like 3 years ago.
- rvz 6y ago> People who say AGI is a hundred years away also said GO was 50 years away and they certainly didn’t predict anything even close to what we are seeing now so why is everyone believing them? Do you why AlphaGo decided to perform move 37 in Game 2 with Lee Sedol? Can AlphaGo explain itself as to why it did that move? If we don't know why it made that decision, then it is a mysterious black-box hiding it's decisions and taking in an input to produce and output, which is still a problem. This isn't useful to researchers in understanding decisions of these AI systems, especially for AGI. Hence, this problem also applies to GPT-3. While it is still an advancement in NLP, I'm more interested in getting a super accurate or generative AI system to explain itself than one that cannot.
- ZephyrBlu 6y ago>While it is still an advancement in NLP, I'm more interested in getting a super accurate or generative AI system to explain itself than one that cannot Why? People can explain ourselves because we rationalize our actions, not necessarily because we know why we did something. I don't understand why we hold AI to such a high standard.
- MiroF 6y agothere seems to be an overabundance of negative sentiment towards deep learning among hn commentators, but whenever i hear the reasons behind the pessimism i'm usually unimpressed.
- LoSboccacc 6y agofor the same reason why we have psychiatrist, for when the AI does a mistake, you need to fix it, work around it, prevent it or if all else fail to protect others from it. it's all fun and games when AI do trivia. when AI get plugged into places that can result in tangible real world consequences (i.e. airport screening) you need to be able to reason about the system so it gets monotonically better over time.
- joe_the_user 6y agoI think there's a divide between "impressive" and "good". I think deep learning will keep creating more impressive, more "unintuitive and unexpected", more "wow" results. The "wow" will get bigger and bigger. Gpt-3 is more impressive, more "wow"-y than Gpt-2. Gpt-3 very impressively seems to demonstrate understanding of various ideas, Gpt-3 indeed very impressively develops ideas over several sentences. No argument with the "unintuitive and unexpected" part. The problem is the whole thing doesn't seem definitively good (in Gtp-3's case, doesn't produce good or even OK writing). It's not robust, reliable, trustworthy. The standard example is the self-driving car. They still haven't got those reliable but with more processing power, a company could probably add more bells and whistles to the self-driving process but still without making it safe. And GPT-3 seems in that vein - more "makes sense if you're not paying attention", the same "doesn't really say coherent things". I'm trying to trace a middle ground between the two reactions. I'm perhaps laughing a little at those just looking at impressive but I acknowledge there's something real there. Indeed, the more you notice something real there, the more you notice something real missing there too.
- deleted 6y ago[deleted]
- Polylactic_acid 6y agoThats similar to my thoughts. That demo video of generating html was very impressive, I have never seen anything that can do that, but its also 1000x less useful than squarespace or wordpress. The tool in its current state is totally useless even if it is very impressive.
- p1esk 6y agoIt's not robust, reliable, trustworthy Is human writing robust, reliable, trustworthy? Would you agree that some humans produce vastly better writing than others? Have you never read comments here on HN that appeared to be incoherent rambling, logically faulty, or just shallow, trite and cliched? GPT-1 is a significant improvement over earlier RNN based language models. GPT-2 is a significant improvement over GPT-1. GPT-3 is a significant improvement over GPT-2, especially in terms of "robustness". All these achievements appeared in the course of just 3 years, and we haven't yet reached the ceiling of what these large transformer based models can do. We can reasonably expect that GPT-4 will be a significant improvement over GPT-3 because it will be trained on more and better quality data, it will be bigger, and it might be using better word encoding methods. Aside from that, we haven't even tried finetuning GPT-3, I'd expect it would result in a significant improvement over the generic GPT-3. Not to mention various potential architectural and conceptual improvements, such as an ability to query external knowledge bases (e.g. Wikipedia, or just performing a google search), or an ability to constrain its output based on an elaborate profile (e.g. assuming a specific personality). There are most likely people at OpenAI who are working on GPT-4 right now, and I'm sure Google, Microsoft, Facebook, etc are experimenting with something equally ambitious. I agree that GPT writing is not "good" if we compare it to high quality human writing. However, it is qualitatively getting better and better with each iteration. At some point, as soon as a couple years from now, it will become consistent and coherent enough to be interesting and/or useful to regular people. Just like self-driving cars in a couple of years might reach the point where the risk of dying is higher when you drive than when AI drives you.
- paulie_a 6y agoWhat's impressive about it? It's bigger, that's cool. What's it actually mean in the real world. I see nothing to get excited about at this point.
- sama 6y agoI think people should be impressed, but also recognize the distance from here to AGI. It clearly has some capabilities that are quite surprising, and is also clearly missing something fundamental relative to human understanding. It is difficult to define AGI, and it is difficult to say what the remaining puzzle piece are, and so it's difficult to predict when it will happen. But I think the responsible thing is to treat near-term AGI as a real possibility, and prepare for it (this is the OpenAI charter we wrote two years ago: https://openai.com/charter/ https://openai.com/charter/). I do think what is clear is that we are, in the coming years, going to have very powerful tools that are not AGI but that still change a lot of new things. And that's great--we've been waiting long enough for a new tech platform.
- icebergwarrior 6y agoOn a core level, why are you trying to create an AGI? Anyone who has thought seriously about the emergence of AGI equates the chance that AGI causes a human extinction level event ~20%, if not greater. Various discussion groups I am a part of now see anyone who is developing AGI to be equivalent to developing a stockpile of nuclear warheads in your basement that you're not sure won't immediately shoot off on completion. As an open question. If one believes that 1. We do not know how to control an AGI 2. AGI has a very credible chance to cause a human level extinction event 3. We do not know what this chance or percentage is 4. We can identify who is actively working to create an AGI Why should we not immediately arrest people who are working on an "AGI-future" and try them for crimes against humanity? Certainly, In my nuclear warhead example, I would immediately be arrested by the government of the country I am currently living in the moment they discovered this.
- TOKYORACER99 6y agoThe problem is that if the United States doesn't do it, China or other countries will. It's exactly the reason why we can't get behind on such a technology from a political / national perspective. For what it's worth though, I think you're right that there are a lot of parallels with nuclear warheads and other dangerous technologies.
- api 6y agoI am really impressed with it as a natural language engine and query system. I am not convinced it "understands" anything or could perform actual intellectual work, but that doesn't diminish it as what it is. I'm also really worried about it. When I think of what it will likely be used for I think of spam, automated propaganda on social media, mass manipulation, and other unsavory things. It's like the textual equivalent of deep fakes. It's no longer possible to know if someone online is even human. I am thinking "AI assisted demagoguery" and "con artistry at scale."
- totetsu 6y agoI can't help but feel what gpt is really teaching us about is language not AI.
- TOKYORACER99 6y agoIMO, language is one of the purest forms of thinking / consciousness. What is our brain doing that makes it different?
- totetsu 6y agoThis brings to mind the debates between Frank Ramsey and Ludwig Wittgenstein. Episode: https://philosophybites.libsyn.com/cheryl-misak-on-frank-ramsey-and-ludwig-wittgenstein https://philosophybites.libsyn.com/cheryl-misak-on-frank-ram... Media: https://traffic.libsyn.com/secure/philosophybites/Cheryl_Misak_on_Frank_Ramsey_and_Ludwig_Wittgenstein.mp3?dest-id=14010 https://traffic.libsyn.com/secure/philosophybites/Cheryl_Mis...
- burtonator 6y agoI mean I find the fact that a human can actually build and work with a tool that it can't actually understand? Even now, you could, if you wanted to, rip apart your computer even to the CPU level and understand how it works. Even analyzing the code. Sure, it might take you ten years. But you would NEVER be able to understand how GPT3 works... it's just too complex.
- ladberg 6y agoReally? I bet in a few years we'll have tools that can inspect a model and tell you exactly what parts do what function and how they do it.
- TOKYORACER99 6y agoAgreed. Even if we put research into deconstructing and attempting to understand how deep neural networks work in tasks such as autonomous driving, the fact is that these tasks are too complex to even logically describe. That said, I do think it is possible to come up with robust guarantees to these methods.
- f6v 6y agoI’m no expert, but tools such as SHAP and DeepLift can give you insight into what activates a network. It’s probably not possible to inspect a network with billions of parameters, however it’s to be expected since I don’t think that explainable ML is an established field yet. But also think about it from another angle: it doesn’t seem too hard to explain why people say what they say. We can usually get into the shoes if the other person if we try hard enough. However, if we say there’s no way for us to explain GPT-3, it just shows how fundamentally different it is from human mind.
- missosoup 6y agoIt's something pretty unique to ML research. The goal posts keep moving whenever an advancement is made. Every time ML achieves something that was considered impossible X years ago, people look at it and say, actually, that's nothing special, the real challenge is <new goalpost>. I'm pretty sure even as we cross into AGI, people will react the same way. And only then will some stop and realize that we just wrote off our own intelligence as nothing special, a parlour trick.
- vladislav 6y agoImpressive to a human is a highly subjective property. Humans generally consider the understanding of language to be an intelligent trait, yet tend to take basic vision which took much longer evolutionarily to develop for granted. Neural networks can approximate arbitrary functions, and the ability to efficiently optimize neural network parameters over high dimensional non-convex landscapes has been well established for years. What typically limits pushing the state of the art is the availability of "labeled" data and the finances required for very large scale trainings. With NLP, there are huge datasets available which are effectively in the form of supervised data, since humans have 1) invented a meaningful and descriptive language and 2) generated hundreds of trillions of words in the form of coherent sentences and storylines. The task of predicting a missing word is a well-defined supervised task for which then there is effectively infinite "labeled" data. Couple these facts with a large amount of compute credits and the right architecture and you get GPT3. The results are really cool but in my opinion scientifically unsurprising. GPT3 is effectively an example of just how far we can currently push supervised deep learning, and even if we could get truly human level language understanding asymptotically with this method, it may not get us much closer to AGI, if only because not every application will have this much data available, certainly not in a neatly packaged supervised representation like language (such as computer vision). While approaches like GPT3 will continue to improve the state of the art and teach us new things by essentially treating NLP or other problems as an "overdetermined" system of equations, these approaches are subject to diminishing returns and the path to AGI may well require cracking that human ability to create and learn with a vastly better sample complexity, effectively operating in a completely different "under-sampled" regime.
- LoSboccacc 6y agogp3 is an impressive technical feat and the pinnacle of the current line of research however, if you remove the technical colored glasses and boil down what it is and what it does, it's a regurgitation of existing data that it had been fed, it has no understanding of the data itself beyond linguistic patterns. it's not going to find correlations where there were none, it's not going to actually discover new data, it will find unexpected correlations between data but there's zero indication whether these correlation bear any significance until a human goes validate the prompt, and it can generate infinite of these, making the discovery of significant new ideas pretty slim.
- Bx6667 6y agoAnd precisely none of what you just said addresses my point.
- LoSboccacc 6y ago> The point is that we are getting unintuitive and unexpected results. > it will find unexpected correlations between data but there's zero indication whether these correlation bear any significance until a human goes validate the prompt, and it can generate infinite of these, making the discovery of significant new ideas pretty slim seems a pretty direct response tbh
- Bx6667 6y agoNo, you are missing the point completely
- LoSboccacc 6y agoit isn't really helpful or conductive of an interesting discussion that you are not detailing what the point is, even when the "the point is this" gets directly quoted, while not clarifying neither the point nor why the reply don't apply.
- 6y ago
- baryphonic 6y agoI'll tell you why I'm not impressed. We can't keep doubling, er, increasing model size by two orders of magnitude forever for iterative improvements in quality. (Maybe this is a Malthusian law of the nothing-but-deep-learning AI approach: parameters increase geometrically, quality increases arithmetically.) This is an achievement, but is not doing more with less. When someone refines GPT-3 down to a form that can be run on a regular machine again (hint: probably a new architecture), then that will be genuinely exciting. I also want to address this point directly: > We are digging deeper and deeper into “algorithm space” and we keep hitting stuff that we thought was impossible and it’s going to keep happening and it’s going to lead very quickly to things that are too important and dangerous to dismiss. I hope the above convinced you that this is basically not possible with current approaches. OpenAI spent approximately $12M just in computation cost training this model (no one knows how much they spent on training previous iterations that did not succeed). Running this at scale only for inference is also an extremely expensive proposition (I've joked with others about many tenths of a degree Celsius GPT-3aaS will contribute to climate change). If we extrapolate out, GPT-4 will be a billion dollar model with tens of trillions of parameters, and we might get a dozen pages of so of generated text that may or maybe not resemble 8chan! > People who say AGI is a hundred years away also said GO was 50 years away and they certainly didn’t predict anything even close to what we are seeing now so why is everyone believing them? Isn't this a bit too ad hominem? And not even particularly good ad hominem. I'm sure there existed people on the eve of AlphaGO saying it would be another 50 years, but there's no evidence that the set of these people is the same as those saying AGI is 50-100 years away. How many people made this particular claim? I, for one, made no predictions about Go's feasibility (mostly because I have never thought that playing games is synonymous with intelligence and so mostly didn't find it interesting) but absolutely subscribe to the 50-100 year timeline for AGI. Think about it like this: Go is a well-defined problem with a well-defined success criterion. AGI has neither of those properties. We don't even understand what intelligence is enough to answer those questions. Life took billions of years to achieve landing on the Moon and building GPT-3. It's not far-fetched that it'll take us at least 100 more using directed research (as opposed to randomness) to learn those same lessons.