13 ms·
Turing-NLG: A 17B-parameter language model
- galkk 7y agoThose summaries look impressive, although a bit repepetive
- Schiphol 7y agoI see what you did there?
- aabhay 7y agoI see what you see what I did there
- RobertDeNiro 7y agoWhat they don't tell you is that these summaries are always hand picked from a few that were generated.
- XnoiVeX 7y agoQuite possible but that also means that there is an opportunity to implement some sort of RL to choose the best possible summary.
- corporateslave5 7y agoPeople are vastly underestimating the changes that are about to come from NLP. The basic ideas of how to get language models working are just about in place. Transformer networks, and recent innovations like GPT-2, googles reformer model, etc are precursors to the real machine learning boom. Machine learning as we have known it, has been stuck as an optimization tool, and used for computer vision here and there. NLP, and with it, the ability to create, synthesize, and understand content, will change the internet. More than that, I think NLP will unlock new ways of interacting with computers. Computers will be able to handle the ambiguity of human language, transcending their rigid “only do exactly what you tell them” models of the world. Edit: Adding this to give more technical context. I think most people don’t know where the line is currently between what possible, and what’s not, but also what we are on the cusp of. And we are on the cusp of a lot. A quick explanation of one area is here: Basically, transformer models are the best for NLP. They use something called attention based mechanisms, which allows the model to draw correlations between pieces of text/tokens that are far apart. The issue is that this is an O(n^2) operation. So the model is bounded by the context window, which is currently mostly at 512 tokens, and is thus, bounded in how much it can understand. Recent innovations, and further study, will broaden the context window, and thus unlock better reading comprehension and context understanding. For instance, the ability to answer a question using a piece of text is mostly stuck at just finding one paragraph. The future will see models that can find multiple different paragraphs, understand how they relate, pull the relevant information, and synthesize it. This sounds like a minor step forwards, but its important. This will unlock better conversational abilities, but also, better ways to understand how different pieces of textual information relate. The scattershot of information across the internet can go away. Computers can better understand context to act on human intention through language, unlocking the ability to handle ambiguity. This will change the internet. Again to empathize, these models only started showing up in 2017! The progress has been rapid.
- gzu 7y agoI can't wait for the day we see "deep-dream" styled literary works.
- jandrese 7y agoThere have been plenty of "AI written" textual works, like the AI Dungeon for a D&D game, or the AI Recipe Generator that attempts to make something that looks like a cooking recipe. For the most part they aren't successful because the "AI" isn't smart enough to have a goal in mind, so they end up just monkey-cheesing everything. Pasting together snippets in ways that are usually grammatically correct but make no sense.
- okigan 7y agoany books/papers that address "goal" oriented NLP or hybrid based discussion?
- jandrese 7y agoThe ability to set your own goals and task yourself to achieve them is the essence of AI. Not "AI" as we know it today, but Sci-fi AI where it's a machine person.
- joe_the_user 7y agoIndeed, but that kind of describes the gulf between current language processing and present AI. Present AI generates tokens that seem to have meaning or seem like an appropriate response to a statement but where it become evident after 2-3 paragraphs, there's no substantial relation to either underlying meaning or underlying goals. Part of this is "underlying meaning" is an intuitive way to describe things but whatever is underlying here is more tenuous than a classical logic/GOFAI model of the world but more "solid" than a long, clever stream of associations.
- sgt101 7y ago
- saurkt 7y agoOne of the team members from Project Turing. Happy to answer any questions.
- osipov 7y agoWhy the lack of number on the more popular SQuAD and Glue benchmarks?
- saurkt 7y agoSQUAD and GLUE are tasks for language representation models -- aka BERT-like. This is a language generation model -- GPT-like. Hence, SQUAD/GLUE test sets are not really applicable. We are reporting on the wikitext and lambada sets that openAI also uses for similar models (numbers are in the blogpost).
- igravious 7y agoWhat's the difference between the two models?
- octbash 7y agoOne is a language generation model, the other is a fill-in-the-blank model. It sounds like they might be similar, but in practice they are different enough objectives (and in particular the "bi-directional" aspect of BERT-type models) that the models learn different things.
- Voloskaya 7y ago* BERT & language representation models: They basically turn a sentence into a compact vector that represents it so you can then do some downstream task on it such as sentiment detection, or matching the similarity between two sentences etc. * GPT & language generation models: Given some context (say a sentence), they can generate text to complete it, or to summarize it, etc. The task here is to actually write something.
- riku_iki 7y agounfortunately they abstained from participation in more popular SQuAD and Glue benchmarks..
- deleted 7y ago[deleted]
- octbash 7y agoThose are question-answering and language-understanding benchmarks respectively, neither of which has been suitable for language generation mode evaluation since GPT-1 was roundly beating by BERT. GPT-2 didn't evaluate on them either.
- 01100011 7y agoHow long until the language models stabilize enough that we can bake them into a low-cost, low-power chip for edge uses?
- foota 7y agoI think this is largely unnecessary, can't things like TPUs handle the inference?
- slashcom 7y agoThere’s lots of work on distillation, smaller models, approximations, etc. People already have simpler forms of these running on smartphones. Models seem to be growing faster than we can make them small though :D
- saurkt 7y agoYes, these things keep us up at night as well :-).
- 0xff00ffee 7y agoB = Billion, not Byte. For second I was like, WTF?
- deleted 7y ago[deleted]
- Tenoke 7y agoI expect we'll see some very interesting, very big models following it. I didn't dig too far into the code but the library looks very easy to use and will open up a lot of doors for people who have a few or a few thousand GPUs.
- lowdose 7y agoThis does GPT-2 X 10. For anyone wondering what GPT-2 is doing look at this baffling subreddit and marvel at how one GPT-2 model trained for $70k spits out better comedy than everybody on the payroll of Netflix combined. https://www.reddit.com/r/SubSimulatorGPT2/ https://www.reddit.com/r/SubSimulatorGPT2/
- jamiek88 7y agohttps://www.reddit.com/r/SubSimulatorGPT2/comments/f7dlec/do_you_know_what_the_world_would_be_like_if/ https://www.reddit.com/r/SubSimulatorGPT2/comments/f7dlec/do... That thread title and posts were stunning! This is a great niche for GPT-2
- jeromebaek 7y agogosh, this is unsettling. my brain is literally getting stuck on an infinite loop trying to read this and coherently put them together
- woodgrainz 7y agoComedy is right. "I'm really starting to get worried about my Higgs Boson (HBN) after watching some videos on YouTube" [0] "This repost and all of your posts are garbage." [1] "It's the most random and unoriginal shit I've ever read." [2] [0] https://www.reddit.com/r/SubSimulatorGPT2/comments/f1sqyh/any_idea_if_the_new_higgs_boson_will_be_stable/ https://www.reddit.com/r/SubSimulatorGPT2/comments/f1sqyh/an... [1] https://www.reddit.com/r/SubSimulatorGPT2/comments/f1vfnv/why_did_the_doctor_put_the_cat_back_into_the_bag/ https://www.reddit.com/r/SubSimulatorGPT2/comments/f1vfnv/wh... [2] https://www.reddit.com/r/SubSimulatorGPT2/comments/f1o83a/found_this_gem_on_a_facebook_post_in_the_middle/ https://www.reddit.com/r/SubSimulatorGPT2/comments/f1o83a/fo...
- SkyBelow 7y agoMy favorite so far is the science bot. 75% of the top level posts are explaining why the submission was removed.
- speedgoose 7y agoIt's definitely better than the original using Markov chains. It fits very well this use case, and in my opinion only this use case. GPT2 is still very random and quite stupid. You start it with your love for your girlfriend as a context, she becomes a cam girl into hard core anal two paragraphs later. You start with religion, "Muslims must be exterminated". You start with software and you get a description of non existent hardware with instructions about how to setup a VPN in the middle. You start with news, and you can read than China supports the Islamic state. That's cool because it has more context than Markov chains which usually have only 3 words of context, but it's still a long way to go before I trust anything generated by this kind of algorithm.
- deleted 7y ago[deleted]
- danharaj 7y agoIf your model has 17 billion parameters, you missed some.
- undoware 7y agoAs a trans woman I can't help but notice that you rank me being your girlfriend at about the same level of beyond-the-pale-ness as a genocide. The feeling is mutual
- speedgoose 7y agoSorry I didn't think one could interpret it like this, this was inappropriate indeed. I edited.
- undoware 7y agoThank you <3 For reference, this reasonable, human exchange cost me 4 karma points, because as everyone knows, being a human on Hacker News is a 403.
- bitL 7y agoWhat GPU do I need to train it? Titan Mega RTX with 240GB of RAM?
- slashcom 7y agoA DGX-2 will do just fine.
- freediver 7y agoI can uderstand announcing this without code, but without a demo so anyone can try it in different scenarios?
- saurkt 7y agoIf you want access, please send an email to [turing_ AT _microsoft _DOT_ com]. Remove underscores and spaces.
- rjeli 7y agoI have been bearish on AGI, but GPT2 surprised me with the lucidity of its samples. My take from the past few years is that we're 99% done with the visual cortex - convolutional nets can be trained to perform any visual task a human can in <100ms. Now I'm mostly convinced that GPT2 has solved the language cortex, and can babble as well as we will ever need it to. We just need a prefrontal cortex (symbolic processing / RL / whatever your pet theory is) to drive the components, which is a problem we have not even started to solve. I am 90% sure it is a different class of problem and we won't knock it out of the park in 5 years like the visual/language cortexes, but we can hope. edit: it's possible cognition follows from language, which would be convenient. is GPT2 smarter than a dog? I don't think so but I could be wrong ¯\_(ツ)_/¯
- kragen 7y agoI have been bearish on AGI, but GPT2 surprised me with the lucidity of both paths. I still maintain my support for the basic metric of the GPT-I. However, I have a number of requirements on how my proposal is to be funded to resolve concerns. First, I strongly believe that academic research should be the method of choice (that is, if we are to figure out how to make AGI possible), and I advocate funding to support results from the central bank community. Second, given that the GPT can be articulated in mathematical terms, this should be reflected in funding policy. A very serious concern is that if funding of GPT is disincentivized, investors may react similarly to the way they reacted to AGI. This is
- eyegor 7y agoI've always been interested in techniques to try to minimize parameters or alternate approaches to learning. Meanwhile, state of the art is over here just finding clever ways to make everything bigger. I have a feeling we're going to end up with a very different landscape in 5-10 years, much like the automotive industry never started mass producing inline 12s and instead moved to turbos and superchargers.
- deleted 7y ago[deleted]
- ragebol 7y agoAll these language generation models, in short, base their next word solely on the previous words, right? I'd expect that these generators can be conditioned on e.g. some fact (like in first order logic etc) to express something I want. This is roughly the inverse of for example Natural Language Understanding. Does anything like this exist?
- gchq-7703 7y agoI'm fairly sure that these models don't work solely on the previous word, but instead are able to remember some level of information from history. Otherwise, you'd reach a word like 'and' and couldn't possibly follow it with a logical statement that follows on from the previous part.
- ragebol 7y agoThis is why I said 'words', multiple :-). My point being that these generation models should be conditioned on something more than just word history, like something they want/are instructed to express.
- FlyingCocoon 7y agoAt what stage of throwing compute & data at the problem, diminishing return sets in?
- tuxguy 7y agohttps://news.ycombinator.com/item?id=22291417 https://news.ycombinator.com/item?id=22291417