11 ms·
OpenGPT-2: We Replicated GPT-2 Because You Can Too
- p1esk 7y agoThey spent $500k replicating it. But sure, you can do it too /s
- gwern 7y agoThey used research credits, and even that aside, with their code and training tips, you can redo it for $50k on cloud instances or less on dedicated hardware + patience. And look at ImageNet training progress: you can train a near-SOTA ImageNet CNN in like a minute for $20-40 after a lot of optimization work. We've already seen a lot of improvements in LMs over the past 2 years... (For example, the main barrier to training GPT-2 is just the bloody memory use from the Transformers exploding at runtime, which pushes you into high-end hardware like cloud TPUs on GCP. Do Sparse Transformers fix that?)
- p1esk 7y agoWait, how can I get to near SOTA on Imagenet in a minute (!) for $40?
- gwern 7y agoOK, I exaggerated a little because I was recalling from memory: the old fast.ai approach actually takes <18 minutes (https://www.fast.ai/2018/08/10/fastai-diu-imagenet/ https://www.fast.ai/2018/08/10/fastai-diu-imagenet/). My bad. (I'm sure it's improved since then but I don't know how much.) I was also thinking of https://myrtle.ai/how-to-train-your-resnet-8-bag-of-tricks/ https://myrtle.ai/how-to-train-your-resnet-8-bag-of-tricks/ which does CIFAR-10 in 26s but I'm not sure offhand what CIFAR-10's SOTAs look like so not sure how far away that is.
- p1esk 7y agoActually this is still very good. Thanks for the links. I'll be timing some of these tricks tomorrow for my Imagenet experiments. By the way, I believe this is the current SOTA for Imagenet: https://arxiv.org/abs/1905.11946 https://arxiv.org/abs/1905.11946 (84/97%). CIFAR10 appears to be essentially solved (99%).
- ZhuanXia 7y agoAre you going to update your poetry engine now that we have this?
- gwern 7y agoHm, maybe. It depends on how easy their training code is to use and how long retraining would take. It presumably will take at least a week because 345M took about a week, but I'm not sure I want to spend the money on a week of a very large cloud instance (which would be what, $300?) for what is probably a substantial but not stunning improvement in generation quality. I might rather wait for the next leap, from something like a Sparse Transformer approach which can get global coherency by having a lookback over the entire poem or getting a better poetry corpus with delimited poems (rather than entire books).
- the8472 7y agoYou're off by an order of magnitude and omit the caveats to that cost estimation. From the article: > The cost of training the model from scratch using our code is about $50k. It’s important to note this figure is the estimated value of the cloud compute, and does not reflect the much smaller intrinsic costs involved
- p1esk 7y agoThey edited the article after I left a comment there. The original text stated they spent $500k to run all the hyperparameter search experiments to replicate OpenAI results. Only after they did all that work you can run their code for $50k.
- kuzehanka 7y agoWhere did you get 500k from? They said 50k. In estimated cloud compute costs.
- p1esk 7y agoThey removed this info for some reason. It takes $50k per training run, and they initially said they spent $500k total on experiments. Only after they did all that work you can run their code for $50k.
- Felz 7y agoWhat's it take to actually run a model like this, hardware-wise? I've been toying around with a gpt2 discord bot (https://github.com/ScottPeterJohnson/gpt2-discord https://github.com/ScottPeterJohnson/gpt2-discord) using just a CPU calculation, and already it takes up 2 GB RAM (and is slow obviously) on the 345M model. I might be able to get the 774M model running, but there's no way I can afford the full model, assuming linear RAM use. And that's just for CPU compute, I can't even begin to imagine how expensive GPU would be.
- ageitgey 7y agoInferencing on this model works fine on Google Colab which gives Tesla K80 GPU with access to 12GB of GPU RAM. You can buy a used K80 for probably about $850, but it's not really ideal for putting in a home computer because of the cooling requirements. [ deleted reference to 2070 Super ]
- acd10j 7y agosince when RTX 2070 ship with 14GB of GPU Ram, Max memory for RTX 2070 super is 8 GB.
- ageitgey 7y agoOops, you are correct. I mis-read the spec sheet.
- p1esk 7y agoUsed K80 can be had for $350 [1] Not bad actually (it's probably as fast as 1080Ti, and has 24GB of memory). https://www.ebay.com/itm/NVIDIA-Tesla-K80-GDDR5-24GB-CUDA-PCI-e-GPU-Computing-Accelerator-Card/193033951488 https://www.ebay.com/itm/NVIDIA-Tesla-K80-GDDR5-24GB-CUDA-PC...
- happycube 7y agoK80 is 2 GPU chips with 12GB, so it's not always as good as one newer/larger GPU. Much more affordable though :)
- macawfish 7y agoPrompt: The jig is up. And what now? Where do we go from here? Completion: Where Do We Go From Here: In the aftermath of the fall of the German Republican party, we now have a significant degree of instability across the earth’s systems of government and finance. The almost complete collapse of systemic forces in the Eurozone and limited success at stabilizing the system means the question is not if but when, what do we do next? The answer is simple. We must move beyond the localized, bubble-like, and short-termist “get involved,” tactic of getting into the scene and trying to control it in some way. We have to come up with a way of shifting the socio-political power in the world, the prime place for transformation is worldwide at the supra-system level and not just the economy and finance. We must cast out the old dominated system, of which we have been just a part and recognize that we need a new dominant system that serves human interests, and the meta-level global system must serve human interests. The fact that the global status quo is collapsing of its own weight shows us that the system is structured in such a way that the group of big players who have dominated and still dominate, are in ever-decreasing danger of losing both power and integrity. The question now is, how? How do we avoid degeneration into chaos and conflict when the anarchic nature of the system leads inevitably to greater and greater competition among and frustration and anger in the younger generations? This is a society, this is a planet and we live in the first global century of human history, which the young will pass from generation to generation in the next twenty years, or perhaps not. When we see the events of the last weeks and days, you can just imagine what will happen to this planet, to this planet and human society in the century ahead, and you can just imagine what the future will bring to this and subsequent generations. When history describes the past, it sees the collapse of an old political establishment, of the traditional hierarchies of power, of society, of economics and finance. It sees a collapse in the old order of power and in the equilibrium it has created, which was grounded in constant growing jobs and the prosperity it produced. We are in the middle of a permanent expansion of capitalism, which also creates ever-growing wealth and prosperity for a small population of wealthy earners, while social polarization and inequality increase and older people depend on each other more and more desperately. The forward and downward momentum of all these forces has created a situation in which there is almost no limit to the volume of the day to day, or minute to minute production and consumption, and in which there is no single concern about the future of the planet Earth. We have become so insatiable, in need, addicted to this ever increasing appetite for consumer goods, that we destroy the planet with it. You see this just from what we feed our children, the choices we make, and the products we consume. You see it in our greedy attempts to buy as much as we can, even if it leads to ecological ruin. You see it in our drive to consume new and ever more lavish luxury products, materials, tools, devices, insatiable lifestyles, modern-day imperialism, racism, cynicism, competition, greed, consumerism, hubris, and endless pursuit of personal ambitions and leisure. See how when push comes to shove, the social and economic growth created by the continued expansion of capitalism is now a life or death matter. See how the political establishment has failed us, all of us, and how we turned in desperation to another self-serving-self-protective-petty, self-interested-philistine mass-mediator, in the form of Mr. Romney, in order to maintain the old sources of power, to make the old social structures fit to serve human needs and the system could be kept going. And now he has packed his bags and wants to leave, so there we are, stuck here with those of us who have found a way to provide for ourselves and live peacefully and prosperously, without the brutality and violence visited on us by politicians and corrupted systems. That is, unless we fix these broken systems and deliver an alternative based on human needs and human compassion. How do we do it? How do we get there? Stay tuned and we’ll let you know.<|endoftext|>This exciting book is an overview of a phenomenon that started in the 1970’s, and became the most spectacular of all the urban myths. It combines all things the paranormal in this feature length book, from scientists to aliens to experimental reports. William Kean is an astronomer working for NASA. One evening he is on his way to a remote overlook on a Martian hill. Suddenly, he is teleported to the top of a fifty story building, two thousand feet in the air. The building
- macawfish 7y agoIt's fun to feed it the first paragraph of the Sokal affair paper: Prompt: There are many natural scientists, and especially physicists, who continue to reject the notion that the disciplines concerned with social and cultural criticism can have anything to contribute, except perhaps peripherally, to their research. Still less are they receptive to the idea that the very foundations of their worldview must be revised or rebuilt in the light of such criticism. Rather, they cling to the dogma imposed by the long post-Enlightenment hegemony over the Western intellectual outlook, which can be summarized briefly as follows: that there exists an external world, whose properties are independent of any individual human being and indeed of humanity as a whole; that these properties are encoded in ``eternal'' physical laws; and that human beings can obtain reliable, albeit imperfect and tentative, knowledge of these laws by hewing to the ``objective'' procedures and epistemological strictures prescribed by the (so-called) scientific method. Completion: Before one can explore these ``objective'' principles and procedures, one must do something important, namely, invent the objectivity -- or, more precisely, the objectivity-transcendent properties -- of the ``subjective'' functions of ``man" (as a species of ''objective'', or something-in-itself-independent, human, being). Since the 19th century, this has been done on a literal or phenomenological level, the metanarrative of psychoanalysis. However, the ultimate test of what constitutes an objective fact or structure is the statistical discovery of reliable facts or structures residing in the human mind or body. That is what the discipline of psychology is for. The distinction between the subjective and the objective is therefore necessary for the discipline of psychological science, which is by no means a waste of time, which is today a flourishing field of scientific inquiry, and which might contribute a great deal toward furnishing a genuinely complete theoretical understanding of human behavior and psychical experience. In short, the philosophical implications of the post-Enlightenment philosophy of the sciences and of psychical psychology are -- when considered in the context of both human existence and human thought -- deeply unsatisfactory. Once we abandon the most basic postulate of scientific rationality -- that subjective knowledge and belief and behavior are not to be said to be existent but to be, not to be, knowable -- there is nothing left to suggest that the content of human knowledge and experience is dependent upon our conscious perceptions or personalities. We might describe that content as ineffable (it is irreducible to language, outside the experience of individual thinkers), or inconceivable (that which is not one's own consciousness, but is beyond the realm of conscious experience). In either case, there is no genuine need for an objective world, or, at least, no need for an objective foundation. In that sense, it is erroneous to assert that the scientific standpoint is a paradigm of objective reality; rather, it is the antidote to the paradigm of irreducible subjectivity and self-consciousness. Unsupported by this analysis of the problem of scientific knowledge, the field of psychical science is often rejected by younger scholars, for whom the study of the human condition and its ineffable psychology is regarded as a particularly enticing but very unapproachable form of mysticism. They regard this section of the theory and practice of modern scientific inquiry as a kind of metaphysical theory-cracking rather than as a vehicle for the discovery of objective facts. Such a view can only be sustained on the basis of the strictest adherence to the doctrines of empiricism and fundamental materialism, respectively, and it is by no means fruitful for the real progress of science. Let me explain. The residual emphasis on irreducible subjectivity in modern Western thought (and only by the residual emphasis) is a legacy of the metaphysical concept of objective reality which, to my mind, is still embraced by the cognitivist movement of the 20th century. It is an unfortunate legacy. Irreducible subjectivity and subjectivity-dependent, subjective perceptions were at the root of the metaphysics of Buddhism, and these beliefs now constitute the dominant conception of reality in contemporary Western philosophy, particularly Kantian, Descartesian, and Hegelian thought, which are each deeply indebted to Aristotelian theories of subjectivity. While both these views of reality are patently absurd, and since the scientific revolutionary of the middle 19th century was able to disregard or repudiate them, the followers of Kant and Descartes were able to maintain that there is no need for a foundation for the science of knowledge. They could do this because they held to a primitive, problematic conception of objectivity, based on the notion of an objective, external world, in which human consciousness, thus independent of any particular body, mind, or culture, was inchoate, mutable, and subject to change or speculation. There was therefore no need to search for a theory of experience. Science and experience were simply different approaches, of which each was as good as the other, and they both...
- steve19 7y agoIt sounds like all the drama OpenAI made about not releasing the model was all just marketing. $50,000 is nothing for a nation-state or even just a motivated third party. I had always assumed OpenAI had spent well into the 6 or even 7 figures to train the full model. MSFT has sort of invested $1 billion into OpenAI so I guess it worked!
- skrebbel 7y agoI have a beef with this. OpenAI is run by people who believe and spread horror stories about how AI can totally change society, for better or worse. A lot of what AI can and cannot do depends on cost and computational power as much as anything else. If I understand correctly, there's a whole bunch of "you could but it's prohibitively expensive" stuff going around. I'm convinced that the people behind OpenAI understand this nuance, even if most non geeks struggle with the difference between possible and feasible. This means that OpenAI mixed the two up on purpose. If the people warning us of the robot apocalypse can't be trusted to communicate openly and honestly about odds, then what do we do? They clearly have no qualms about misinforming the public to serve a hidden agenda. It might be pretty harmless in this case, but it's indicative of a cultural pattern. Basically I just hope (and kind of believe) that we'll never achieve AGI because that would make all this talk unimportant.
- repolfx 7y agoOpenAI is run by people who believe and spread horror stories about how AI can totally change society, for better or worse. I think this is reflective of a substantial disconnect and polarisation in society, because OpenAI's position here was driven by a very fundamental intuition about human nature which is not universally shared. Namely, their rationale was something like this: the world is dominated by people whose opinions are fundamentally shaped by what others are saying, and moreover, by how frequently other people seem to be saying it. They don't really think about things or rely on their own experience: they're just mimics. In such a world being able to auto-generate fake messages of support for some political position or another at scale would give you immense power, because the population would automatically swing behind you based merely on the perception that everyone was swinging behind you. So that's their fear. But is it realistic? Well, this is where we get into the polarisation. "You aren't smart enough to have an opinion" is the sort of viewpoint that leads to an elitist vs populist conflict. There have been reams of analysis about this and it's not really AI specific - e.g. did people vote for Brexit because of Twitter, or because of things they saw in the press, or did they vote based on their own experiences, or what their close friends/family thought, or what mix of "all of the above" is the truest mix? When they refused to release GPT-2 the OpenAI researchers took an extreme on that spectrum, asserting that essentially they had built a mind control device. But of course which didn't work on themselves, only on lesser minds. Not surprisingly this was very controversial, albeit I think the Valley/Hacker News set would be surprised at just how controversial it would have been if anyone outside the AI world had really noticed. You can't tell most of the world their opinions aren't really their own and not expect pushback. Personally I don't share those intuitions at all. The controversy over DeepFakes is similar. Photoshop has existed for decades and hasn't led to a dystopia, but suddenly DeepFakes is going to create fundamental social change? I don't think so. AI researchers, especially in academia, need big narratives to keep up the funding and suggest they're on a mission to save the world vs e.g. making slightly better playlist recommendations. GPT-2 seems like a classic case of this being taken to absurdity.
- Chris2048 7y agoI skimmed the article and stated reading the last paragraph to get an idea of what it was about. I was v confused..
- high_derivative 7y agoIt may be high time to discuss what AI policy has actually done so far. From what I can tell, not much other than letting social scientists get in on the deep learning gravy train. Meanwhile, misuses of ML are proliferating without limits, and 'AI policy' is apparently mostly used as a fig-leaf to collect good-will, marketing, and buy a seat at the table for future regulations. As usual, regulations will protect incumbents, so my as-usual cynical read is that OpenAI's policy interests are about protecting its own future interests. From that perspective, the entire GPT-2 stunt was highly effective. Now depending on your outlook, that may be an argument that we need more people in policy, or fewer. Or different ones.
- buboard 7y agoThe biggest offenders of AI are of course governments. AI is military technology. Yet we re arguing about the merits of gender inclusion.
- Dobbs 7y agoAnd banks, and large rental companies, and credit lenders, and Amazon, and Facebook, and Google, and so many other people. So many of which easily encapsulate their biases into their algorithms, which then perpetuate the biases into society as a whole.
- ethbro 7y agoWhich is why regulation should be model transparency in sensitive areas: full stop. Sure, use it to decrease your operating costs. But it doesn't get to be secret sauce.
- godelski 7y ago> Which is why regulation should be model transparency in sensitive areas: full stop. It's not the model that makes the bias. It's the data. Garbage in, garbage out. The tricky situation is that's the magic sauce to most ML algos (as really GPT2 demonstrates). I don't care about what model they use (CNN, transformer, whatever) or their architecture (layers and how things are connected). What I'm concerned about is the implicit bias in the data. This is the hardest things to figure out and so the issue is how do you build trust of that data? Handing out that data makes you lose your competitive edge in a marketplace. Not only that, in many cases that data contains sensitive information that people wouldn't want public. So it's hard to say what to do. 3rd party auditors? But how do you prevent them from becoming corrupt and complicit like the credit agencies. I think it sounds like there's easy answers, but honestly this seems really difficult to me.
- anon1253 7y ago"The cost of training the model from scratch using our code is about $50k." Still a substantially steep curve for a bootstrapping startup. It's something I continually run into myself. I have somewhat of a weekend project trying to build a search engine but man ... the cost of just the SSDs and GPUs is daunting on a regular salary. As the complexity of these models grows, so does the barrier to entry for a regular joe like me; which is a shame I think. I know in the US it's fairly normal for a data scientist to pull 100k+ / year, but in the Netherlands salaries pretty much stall at 40k (and angel investment in IT/AI is at an all time low). More generally I fear this will become a bit of a sociotechnical issue if complex AI models will be out of reach for entire economies (especially for cases like language because not everyone speaks English and "minor" languages like those in EU countries are a massive market to explore, yet hard to get into).
- buboard 7y agoThere is no need to train the model, they already provide the parameters. These transformer models , like BERT are pretty adaptable to reuse.
- anon1253 7y agoThere is no need to train this particular model, but adopting this (or any novel model, the field is moving fast) to say Dutch, Italian, Hungarian, Icelandic or whatever still requires training. Luckily for most languages they are provided (at least in the case of BERT, FastText, or regular skipgram). But there is also still quite a bit of leeway in domain specific adoption (for example SciBERT for scientific texts, or legal and financial documents) reddit/wikipedia does carry a bias. Each of which not only requires pretraining the model, but also generating a huge and fairly well formatted corpus. And, although the parameters are usually finetunable, it does break down sometimes on the various sub-word tokenizations used.
- tzapzoor 7y agoAnd they're just "two masters students, with no prior experience in language modeling" with $50k lying around for training a huge model.
- exabrial 7y agoWithout context, this article reads like something generated by machine learning.
- 6gvONxR4sf7o 7y agoIt's worth noting that, like other attempted replications, the perplexities of this model mostly aren't as good as GPT-2. Given that the title of the GPT-2 paper was "Language Models are Unsupervised Multitask Learners," I'd be interested in a lot more metrics before I'd believe GPT-2 has actually been replicated. Especially because every other time someone says this, metrics show otherwise. Until then, this is just a really big model.
- deleted 7y ago[deleted]
- happycube 7y agoThe results are (mostly) between the large and extra large GPT2 models, and it's possible to reproduce if you have the resources. Between this and the 774M GPT2 release, it's been a pretty good week :)
- doctorpangloss 7y agoI suppose you can go and shit on people trying to replicate scientific work, and telling them, extremely reductively, “Well your almost as big number isn’t as big, so fuck you,” like a soccer ultra comparing their team’s score of 3 versus the opposing team’s score of 2, as though the only question that matters is “Whose Soccer Team is Best?” Is that even the right question to ask? People who actually do research, they don’t just look at the absolute comparison of published numbers! Do you think that’s how research is done, by chasing whatever has the biggest number? No repeat innovator does that. It’s an interesting collision of world views for sure. This is a social media forum for a venture capital firm. They’d hate for anyone to discover that the numbers don’t tell the whole story, that actually everybody starts at zero, and that being second, because of the price premium put on first, is a huge opportunity. So even in some narrow, cynical interpretation, your point of view would lose people a ton of money. But I don’t really know anything about that.
- 6gvONxR4sf7o 7y agoHoly cow, dude. That's not even remotely like what I said. If you read the original paper, they show that GPT-2 does all sorts of cool things. In OP, they show that it's almost as good at one thing on a couple data sets. It's like if I wrote a paper showing that widgets improve liver health in young men, young women, and adult men. Additionally, widgets make you happy and taller and turn blue. Then you come along and try to replicate my results, showing only that your version of a widget makes young men and women's livers almost as healthy as mine did. But sure, pretend I said whatever you want to argue against.
- minimaxir 7y agoTwitter thread by a Research Scientist at OpenAI addressing OpenAI's policies in response to this discussion here: https://twitter.com/Miles_Brundage/status/1164959322633318400 https://twitter.com/Miles_Brundage/status/116495932263331840...