17 ms·
OPT: Open Pre-trained Transformer Language Models
- coding123 4y agoCan someone open a Bittorrent seed if you get it
- MasterScrat 4y ago"We are releasing all of our models between 125M and 30B parameters, and will provide full research access to OPT-175B upon request. Access will be granted to academic researchers; those affiliated with organizations in government, civil society, and academia; and those in industry research laboratories." GPT-3 Davinci ("the" GPT-3) is 175B. The repository will be open "First thing in AM" (https://twitter.com/stephenroller/status/1521302841276645376 https://twitter.com/stephenroller/status/1521302841276645376): https://github.com/facebookresearch/metaseq/ https://github.com/facebookresearch/metaseq/
- ALittleLight 4y agoI don't like "available on request". I just want to download it and see if I can get it to run and mess around with it a bit. Why do I have to request anything? And I'm not an academic or researcher, so will they accept my random request? I'm also curious to know what the minimum requirements are to get this to run in inference mode.
- javchz 4y agoMy bet it's probably a filter, trying to prevent create a even more realistic farmbots in social media, as they are already bad as they are now.
- robonerd 4y agoBut they'll consider requests from government and industry.. both greater threats in the information war than any private individual.
- xxpor 4y agoNot from their perspective
- robonerd 4y agoOf course. To somebody in Zuck's position, shoring up the power of the status quo is common sense.
- alimov 4y agoI don’t know whether this is true and have no way of knowing this with any degree of certainty, but to me it seems unlikely that Mark had anything to do with this stipulation (requesting access). Although it’s not unimaginable.
- IAmEveryone 4y agoSince “everyone” would include governments and industry as well, their restriction is guaranteed to not contain more bad actors than no restriction.
- bogwog 4y agoThat is dumb when you consider that this thing is likely going to leak anyways. It’s inevitable, and when it does happen, it will just end up in the hands of criminals/scammers and not the general public.
- londons_explore 4y agoIt's super easy to watermark weights for ML models. Just add a random 0.01 to a random weight anywhere in the network. It will have very little impact on the results, but will mean you can identify who leaked the weights.
- ipaddr 4y agoCompare two copies.
- vladf 4y agoSlightly modify a million random weights by changing the least significant bit up or down.
- ipaddr 4y agoCompare three copies.
- ALittleLight 4y agoOr slightly randomly modify all the parameters on the copy you distribute, then it will be a match for nobody.
- ipaddr 4y agoYou compare all three and average the variance of each value. So the more copies the better.
- ZephyrBlu 4y agoCouple of random ideas: - They are concerned about the usage of the largest model, so want to vet people - The 175B parameter model is so large that it doesn't play nice with GitHub or something along those lines
- metadat 4y agoEnding up in the wild is an eventuality, whether FB creates it or someone else, why draw it out? Bandwidth concerns is nonsensical these days, fb has nearly unlimited resources in that department. Set it free! It wants to be free.
- deleted 4y ago[deleted]
- jquery 4y agoThis is an ideal use case for a torrent.
- sanxiyn 4y ago"It wants to be free" is a ridiculous statement, considering that after full two years (GPT-3 was published in May 2020), there is no public release of anything comparable. In May 2020, was your estimate of time to public release of anything comparable shorter or longer than two years? I bet it was shorter.
- HWR_14 4y ago> "It wants to be free" is a ridiculous statement "It wants to be free" is based on the standard line "code/data wants to be free". It doesn't mean this cost nothing to produce or isn't valuable.
- londons_explore 4y agoIn big companies, something as simple as "host it on facebook.com/model.tar.gz" can be mountains of approval and paperwork.
- 4y ago
- dukeofdoom 4y agoTo prevent someone from building something that returns certain inferences that might be true but are politically taboo.
- bestcoder69 4y agoYou think GPT-3 generates text that's truthful? Have you used it even once?
- dukeofdoom 4y agoI haven't used GPT-3, but I did try out a site that was based on GPT2. I believe it was called "talk to transformer". But I never tried quarrying anything controversial. However, I bet this a concern and certain queries will be filtered or "corrected" to be more politically correct. To give you an example, a few days ago I made a comment one Alex Jones, and wanted to google him. The second link returned on him was from ADL. No way that's an organic result. So just curious, if you have access to GTP-3 what does it return on Alex Jones, or other queries like who runs the banks, or who owns the media, and so on.
- deleted 4y ago[deleted]
- CrispinS 4y ago> The second link returned on him was from ADL. No way that's an organic result. It might be, actually. I understand why you'd think that, but look at the results for other search engines. Kagi: ADL in 2nd place Bing: ADL in 3rd place Yandex: ADL not on the first page, but SPLC[1] is the the 6th result [1]: https://www.splcenter.org/fighting-hate/extremist-files/individual/alex-jones https://www.splcenter.org/fighting-hate/extremist-files/indi...
- dukeofdoom 4y agoThis logic kind of fails quickly. I bet you wouldn't use it to show that Tiananmen Square did not happen, by showing all Chinese Search Engine are in apparent agreement on it not happening.
- alar44 4y agoGimme gimme. I want all your research and man hours for free. Gimme gimme. They are a for profit company and don't need to release anything. It's not that hard to understand.
- ninjin 4y agoTrue; they are free to do as they see fit. But how about not leeching on the word “open” in that case? DeepMind is essentially the NSA (or Apple), OpenAI is paid-for cloud services with paper-based marketing, and FAIR may be the best of the bunch, but it still annoys the hell out of me that they push code with non-commercial clauses as their current default (these are legally complicated in a university context) and now a model that they label “open” despite not honouring the accepted meaning of the word. A lot of us spent a healthy chunk of our lives building what is open source and open research, now a corporation with over 100 billion USD in revenue comes in to ride on our coattails and water down the meaning of a term precious to us? How about you spend the time and money to build your own terminology? “Available”, perhaps?
- ALittleLight 4y agoSure, but I'm an individual and free to say what I do and don't like. Why is that hard to understand?
- acchow 4y agoI’m thankful they’re offering anything at all openly. Is it such a big deal a gigantic download is hidden behind a request form?
- HWR_14 4y ago> Why do I have to request anything? I'm guessing it could be one or a mix of these: They want to build a database of people interested in this and vetted by some other organization as worth hiring. Just more people to feed to their recruiters. To see the output of the work. While academics will credit their data sources, seeing "XXX from YYY" requested, and then later "YYY releases product that could be based on the model" is probably pretty valuable vs wondering which ML it was based on. A veneer of responsible use, maybe required by their privacy policy or just to avoid backlash about "giving people's data away".
- lumost 4y agoA 175 billion parameter model might be a couple hundred gigs on disk. The file is probably just too big for GitHub/other standard FB services.
- fxtentacle 4y agoI'd guess they want to limit traffic. Once Huggingface links to you, your bandwidth bill 100x-es.
- saynay 4y agoIf it is like many other models, part of the reason would just be to reduce their bandwidth costs. The models can be huge, and they want to limit those who just want to download it on a whim so they don't rack up $10k+ is bandwidth charges, as has happened to many others who hosted big models out on S3 or something.
- levesque 4y agoIf only there was a way to distribute large files in a peer-to-peer manner, thus reducing the load on facebook's servers to effectively nothing. That would likely result in a torrent of bits being shared without any issues!
- JackC 4y ago> Why do I have to request anything? And I'm not an academic or researcher, so will they accept my random request? Just a guess: you will have to contractually agree to some things in order to get the model; at a minimum, agree not to redistribute it, but probably also agree not to use it commercially. That means whatever commercial advantage there is to having a model this size isn't affected by this offer, which makes it lower stakes for Facebook to offer. And then the point of "academics and researchers" is to be a proxy for "people we trust to keep their promise because they have a clear usecase for non-commercial access to the model and a reputation to protect." They can also sue after the fact, but they'd rather not have to. Not saying any of this is good or bad, just an educated guess about why it works the way it does.
- JoeyBananas 4y agoA 175B parameter language model is going to be huge. You probably don't want the biggest model just for messing around.
- gwern 4y agoI expect they will release the models fully, perhaps even under nonrestrictive licenses. Most researchers aren't too happy about those sort of restrictions, and would know that it vitiates a lot of the value of OPT. They look like they are doing the same sort of thing OA did with GPT-2: a staggered release. (This also has the benefit of not needing all the legal & PR approvals done upfront all at once; and there can be a lot of paperwork there.)
- schleck8 4y agoRepo down?
- scrollbar 4y agoDoes this make Meta AI more “open” than OpenAI? Oh, the irony.
- axg11 4y agoThey always have been. Meta has made a number of open contributions for the ML/AI community, one of which is PyTorch.
- px43 4y agoNot surprising at all since OpenAI is basically run by Microsoft now.
- mtlmtlmtlmtl 4y agoDon't worry, I'm sure they have some nefarious plans down the road. They're just being "open" to corner the market first.
- nomilk 4y ago> to corner the market first Is Meta's model going to be open source or paid?
- deleted 4y ago[deleted]
- sanxiyn 4y agoThe linked paper makes it clear it will be released under a non-commercial license. You will download it gratis (so it won't be paid), but it won't be open source.
- mtlmtlmtlmtl 4y agoSo they make a more available alternative, but they maintain control over it, and in turn gain control over the people and companies using it. Similar to what Microsoft did by bundling Windows with PCs[1]. I already have a multitude of ideas on potential nefarious plans based on this, but I'll keep them to myself. [1]: Sure they got a licence payment, but since it was built into the price and non-optional, it was effectively equivalent to free from the customer POV. It effectively became a tax. I have to admit, Gates might not be a genius programmer but he sure knows how to design dark patterns :)
- aleks5678 4y agoThanks Meta AI
- causality0 4y agoOut of curiosity, what's the file size on that?
- learndeeply 4y agoDepends which model, but assuming the largest: 175B * 16 bits = 350GB. Half of that if it's quantized to 8 bits. Good luck finding a GPU that can fit that in memory.
- faebi 4y agoDoes the model need to be in memory in order to run it with current tooling?
- PeterisP 4y agoTo run it at a reasonable speed, yes. Computing a single word requires all of the parameters; if you don't have them in memory you'd have to re-transfer all those gigabytes to the GPU for each full pass to get some output, which is a severe performance hit as you can't fully use your compute power because the bandwidth is likely to be the bottleneck - running inference for just a single example will take many seconds just because of the bandwidth limitations. GPT-3 paper itself just mentions that they're using a cluster V100 GPUs with presumably 32GB RAM each, but does not go into detail of the structure. IMHO you'd want to use a chain of GPUs each having part of the parameters and just transfering the (much, much smaller) processed data to the next GPU instead of having a single GPU reload the full parameter set for each part of the model; and a proper NVLink cluster can get an order of magnitude faster interconnect than the PCIe link between GPU and your main memory. So this is not going to be a model that's usable on cheap hardware. It's effectively open to organizations who can afford to plop a $100k compute cluster for their $x00k/yr engineers to work with.
- thrtythreeforty 4y agoExactly! This is called "model parallelism" - each layer of the graph is spread across multiple compute devices. Large clusters like the V100s or the forthcoming trn1 instances (disclosure, I work on this team) need _stupid_ amounts of inter-device bandwidth, particularly for training.
- crazypython 4y agoBigScience (a coalition including Huggingface) is training and releasing a 175B language model and finishes in 2 month.
- p1esk 4y agoWe are also releasing our logbook detailing the infrastructure challenges we faced Where’s the logbook?
- moyapchen 4y agohttps://twitter.com/stephenroller/status/1521302841276645376 https://twitter.com/stephenroller/status/1521302841276645376? Have patience it’s coming. :)
- moyapchen 4y agoAnd it’s live! https://github.com/facebookresearch/metaseq https://github.com/facebookresearch/metaseq Logbook links in specific: https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/README.md https://github.com/facebookresearch/metaseq/blob/main/projec...
- qgin 4y agoIf we’re already at the level of truly dangerous ml models… I don’t have a lot of hope for how the next decades are going to play out.
- thorum 4y agoA quick summary of the Limitations section: - "OPT-175B does not work well with declarative instructions or point-blank interrogatives." - "OPT-175B also tends to be repetitive and can easily get stuck in a loop. While sampling can reduce the incidence rate of repetitive behavior (Holtzman et al., 2020), we anecdotally found it did not eliminate it entirely when only one generation is sampled." - "We also find OPT-175B has a high propensity to generate toxic language and reinforce harmful stereotypes, even when provided with a relatively innocuous prompt (Gehman et al., 2020), and adversarial prompts are trivial to find." - "In summary, we still believe this technology is premature for commercial deployment." With regard to stereotypes: - "When compared with Davinci in Table 4, OPT175B appears to exhibit more stereotypical biases in almost all categories except for religion. Again, this is likely due to differences in training data; Nangia et al. (2020) showed that Pushshift.io Reddit corpus has a higher incidence rate for stereotypes and discriminatory text than other corpora (e.g. Wikipedia)." - When testing with the RealToxicityPrompts data set, "OPT-175B has a higher toxicity rate than either PaLM or Davinci"
- speed_spread 4y agoReminds me a lot of "Do not taunt Happy Fun Ball".
- ad_hominem 4y ago> Pushshift.io Reddit corpus Pushshift is a single person with some very strong political opinions who has specifically used his datasets to attack political opponents. Frankly I wouldn't trust his data to be untainted. These models really need to be trained on more official data sources, or at least something with some type of multi-party oversight rather than data that effectively fell off the back of a truck. edit: That's not even to mention I believe it's flat-out illegal for him to collect and redistribute this data as Reddit users did not agree to any terms of use with him. Just look at the disastrous mess of his half-baked "opt-out" thing that flagrantly violates GDPR: https://www.reddit.com/r/pushshift/comments/pat409/online_removal_request_form_for_removal_requests/ https://www.reddit.com/r/pushshift/comments/pat409/online_re...
- 4y ago
- etaioinshrdlu 4y agoWhat type of hardware would you need to run it?
- robbedpeter 4y agoA cluster of many $8000+ gpus. You're looking at around 350GB of vram, so 30 12gb gpus - a 3090 will cost around $1800, so $54k on the gpus, probably another $15k in power, cooling, and infrastructure, $5k in network, and probably another $20k in other costs to bootstrap it. Or wait 10 years, if gpu capacity scales with Moore's law, consumer hardware should be able to run a ~400GB model locally.
- adamsmith143 4y agoAlmost no one does this on prem. What would this cost on AWS?
- cardine 4y agoThis is not true. On prem is extremely common for things like this because after ~6 months you'll have paid more in cloud costs than it would have cost to purchase the GPUs. And you don't need to purchase new GPUs every 6 months. AWS would cost $50-100k/mo for something comparable.
- coolspot 4y ago> so 30 12gb gpus - a 3090 will cost around $1800 3090 has 24Gb, thus 15 GPUs X $1800 = $27,000 in GPUs
- etaioinshrdlu 4y agoCan 3090 GPUs share their memory with one another to fit such a large model? Or is the enterprise grade hardware required?
- coolspot 4y ago
- LeicaLatte 4y agoAs someone who finds openai patronizing, this is welcome.
- bestcoder69 4y agoI love text-davinci-002, but they need competition, badly. Their ToS is preventing me from releasing the world's greatest chatbot :P https://old.reddit.com/r/GPT3/comments/ubm0hm/my_customizable_gpt3_chatbot_as_karl_marx_with_a/ https://old.reddit.com/r/GPT3/comments/ubm0hm/my_customizabl...
- lumost 4y agoI often wonder if OpenAIs decision not to open gpt-3 was because it was to expensive to train relative to its real value. They’ve hidden the model behind an api where they can filter out most of the dumb behaviors, while everyone believes they are working on something entirely different.
- teaearlgraycold 4y ago> They’ve hidden the model behind an api where they can filter out most of the dumb behaviors What do you mean by this?
- codebolt 4y agoThings like cobbling on a bunch of heuristic rule-based behaviours that wouldn't look good in the public repo of a supposed quasi-AGI system?
- lumost 4y agoThere is some evidence that the OpenAI GPT-3 APIs have a human in the loop for bad examples. They may also have a number of filters to exclude certain words/patterns/other rules. The challenge with such rule and human in the loop systems is that the long-tail of these problems is huge, and fat. Meaning that you generally can't make a product which doesn't have full generalization. That it took ~1.5 years to open the GPT-3 API inclines me to think that they've run into similar problems. We're also not seeing the long pitched swarm of GPT enabled content despite the API being open for ~10 months.
- teaearlgraycold 4y agoThere’s no way they have a human in the loop. The model spits out tokens one at a time. You can see that with the stream flag set to true. The latency doesn’t allow for human intervention. They do have API parameters for tweaking repetitiveness. That might be what you’re talking about - but it’s fair to call the model and an external repetition filter part of the same product. As for word filters - no. If they did they’d not be sending back explicit content. But they do. If you have a gpt-3 product you’re obligated to run each result through their content filter to filter out anything nsfw. We don’t see a ton of gpt-3 enabled content because writing good gpt-3 prompts is hard. You’re trying to learn how this black box works with almost no examples to go off of. I worked for a gpt-3 startup and we put someone on prompt writing full time to get the most out of it. Most startups wouldn’t think to do that and won’t want to.
- mikolajw 4y agoThe big one, OPT-175B, isn't an open model. The word "open" in technology means that everyone has equal access (viz. "open source software" and "open source hardware"). The article says that research access will be provided upon request for "academic researchers; those affiliated with organizations in government, civil society, and academia; and those in industry research laboratories.". Don't assume any good intent from Facebook. This is obviously the same strategy large proprietary software companies have been using for a long time to reinforce their monopolies/oligopolies. They want to embed themselves in the so-called "public sector" (academia and state institutions), so that they get free advertising for taxpayer money. Ordinary people like most of us here won't be able to use it despite paying taxes. Some primary mechanisms of this advertising method: 1. Schools and universities frequently use the discounted or gratis access they have to give courses for students, often causing students to be only specialized in the monopolist's proprietary software/services. 2. State institutions will require applicants to be well-versed in monopolist's proprietary software/services because they are using it. 3. Appearance of academic papers that reference this software/services will attract more people to use them. Some examples of companies utilizing this strategy: Microsoft - Gives Microsoft Office 365 access for "free" to schools and universities. Mathworks - Gives discounts to schools and universities. Autodesk (CAD software) - Gives gratis limited-time "student" (noncommercial) licenses. Altium (EDA software) - Gives gratis limited-time licenses to university students. Cadence (EDA software) - Gives a discount for its EDA software to universities. EDIT: Previously my first sentence stated that the models aren't open - in fact, only OPT-175B is not (but the other ones are much smaller).
- mwint 4y agoAdd Mathematica to that list, too. Pretty cool to play with and I would have bought a license if I had a good excuse to; the tactic works.
- jrockway 4y agoMathematica has been on my mind since high school because we got it for free. I went through the free trial process recently and tried a couple of things I have been too lazy to manually code up (some video analysis). It was too slow to be useful. My notebooks that were analyzing videos just locked up while processing was going on, and Mathematica bogged down too much to even save the notebook with its "I'm crashing, try and save stuff" mode. I ultimately found it a waste of time for general purpose programming; the library functions as documented were much better than library functions I could get for a free language, but they just wouldn't run and keep the "respond to the UI" thread alive. So basically all their advertising money ended up being wasted because they can't fork off ffmpeg or whatever. Still very good at symbolic calculus and things like that, though.
- sjg007 4y agoHow about smaller more performant models? There’s so much redundancy in language that it should be possible.
- langsoul-com 4y agoI hope someone released a DALLE model. That seems far more interesting to play with.
- teaearlgraycold 4y agoIt'll happen eventually. And when it does, and if it's good enough, the world will be a different place afterwards.
- deleted 4y ago[deleted]
- anubhav200 4y agoDownload link?
- d--b 4y agoRemember when OpenAi wrote this? > Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT-2 along with sampling code. We are not releasing the dataset, training code, or GPT-2 model weights Well I guess Meta doesn’t care. https://openai.com/blog/better-language-models/ https://openai.com/blog/better-language-models/
- sigmoid10 4y agoEver since OpenAI transitioned away from the non-profit model, I'd take these statements with a grain if salt. Yes, there may also be some truth in that opinion, but don't underestimate monetary interests when someone has an easy ~12 month industry lead. Meta's existence and financial wellbeing on the other hand doesn't depend on this stuff, so they have less incentive to keep things proprietary. It seems ironic and almost bit sad that the new commercial circumstances have basically reversed these companies' original roles in AI research.
- SteveDR 4y agoI feel the same way. It does seem odd, though, that Meta would release this despite the precedent set by OpenAI with statements like this. What does Meta gain by releasing this for download?
- wdroz 4y agoI hate the nanny point of view of OpenAI. IMO trashing Meta because theirs models may be misused isn't fair. I think that hackers should advocate to have the freedom to toy/work with these models.
- jdrc 4y agohint: openAI didn't care either
- learndeeply 4y agoOpenAI released their large GPT-2 models weights a couple months after making that post: https://openai.com/blog/gpt-2-1-5b-release/ https://openai.com/blog/gpt-2-1-5b-release/
- deleted 4y ago[deleted]
- Pragati_08 4y agosome of the hardware Meta is working on to deliver it, https://www.theverge.com/2022/5/2/23053888/meta-virtual-reality-headset-cambria-quest-vr-mr https://www.theverge.com/2022/5/2/23053888/meta-virtual-real...
- jfmc 4y agoWe need robopsychologists.
- WithinReason 4y agoDr. Susan Calvin?
- f311a 4y agoJust curious, will I be able to use it using my Nvidia card with 10GB of memory? Does it require multiple graphic cards?
- woodson 4y agoAs the model weights (even quantized) would be several hundred GBs, it’s unlikely, unless special inference code is written that loads and processes only a small subset of weights and calculations at a time. But running it that way would be painfully slow.
- lostmsu 4y agoThe code is already there: DeepSpeed
- robbedpeter 4y agoThe smaller models, yes. I'd bet dollars to donuts that gpt-neo and EleutherAI models outperform most, if not all, of Facebook's. Check out huggingface, you'll be able to run a 2.7b model or smaller. https://huggingface.co/EleutherAI/gpt-neo-2.7B/tree/main https://huggingface.co/EleutherAI/gpt-neo-2.7B/tree/main
- imtemplain 4y ago
- ctreseler123 4y agoannounced because GPT4 makes this so very obsolete.
- einpoklum 4y agoI don't want to be a Luddite, but every time one of these FAANG companies makes advances in this domain my mind immediately goes to how they will use it to better spy on people, for commercial and government interests.
- urthor 4y agoIs the model of using an asterisk after first author's names to signal equal contribution common? Don't read many papers, but that's a new one.
- axg11 4y agoVery common.
- lol1lol 4y agoI appreciate that they are releasing their log book detailing the challenges faced.
- chrisMyzel 4y agoDoes anyone else think closed AI is turning into it's most weirdest forms and becoming a trend?
- israrkhan 4y agoI am afraid NLP is becoming a game of scale. Large scale models improve the quality but makes it prohibitively expensive to train, and even host such models.