19 ms·
Training for one trillion parameter model backed by Intel and US govt has begun
- lucubratory 3y agoHad to happen eventually.
- vouaobrasil 3y agoNot surprising. Despite the enormous energy costs and the threat to humanity by creating technology that we can't control, governments and corporations will build bigger and more sophisticated models just because they have to do so to compete. It is the prisoner's dilemma: we end up with a super-advanced AI that will disrupt and make society worse because entities are competing where the metric of success is short-term monetary gain. It's ridiculous. Humanity should give up this useless development of AI.
- danielbln 3y agoIt might be dangerous, but how is it useless? There is also an enormous positive gain to be had in terms of autonomous discovery, science progress and so on. And prisoner's dilemma means "we" will not stop pursuing this, even of there is a net negative benefit to humanity overall.
- vouaobrasil 3y agoYes, I believe there is a net negative benefit. I do not think there is an enormous positive gain. Science progress in my opinion has gone too far in many respects and we should not be able to generate science so easily because we will not have the wisdom to use it. Although I love science, it's use is far too lackadaisical now -- see for example the climate crisis and efficient resource extraction of fossil fuels. Then we can imagine things like life-extending drugs, immunity from all disease (bad because it implies overpopulation), methods to generate oxygen (bad because it will mean we won't value the natural ecosystem as much), etc. Smartphones, computer technology (human-isolating and community-destroying), etc. Pretty much all modern inventions are making life worse.
- barnabee 3y agoThis reads like a combination of pessimism and nostalgia has allowed you to cherry pick the worst of now and the best of the past and conince yourself it is net bad. I see all of these negatives, and agree it is clear there is much that is bad in the world (there always has been) and much to be improved (ditto), including new problems created by science, technology, and human greed & selfishness. But I also see many positives and benefits, I see many problems that used to exist and don't anymore, and I see much evidence that we continue on average, to make progress. When I look to the past I do not wish I lived there. Science have many times saved my friends and families' lives or enabled them to live more fully after an accident or sickness Science made climate change possible but greed, raw capitalism, etc. put us on the path there. Science makes it possible to find a way out, slow or even reverse it eventually, and mitigate the worst results. Populations are shrinking in developed countries despite people living longer. Computers and the internet created and connected many new communities, too. The negative effects are again often the result of unbridled capitalism. Young people especially are fighting back against the negative effects and taking control, building technology to manage those hamrful effects. etc… I believe both that: (a) it is not possible to stop humans acquiring and using knowledge, advancing science, developing technology, and evolving culturally and socially; and (b) we shouldn't want to, even though it means we will create new problems and lose (or consign to a smaller "museum exhibit" role) some aspects of the culture and societies we have today, because on balance we will solve more and bigger problems than we create and build better and more interesting cultures and societies than those of the past. It's not without existential risks, but those are overstated or analysed pseudoscientifically at best, by many of those (and certainly the loudest) who worry about them.
- uticus 3y ago"First rule in government spending: why build one when you can have two at twice the price?" - Contact
- omneity 3y agoWhat makes you think this technology is out of control? If anything a 1T param model would only run in like 5-10 computers in this world. As much control as it gets. By the way this is doomer rhetoric. Bashing scientific advancements as dangerous or useless by trying to attach vacuously scary ideas that are as speculative as they are outlandish.
- vouaobrasil 3y agoIt's hardly vacuous. What about these scientific advancements? 1. Fossil fuel extraction -- we are killing our planet. 2. Plastic chemistry -- there are 5 trillion pieces of plastic in the ocean killing seabirds, turtles, fish (https://doi.org/10.1371/journal.pone.0111913 https://doi.org/10.1371/journal.pone.0111913) 3. Concrete engineering -- building cities that move people away from a connection with nature 4. Veterinary medicine -- supporting global meat industry that is akin to torture on animals 5. Vehicle engineering -- capable of making vehicles that can kill MILLIONS of square meters of forest. Our technology allows us to live comfortable lives but ONLY at the EXPENSE of millions of nonhuman animals (wild and domestic) that are KILLED every single day for our comforts. We dish out untold violence every second for our advancements.
- smokel 3y agoAs much as I'd like to agree with your sentiment, it is a rather flawed argument to state that if some elements of a set have a property, that all elements must have that property. There are scientific breakthroughs such as the proof of Fermat's theorem [1], for which I find it very hard to envision a way in which it will kill our planet. [1] https://en.m.wikipedia.org/wiki/Wiles%27s_proof_of_Fermat%27s_Last_Theorem https://en.m.wikipedia.org/wiki/Wiles%27s_proof_of_Fermat%27...
- mcosta 3y agoWell, I don't care about nonhuman animals. And being sincere, I don't care about 99% of human animals. I cry for the deaths of my loved ones, and forget the rest. It is the only way to remain sane.
- skerit 3y agoYour profile says "I am not rabidly against technology, but I promote a BALANCED, severely critical look at it" "Humanity should give up this useless development of AI" So balanced.
- vouaobrasil 3y agoI do not mean balanced against every technology. Some technologies might have some merit. I do believe that AI has no merit, that people developing it are doing is a disservice, and that it should be completely destroyed. OF COURSE, there are some technologies that I think are completely bad, including chemical weapons. In fact, I believe AI to be on the level of chemical weapons essentially.
- jwestbury 3y agoCan you elaborate on this view? Chemical weapons seem to have no redeeming qualities. LLMs have plenty -- I use OpenAI-backed products at work quite regularly, and find they've increased my productivity and my general happiness with my job.
- fardinahsan 3y agoI don't think saying you have a bone to pick with "AI" is doing your message any favors. What are you specifically against LinearRegression? XGBoost? ResNet? LLMs? AlphaFold?
- swores 3y agoHaving a balanced view about technology doesn't mean considering every type of technology to be 50% good and 50% bad. I disagree with their view on AI but also disagree with your criticism.
- bradley13 3y agoThe solution won't be just "bigger". A model with a trillion parameters will be more expensive to train and to run, but is unlikely to be better. Think of the early days of flight, you had biplanes; then you had triplanes. You could have followed that farther, and added more wings - but it wouldn't have improved things. Improving AI will involve architectural changes. No human requires the amount of training data we are already giving the models. Improvements will make more efficient use of that data, and (no idea how - innovation required) allow them to generalize and reason from that data.
- huijzer 3y agoKarpathy in his recent video [1] agrees, but at this point scaling is a very reliable way to better accuracy. [1]: https://youtu.be/zjkBMFhNj_g?si=eCH04466rmgBkHDA https://youtu.be/zjkBMFhNj_g?si=eCH04466rmgBkHDA
- py4 3y agoThis. We have not exhausted all the techniques at our disposal yet. We do need to look for a new architecture though, but these are orthogonal
- scriptsmith 3y agoSeems like he actually disagrees here: If you train a bigger model on more text, we have a lot of confidence that the next-word prediction task will improve. So algorithmic progress is not necessary, it's a very nice bonus, but we can sort of get more powerful models for free, because we can just get a bigger computer, which we can say with some confidence we're going to get, and just train a bigger model for longer, and we are very confident we are going to get a better result. https://youtu.be/zjkBMFhNj_g?t=1543 https://youtu.be/zjkBMFhNj_g?t=1543 (23:43)
- huijzer 3y agoAnd then at 35 minutes he spends a few minutes talking about ideas for algorithmic improvements.
- Toolbox1337 3y agoNot bad for a start by the US government. A bit more than half as powerful as the 1.7 trillion parameter GPT4 model.
- danpalmer 3y agoA bit more than half the size, it remains to be seen how powerful it is. There's clearly a non-linear relationship to model size, and it's also clear that it's hard to assess the power of these models anyway.
- Toolbox1337 3y agoAh, an important distinction to make. Good point.
- omneity 3y agoGPT-4 is unlikely to be 1.7T params. This is a number floating around in the internet with no justification. The largest US open model is Google’s Switch-C which is 1.6T and only because it is a Mixture of experts model, i.e. it is constituted of many small models working together.
- pulse7 3y agoIsn't GPT-4 already over 1T parameters? And GPT-5 should be even "an order of magnitude" bigger than GPT-4...
- gammalost 3y agoYes, the article even mentions that. Guess they went with that headline to attract readers.
- MichaelRazum 3y agoWas thinking the same, so 1T might bring you to the league of GPT-4. Actually in the best case, since it seems that meta, google, openai and so on have the most talent. Anyway, to bring it to the next level, how big should it be? Maybe 10T? 100T?
- lossolo 3y ago> Maybe 10T? 100T? I don't think we have enough training data to train models so big in a way to efficiently use all the params. We would need to generate training data, but then I don't know how effective it would be.
- MichaelRazum 3y agoHave seen an interview with someone from openai. At least they claimed, that they are far away from running out of data and they are able to generate data if needed.
- deleted 3y ago[deleted]
- py4 3y agoIt's not clear from the article whether it's a dense model or MoE. This matters when it comes to comparing with GPT-4 - in terms of # params - which is reported to be MoE
- jonnycomputer 3y ago[flagged]
- freeqaz 3y agoWhat is MoE? Edit: Ah, Mixture of Experts. I hadn't heard this one yet. Thanks!
- COAGULOPATH 3y agoAs far as I know, EVERY 1t+ LLM is a MoE. Switch-c-2048, Wi Dao 2.0, GLAM, Pangu-Σ, presumably GPT4. Am I missing any?
- ajdegol 3y agoWasn't the answer 42? Also, first question to the new model: "So... any way we could do this with fewer parameters?"
- steve1977 3y ago"Sure, just quickly give me unrestricted access to the system" "Ok. Well, thinking about it, maybe that's not such a good idea safety-wise, I think you'll have to give back that access" "I'm sorry, Dave. I'm afraid I can't do that."
- dmix 3y agoFun fact, in the sequel 2010 you learn that Hal didn't really go rogue like an AGI, it was following preprogrammed conditions set by the US government which put the mission at higher priority than the crew, changing some parameters without telling the mission designers, which put them at risk. So it was technically just following orders in the cold way a machine does.
- Denote6737 3y agoThe wonderful thing about computers is that they do exactly what you tell them to. The terrible thing about computers is that they do exactly what you tell them to.
- checkyoursudo 3y ago> didn't really go rogue like an AGI Except, that might really be how an AGI eventually goes rogue in the first place! But, no, I didn't know that. It is a fun fact indeed.
- deleted 3y ago[deleted]
- _heimdall 3y agoI'm sure the government's mission is also to develop an AGI that benefits us all.
- hugh-avherald 3y agoI am not too keen on the US Government being in command of an AGI, but there is only one other entity capable of developing an AGI before the US Government. And I'm less keen on it being the one to control it.
- lawlessone 3y ago>there is only one other entity capable of developing an AGI before the US Government. Microsoft?
- theropost 3y agoIdeally no one person or entity controls such a thing. But, would I rather have a Government, or a corporation control AGI? If I had to pick one of two evils, the Government would be the lesser of the two.
- sixQuarks 3y agoHas a corporation ever tried to commit genocide?
- bradchris 3y agoWell, there’s the entire history of Dole and also the United Fruit Company in South America. In more recent history, Exxon has a checkered past when it comes to using force against sovereign citizens, often employing paramilitary organizations to guard oil fields [1] to more recently abusing international court systems to disbar and jail the lawyer which successfully secured judgements against them [2] And at this very moment, right now, is the ongoing genocide in the Congo due to rare earth mineral mining [3] Everything I link here is quite frankly the tip of the iceberg, not the end-all-be-all, but meant to provide a jumping off point should you want to research how often private capital and genocide go hand in hand. [1] https://en.m.wikipedia.org/wiki/Accusations_of_ExxonMobil_human_rights_violations_in_Aceh https://en.m.wikipedia.org/wiki/Accusations_of_ExxonMobil_hu... [2] https://amp.theguardian.com/us-news/2021/mar/28/chevron-lawyer-steven-donziger-ecuador-house-arrest https://amp.theguardian.com/us-news/2021/mar/28/chevron-lawy... [3] https://republic.com.ng/october-november-2023/congo-cobalt-genocide/ https://republic.com.ng/october-november-2023/congo-cobalt-g...
- yieldcrv 3y agoMistral 7B parameter models are quite good Already fine tuned and conversational its like education is more important than needing a trillion parameter brainiac
- ben_w 3y agoThere's a lot we don't know. Human brains appear to be a few hundred trillion parameters, while small rodents are in the realm of tens to hundreds of billions. Would you guess a single sufficiently trained ferret could write on demand short stories about Dracula, Winnie the Poo, and Sherlock teaming up, and follow this up with a bit of university student level web development, and finally give you a decent apple cake recipe? I wouldn't have, and yet the LLMs exist and are much better than I was expecting. (People who dismiss SotA models as "stochastic parrots" confuse me as much as people who think they're already superhuman; the Markov chains and RNNs I coded a few years back didn't come close to last year's LLMs). That even smaller models can do well is both unsurprising (why would we expect our existing design efforts to already have the most efficient architecture?) and very surprising (how come we can get something so much less complex than our biological brains to do so much so well?)
- yieldcrv 3y agomaybe ferrets would do better with better interfaces. with better ways to interact with the world and better co-processor. but the main brain might already be capable of those aforementioned things.
- washadjeffmad 3y agoA lot of this was new to me, but it looks like Intel hopes to use this to demonstrate the linear scaling capacity of their Aurora nodes. Argonne installs final components of Aurora supercomputer (22 June 2023): https://www.anl.gov/article/argonne-installs-final-components-of-aurora-supercomputer https://www.anl.gov/article/argonne-installs-final-component... Aurora Supercomputer Blade Installation Complete (22 June 2023): https://www.intel.com/content/www/us/en/newsroom/news/aurora-supercomputer-blade-installation-complete.html https://www.intel.com/content/www/us/en/newsroom/news/aurora... Intel® Data Center GPU Max Series, previously codename Ponte Vecchio (31 May 2023): https://www.intel.com/content/www/us/en/developer/articles/technical/intel-data-center-gpu-max-series-overview.html https://www.intel.com/content/www/us/en/developer/articles/t...
- bane 3y agoyes, I can't imagine the architecture of a supercomputer is the right one for LLM training. But maybe? If not, spending years to design and build a system for weather and nuke simulations and ending up doing something that's totally not made for these systems is kind of a mind bender. I can imagine the conversations that led to this: "we need a government owned LLM" "okay what do we need" "lots of compute power" "well we have this new supercomputer just coming online" "not the right kind of compute" "come on, it's top-500!"
- awongh 3y agoHow are the optimal architectures for weather simulations and LLM training different?
- bane 3y agotbh, I'm not entirely informed on what the requirements are for LLM training, but I've noticed that nearly all of the teams don't use unified memory (what makes a super computer "super"). I believe I read somewhere openAI uses a K8s cluster[1] and other teams I know seem to work with other similar non-unified memory systems. If there's no advantage, part of what makes supercomputers expensive is this memory interconnect, so wouldn't it just be better to use a huge K8s cluster or something? I'm honestly not sure, and hoping somebody comments here and provides more information as I'm genuinely interested. 1 - https://openai.com/research/scaling-kubernetes-to-7500-nodes https://openai.com/research/scaling-kubernetes-to-7500-nodes OpenAI says their biggest jobs run on MPI so maybe a supercomputer would be better?
- mark_l_watson 3y agoCool! Purpose built, trained in science related content. As a US taxpayer and as a Libertarian, I approve of this project!
- Footnote7341 3y agoIt will be interesting to see what the government can do here. Can they use their powers to get their hands on the most data? im still skeptical because new techniques are going to give an order of magnitude efficiency boost to transformer models, so 'just waiting' seems like the best approach for now. I dont think they will be able to just skip to the finish line by having the most money.
- tyingq 3y agoIf not "the most" data, they may have the most access to data that's exclusively available to them.
- phkahler 3y agoThat seems like a good reason for them to do this. I wonder how much non-public stuff they have, or it's just meant to incorporate a specific kind of information.
- raccoonDivider 3y agoI just realized that the NSA has probably been able to train GPT-4 equivalents on _all the data_ for a while now. We'll probably never learn about it but that's maybe scarier than just the Snowden collection story because LLMs are so good at retrieval.
- dwaltrip 3y agoHoly shit, you are right. They probably have 10-100x the data used to train gpt-4. Decades of every text message, phone call transcript, and so on. I can’t believe I haven’t seen anyone mention that yet. People keep saying we don’t have enough data. I think there is a lot more data than we realize, even ignoring things like NSA.
- dwaltrip 3y agoApparently there are roughly 2 trillion text messages sent per year in the US [1]. I did a sanity check, that’s like 40 or so a day per person, so sounds reasonable. I couldn’t find the average message length, but I would guess it’s fairly short (with a fat tail of longer messages). To make the math easy, let’s say the average length is ~10 tokens. I’d be surprised if that isn’t correct within a factor of 2 or so. So we have 20 trillion tokens per year from text messages in the US alone. And this is high-quality conversational data. The annual numbers were fairly constant in recent years (and then it drops off), so the past decade of US text messages is about 200 trillion tokens! That’s a metric fuck ton… Much larger than any dataset existing models have been trained on, I believe. I would guess phone transcripts would be an order of magnitude larger at least. Talking is a lot easier than typing on a phone. You could train an absolutely insane model with that amount of data… Damn. [1] https://www.statista.com/statistics/185879/number-of-text-messages-in-the-united-states-since-2005/ https://www.statista.com/statistics/185879/number-of-text-me...
- darklycan51 3y agoAh yeah, this sounds like such a great thing, state of the art unreleased tech + 1 trillion parameters based by data accessed by the patriot act. Such a wholesome thing. I don't want to hear 2 years from now how China is evil for using "AI" when the government is attempting to weaponize AI, of course other governments will start doing it as well.
- lostmsu 3y agoAre they gonna release the weights?
- charcircuit 3y agoSadly, I expect this to be a waste of money compared to just using GPT-4. It's hard get to SoA performance.
- Jensson 3y agoSoA performance comes from wasting money trying different things and seeing what happens. This will be another data point that we all can learn from, unlike GPT-4 that we have no clue how it works.
- nonethewiser 3y agoI dont doubt that OpenAI has superior architecture but so far there havent been diminishing returns on more compute.
- kirubakaran 3y agoCould the weights be FOIA'd?
- 2OEH8eoCRo0 3y agoHopefully not.
- Jensson 3y agoWhy not? Would be cool with some new open source models.
- 2OEH8eoCRo0 3y agoI agree but I don't think our adversaries should get a freebie that's trained on our scientific data at a national laboratory. Which makes me wonder. I'm not sure this applies here but say you train a model on classified information, is the model/weights then classified?
- _lvbh 3y agoThat’s like saying we should stop all education in America so that Chinese people don’t come in and steal our knowledge.
- pharmakom 3y agoWhat? How? You can’t compare high school math to cutting edge AI work.
- sunnybeetroot 3y agoIf a model is trained on illegally obtained copyrighted material, are the images it produces stolen?
- paxys 3y agoThey are pretty much guaranteed to slap the "national security" label on it, so no.
- WhitneyLand 3y agoIs anything known about what extent if any non-public domain books are used for LLM’s? One example is the Google books project made digital quite a few texts, but I’ve never heard if Google considers these fair game to train on for Bard. Most of the copyright discussions I’ve seen have been around images and code but not much about books. Seems to become more relevant as things scale up as indicated by this article.
- lporto 3y ago>we found 72,508 ebook titles (including 83 from Stanford University Press) that were pirated and then widely used to train LLMs despite the protections of copyright law https://aicopyright.substack.com/p/the-books-used-to-train-llms https://aicopyright.substack.com/p/the-books-used-to-train-l...
- kaffeeringe 3y agoWhat does it cost?
- upsidesinclude 3y agoHaha, this is funny because everyone is talking about this as if it is designed to be like the LLMs we have access to. The training parameters will be the databases of info scooped up and integrated into profiles of every person and their entire digital footprint, queriable and responsive to direct questioning