9 ms·
Consortium launched to build the largest open LLM
- deleted 3y ago[deleted]
- refulgentis 3y agoMeh. PR overreach hoping for a Euro vanity project. University of Turku isnt a "hotbed" or "powerhouse" of LLM experts
- deleted 3y ago[deleted]
- spookie 3y agoIt seems quite a well positioned university. This dismissal seems unwaranted.
- jks 3y agoThe TurkuNLP team does have the best previous-generation language model for Finnish (FinBERT). The current version of ChatGPT speaks passable Finnish, although definitely not at the level it speaks English. None of the open-source models I've tried come close to ChatGPT performance. If they can gather a good multilingual dataset including the smaller European languages, and burn enough money on the compute, they could create a useful model for some specifically European use cases.
- olalonde 3y agoI recall a post on HN a while back about the EU planning to invest billions in AI. I can't find the HN link but found this source: "annual investment volume of €20 billion over the course of the digital decade"[0]. I wonder what happened with that money. It seems like using it to train large scale models would be a relatively effective use of that money. [0] https://digital-strategy.ec.europa.eu/en/policies/european-approach-artificial-intelligence https://digital-strategy.ec.europa.eu/en/policies/european-a...
- outside1234 3y ago[flagged]
- dopidopHN 3y agoOr train phd that ultimately decide to get a 300k offer over a 40k one. Look around and see who has a German, Spanish, French, Romanian accents..
- refulgentis 3y agoIt's looking for projects like this, so you end up with silly things like the University of Turku announcing it's powerhouse of LLM expertise (?) will train the "largest open model" in a consortium with (??). The only concrete factual detail is they'll train on euro-focused data
- avereveard 3y agoYeah, everyone claims sota, but the funniest ones are the ones claiming it before training even starts Also, the focus on "trustworthy" seems problematic. Beyond the obvious who is defining what can be trusted, trustworthy models are significantly harder to work with.
- deleted 3y ago[deleted]
- paxys 3y agoWhat is the bottleneck in building the "largest LLM" today for any interested party? AI expertise? Training data? GPUs?
- xeromal 3y agoMoney?
- olalonde 3y agoGPUs and electricity. GPT-4 reportedly cost 100M$ to train.
- pixl97 3y agoI'd say GPU and people trained in knowing what to do while training. There is a lot of money in the world, as we see Microsoft and other companies handing it out. Power even, if you build in the right place can be cheap, see bitcoin mining.
- patapong 3y agoVery much agree... I think people underestimate how finicky these models are to train. So much so that one of the big announcement around GPT-4 was the fact that OpenAI found a way to make it smoother and more predictable. As evidence by this 114 page log of how engineers resolved problems that came up during the training of OPT-175, which lasted for several months. https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf https://github.com/facebookresearch/metaseq/blob/main/projec...
- mark_l_watson 3y agoGood for them, Europe should compete against US, China, etc. It must be nerve-wracking being in charge of huge multiple month training cycles. In the 1980s, we built our own hardware to speed up back propagation learning and recall and runs could still take a day or two. It was always so disappointing when a long training run went bad. I can imagine how tense the responsibility must be to manage these huge training cycles.
- Jeff_Brown 3y agoAre there no incremental ways to verify that it's working?
- olalonde 3y agoYes, you can. After every training batch, the network's loss function is updated and you can track it over time to verify that it's going down. Sometimes, the loss can get stuck on a plateau, or even go up, and it's not clear whether it's worthwhile to keep training or not.
- brucethemoose2 3y agoTraditionally, training runs can "explode" and fail, but there are methods to incrementally back them up and resume when that happens, see https://www.mosaicml.com/blog/mpt-7b https://www.mosaicml.com/blog/mpt-7b
- mark_l_watson 3y agoOf course, you can measure loss on a dev set, etc. However, LLMs are so complex that until they are done training and many people evaluate them, there will be misses.
- senectus1 3y agoWhat political lean will a LLM have... we've had some really bad politics come out of Europe throughout history. Is this the beginning of a LLLM? (Lenin LLM) /s
- 3y ago
- Ckirby 3y agoAnd then they added their guardrails and made sure only positive, happy answers would be give
- api 3y agoWill it actually be open? As in weights available?
- quickthrower2 3y agoLargeness of the LLM shouldn’t be the goal.
- lappa 3y agoMore data, more parameters, more compute all result in a better model per "Scaling Laws for Neural Language Models" https://browse.arxiv.org/pdf/2001.08361v1.pdf https://browse.arxiv.org/pdf/2001.08361v1.pdf Largeness is a valid goal.
- quickthrower2 3y agoAlso: costs more for inference, uses more energy, less practical for running locally, fewer use cases as a result. Especially for an open model. Being on Github / HuggingFace but needing to be on a AWS or Nvidia wait list to get the resources to run it is not great. In an unlimited energy and chip world I would agree just make em bigger. I guess going bigger has a greater chance of success in being SOTA than looking at architectures. So I get people don’t want to gamble.
- huac 3y agorebuttal: compute optimality matters https://arxiv.org/pdf/2203.15556.pdf https://arxiv.org/pdf/2203.15556.pdf
- kumarvvr 3y agoWatched a debate on major media channels where people were arguing that pictures of killed children were AI pictures (and the opposite side refusing it) What kind of a world we have devolved into, where technological progress makes us less human, on average. I am saddened by the changes brought about by AI, be it generative art or LLMs, where, because of it, on a global scale, no one knows what truth is, even in pictures and videos. And we have the perfect tinderbox of siloed wells on social media, with tools to generate authentic looking content that can be disseminated at essentially 0 cost, automated at 0 effort and be used to indoctrinate people at 0 consequences from governments. It does not matter if people are caught after the fact, the mind that has been indoctrinated, cannot be reverted with the same ease. And we are generating a bunch of children now, who will be adults later, who will essentially reject reality, reject the need for empathy because what they see may or may not be the truth. It used to be that 10 years before, if one saw the pictures of carnage or destruction or dead children, a spark of empathy may be generated, enough to ignite change, no matter how small. In the near future, that may not be the case because that photo or video or content has the probability of being AI generated and that spark of empathy will die without a chance to ignite change. And now we have billions of dollars being invested by the very governments that will feel the effects of these short sighted decisions to unleash technology seldom understood and have no clue as to its long term effects.
- WanderPanda 3y agoI'm wondering if the interconnect between GPUs in this HPC which seems like a more traditional supercomputer good enough for LLM training?
- bratao 3y agoIn my opinion They won't outperform the best LLMs for one problem: Quality of data. A hidden secret that everyone knows is that good LLMs use copyrighted data like books-3. OpenAI itself cites a dataset from books and does not go into detail in its first papers.
- giardini 3y agoSpeaking of "secrets": We didn't create an open-source project for the atomic bomb. Quite the contrary: it was one of the best-kept secrets for years. Sooooo... if LLMs are as revolutionary as some claim, shouldn't there be significant effort to keep new advances and developments secret and lead competitors astray? That is, if I knew something that would provide a 10- or 100-fold improvement in LLMs, my tendency would be to profit/benefit from that rather than publish it. And to mislead competitors? Has LLM technology development gone "underground" yet? And what is the role of the 3-letter agencies and the government in this radical development? Or is LLM technology all unicorns with rainbow farts where the entire world benefits and the creators get only thanks? Somehow I don't believe so.
- giardini 3y agoIs bigger always better? Surely "bigger is better" is at odds with "garbage in, garbage out". But I am not an LLM modeler. IOW can LLMs be improved by improving the quality and/or even *limiting the size of their "corpus"? Could we instead use, as corpus, a select body of knowledge primarily based on Western sources such as those texts a student might read as part of an excellent liberal arts education? Further, would it not be possible to limit the corpus text (and consequently later, queries) to grammatically correct language? That is, must we include all the "errors" of common speech in a corpus, such as occur especially in fiction? Indeed, could we not exclude fictional works altogether, esp. if the model is to be used for scientific work? I'd be curious to see an LLM trained on "The Great Books" plus a college course plan of textbooks. But again, I am not an LLM modeler. Anyone experimenting with such "Small Language Models(SLM)"?
- ronsor 3y agoThe TinyStories paper[0] explores some of the potential of "small language models." I think you don't need as much data if you have really high quality data (really high quality); however, having a lot of data seems to be able to compensate for widely varying quality. [0] https://ar5iv.labs.arxiv.org/html/2305.07759 https://ar5iv.labs.arxiv.org/html/2305.07759
- brucethemoose2 3y agoI think data quality, not just largeness, should be a goal. A huge and excellent dataset is something only a massive organization can afford.
- peddling-brink 3y agoSomeone needs to gamify cleaning up data.
- huac 3y ago15 million GPU hours is interesting, llama 2 had about 1.7 million hours for the 70B (3.3M for all). Their GPU hours are on LUMI which is AMD MI250 GPU, which Mosaic reports is about 75% of the A100 (https://www.mosaicml.com/blog/amd-mi250 https://www.mosaicml.com/blog/amd-mi250) but AFAIK hasn't been tested at this kind of scale. so, validating if the MI250 in a cluster can effectively train a large scale LLM will be useful. not that it really matters to any other group, you won't get these in the cloud anytime soon. either way - say that they have about 10M A100 hours. if they use half their budget for the biggest model as FB did with Llama (rest on testing / smaller models as they scale), then they'll budget 5M A100 hours for the biggest model, or 3x llama-2's 1.7M A100 hours (2k A100 for 21 days). However, why wouldn't Llama-3 also be 3x bigger than Llama-2? FB's H100's are coming online; they apparently have 4k H100 in a cluster (https://www.nextplatform.com/2023/09/26/meta-platforms-is-determined-to-make-ethernet-work-for-ai/ https://www.nextplatform.com/2023/09/26/meta-platforms-is-de...), and each H100 is about 3x the speed of an A100 (https://www.mosaicml.com/blog/coreweave-nvidia-h100-part-1 https://www.mosaicml.com/blog/coreweave-nvidia-h100-part-1). So if they provided the same 'calendar time' of 21 days over a single cluster of 4096 H100, that would be about 6.2M A100 hours for a 70B model. there is an entirely separate question of, LLM's are functions of compute _and_ data, will they also find enough high quality _and_ kosher tokens in their low resource languages?
- vidarh 3y agoIf this gets enough of an official stamp of approval from the EU, then access to enough tokens in smaller languages may not be a big issue. For that matter, even for OpenAI it's likely largely down to how much they care to negotiate access. E.g. for Norwegian the Norwegian national library and state archives sits on at least a couple of magnitudes more of digitized material in Norwegian than OpenAI appears to have found Norwegian material for GPT3. How much exactly would depend on willingness to give access to still in-Copyright data, but it'd still be far more even if you only give access to what is currently on their website. And ChatGPT is well versed enough in Norwegian even without that to be able to convincingly use several regional dialects that are rarely used in writing (I've tested several) so it won't take all that much. But frankly being able to ask it about the contents of these archives would be amazing because they include a couple of centuries of newspapers and almost every book ever published in Norwegian. The level of digitization will differ, and Norway was early with that, but most countries have depositary requirements for printed materials, and countries with smaller languages like are often more anal about them...