7 ms·
Training AI models might not need enormous data centres
- jkuria 2y agohttps://archive.is/kRfd2 https://archive.is/kRfd2
- gnabgib 2y agoRelated: New Training Technique for Highly Efficient AI Methods (2 points, 5 hours ago) https://news.ycombinator.com/item?id=42690664 https://news.ycombinator.com/item?id=42690664 DiLoCo: Distributed Low-Communication Training of Language Models (46 points, 1 year ago, 14 comments) https://news.ycombinator.com/item?id=38549337 https://news.ycombinator.com/item?id=38549337
- metadat 2y agoThanks for the links, some interesting discussion there. The second article you linked indicates there will still be intense bandwidth requirements during training, shipping around gradient differentials. What has changed in the past year? Is this technique looking better, worse, or the same?
- gaogao 2y agoYeah, high bandwidth requirements still remaining. Over the past year, more research has looked from fully async to restrained cases that allow for geographically distributed compute. Async Local-SGD goes for a more standard training objective comparable with a lockstep training, https://arxiv.org/abs/2401.09135 https://arxiv.org/abs/2401.09135. imo technique is looking better.
- m3kw9 2y agoIt’s talking about training 10b parameters “capable models” with less compute using ew techniques, but top models will always need more
- kurthr 2y agoWow, yeah a 10B parameter model is pretty tiny and 300 3-GPU clusters for $18M is not really cheap. I guess enormous is in the eye of the beholder.
- gaogao 2y agoThis approach likely scales to top models. There was a thought for a while among researchers that scaling up Hogwild / LocalSGD to more parameters would be less convergent, but it seems to be that things actually converge a bit better with larger models, especially if a MoE split is used.
- gaogao 2y agoThe intuition is that smaller models are figuring out things like grammar where the whole model comes into play, but larger models, especially in the back half of training, have localized knowledge updates that can merge easier in the AllReduce
- openrisk 2y agoOpen source public models trained on kosher data are substantially derisking the AI hype. It makes a lot of sense to push this approach as far as it can get. Its similar to SETI at home etc. but potentially with far more impact.
- grumbelbart2 2y agoThey could, would and should. But: Training a state of the art LLM costs millions in GPU, electricity alone. There is no "open" organization at this point that can cover this. Current "open source public models" are shared by big players like Meta to undermine the competition. And they only publish their weights, not the training data, training protocols, training code; meaning it's not reproducible, and questionable if the training data is kosher.
- sebmellen 2y agoDoesn’t Deepseek somewhat counter this narrative?
- htrp 2y agoDon't they have something like 10k plus current gen GPUs?
- lz400 2y agoI understood SETI style meaning crowdsourced. Instead of mining bitcoin you mine LLMs. It's a nice idea I think. Not sure about technical details, bandwidth limitations, performance, etc.
- Mountain_Skies 2y agoSETI had a clear purpose that donors of computer resources could get behind. The LLM corps early on decided to drink the steering poison that will keep there from ever being a united community for making open LLMs. At best you'll get a fractured world of different projects, each with its own steering directives.
- whazor 2y agoYou could consider a LLM as a very lossy compression artifact. Where they took terabytes of input data, and ended up with model under the 100 gigabytes. It is quite remarkable what such a model can do, even fabricating new output that was not in the input data. However, in my naïvety, I wonder whether vastly simpler algorithms could be used to end up with similar results. Regular compression techniques work with speeds up to 700MB/s.
- briandear 2y agoIf they could get to a 5.2 Weissman compression score it would probably make a substantial difference.
- topspin 2y agoYou could consider the human mind to be a very lossy compression artifact.
- godelski 2y ago> However, in my naïvety, I wonder whether vastly simpler algorithms could be used to end up with similar results. Almost certainly. Distillation demonstrates this. The difficulty is training. It's harder to train a smaller network and harder to train with less data. But look at humans, they ingest far less data and certainly less diverse data. We are extremely computationally efficient. I guess you have to be when you run on meat
- purplethinking 2y ago> they ingest far less data True in terms of text, but not if you include video, audio, touch etc. Sure, one could argue that there is much less information content in video than their raw bytes, but even so, we spend many years building a world model as we play with tools, exist in the world and go to school. I don't deny humans are more efficient learners but people tend to forget this. Also, children are taught things in ascending order of difficulty, while with LLMs we just throw random pieces of text at it. There is sure to be a lot of progress in curriculum learning for AI models.
- 2y ago
- aimanbenbaha 2y agoThis bottleneck right here is why Open Source is presented with a golden plate opportunity to lead the training of cutting edge models. Federated learning breaks the barrier to entry and expands the ecosystem allowing more participants to share compute and/or datasets for small players to train models. DiLoCo introduced by Douillard minimizes communication overhead by averaging weight updates. What this article misses though is that despite this, each GPU in the distributed cluster still needs to have enough VRAM to load the entire copy of the model to complete the training process. That's where DisTrO comes in which even reduces further the inter-GPU communication using a decoupling technique (DeMo) that only shares the fast moving parts of the optimizer across the GPU cluster. >And what if the costs could drop further still? The dream for developers pursuing truly decentralised ai is to drop the need for purpose-built training chips entirely. Measured in teraflops, a count of how many operations a chip can do in a second, one of Nvidia’s most capable chips is roughly as powerful as 300 or so top-end iPhones. But there are a lot more iPhones in the world than gpus. What if they (and other consumer computers) could all be put to work, churning through training runs while their owners sleep?" This aligns with DisTrO techniques because, according to them it could also allow consumer devices like Desktop Gaming PCs to join the compute cluster and share workloads. Besides there's also an open-source implementation called exo that allows models to be split among idle local devices but it's only limited to inference. Again might still be relevant since in the article it mentions that DiLoCo was able to make the model respond better when faced with instruction prompts or reasoning questions never encountered during pre-training. And Arthur seems to think test-time training will make his approach become the norm. sources: DisTrO: https://github.com/NousResearch/DisTrO https://github.com/NousResearch/DisTrO DeMo: https://arxiv.org/pdf/2411.19870 https://arxiv.org/pdf/2411.19870 Exo: https://github.com/exo-explore/exo https://github.com/exo-explore/exo
- lambda-research 2y ago> What this article misses though is that despite this, each GPU in the distributed cluster still needs to have enough VRAM to load the entire copy of the model to complete the training process. That's not exactly accurate. In the data parallel side of techniques, the Distributed Data Parallel (DDP) approach does require a fully copy of the model on each GPU. However there's also Fully Sharded Data Parallel (FSDP) which does not. Similarly things like tensor parallelism (TP) split the model over GPUs, to the point where full layers are never in a single GPU anymore. Combining multiple of the above is how huge foundation models are trained. Meta used 4d parallelism (FSDP + TP and pipeline/context parallelism) to train llama 405b.
- FrustratedMonky 2y agoAre there not any lessons from Protein Folding that could be used here? There was a distributed Protein Folding project a couple decades ago. I remember there was even Protein folding apps that could run on game consoles when not playing games. But maybe Protein Folding code is more Parallelizable across machines, than AI models.
- neom 2y agoI think the problem is people are going to start playing, everyone is going to train in their own things, businesses are going to want to train different architectures for different business functions etc. I did my first real adventure with training last night, $3,200 and a lot of fun later (whooops) - the tooling has become very easy to use, and I presume will just get easier. If I want to train in even say 10ish gigs, wouldn't I want to use a DC, even with a powerful laptop or DiLoCo? Seems unlikely DiLoCo is enough? (edit: I may also not be accounting enough for using a pre-trained general model next to a fine tuned specialized model?)