6 ms·
I've been contemplating a decentralized model training system for some time using volunteer machines that we all contribute. But, it is astronomically difficult
by palisade 4mo ago
I've been contemplating a decentralized model training system for some time using volunteer machines that we all contribute. But, it is astronomically difficult. The communication speeds are untenable.
And, there is the issue of data poisoning from untrusted nodes. I've almost cracked that last issue with a self-healing checkpointed rollback system that doesn't have to throw out anything that follows the corrupt datum.
But, I'm just one person with an idea and I don't have infinite funds to make this happen. This isn't a small project.
Maybe there would be interest in something like this, now that entire frontier labs are being banned from making further progress.
The total power of all GPUs on the planet dwarf their capabilities, if we had a way to harness them in a distributed way efficiently. We wouldn't be able to train a Fable as fast as them, but eventually having access is better than never having access.
- thomasjeff1 4mo agoI believe we are not the only ones
- Davidzheng 4mo agoIs the total compute capacity outside of meta, google, amazon, anthropic, oai and x is higher than even the capacity of any of them? In any case, there's no chance a public collaboration gets to anthropic levels of compute even if communication were no issue.
- kelnos 4mo agoIs the issue that training with less compute takes more time? Or is it just not possible? I think a collective using distributed training could tolerate the idea that it takes 10x as long as Anthropic to train a model, or whatever.
- mike_hearn 4mo agoIt's possible but it's not linear. A modern AI training cluster is a supercomputer that uses very different architectures and hardware to a bunch of small PCs connected via normal networking. The networking advantage alone kills any chance of decentralized training.
- laserx 4mo agothere are some strong open source groups like NOUS research taking the fight https://nousresearch.com/ https://nousresearch.com/
- edg5000 4mo agoIt seems this project is serious and very promising. They have the Psyche network which seems real and operational. They're able to produce ~50B-class models, this will only grow over time of course. Very cool.
- ai_fry_ur_brain 4mo ago[flagged]
- bot403 4mo agoThe first half of your comment is unnecessarily aggressive and dismissive to op.
- ai_fry_ur_brain 4mo agoOkay
- palisade 4mo agoSomeone with AI psychosis would say it was easy. I'm saying the opposite. I'm stating that it'd be cool, but at the moment I don't see how it is feasible. And, for fun I tried to solve one small aspect of the problem. I also didn't bring up the concept out of nowhere, this is in response to an article about open source AI. The premise of the post is releasing control to the public. What is more open than a decentralized system? And, why wouldn't you brainstorm in a comment on such a thread? I also didn't ask an AI for the idea, it's just an idea I have. There's a difference.
- girvo 4mo ago> The total power of all GPUs on the planet dwarf their capabilities That just isn't true. It misunderstands exactly how much silicon has gone directly to those companies, and exactly how much more powerful said silicon is compared to consumer grade gear.
- sho 4mo agoIf folding@home is a useful yardstick by which we might estimate the amount of GPU-ish capability that civilians might be coaxed into donating to a shared enterprise, yeah, it doesn't look pretty. This is extremely rough napkin math but comparing to xAI's Collosus 2 for example, for training workflows you're probably looking at 4-5 orders of magnitude the capability of all of folding@home combined. That's 100,000 times faster. Very rough math like I said but I doubt it's directionally wrong. And even if you did force literally everyone on earth with some sort of GPU to max it out 24/7 in service of an open source AI training enterprise - you would waste so much power trying to use that inefficient consumer hardware with the worst latency imaginable that it would be cheaper and faster to get everyone to instead chip in some cash to buy a datacenter with blackwell chips instead! So the idea has no legs whatsoever.
- WithinReason 4mo agofolding@home reached 2.43 exaflops by April 12, 2020, which would make it the largest supercomputer on the planet.
- sho 4mo agoit's down 99% since that peak. But let's compare to it anyway. It's pretty useless to compare raw FLOPS, but as a general hand-waving guesstimate, F@H is currently doing about 25 petaflops in a mix of FP16 and 32. AI usually trains at FP8, but to keep things fair the H100 is quoted at 60 FP64 teraflops per unit, so that's 12 FP64 exaflops given its 200k count. So F@H at its peak did 2.43 exaflops@FP16/32. Colossus 1 does 12@FP64. These numbers are very hand-wavy, but I think the point is made. By the way, I'm not trying to crap on F@H - I think it's an outstanding project and I've run it in the past. But a volunteer group simply cannot compete with well-funded, concentrated effort like what's going into AI.
- Catloafdev 4mo agoYa that'd be an awesome project, the only issue is how do you verify it's not being poisoned? To actually validate it would require more analysis than the training took to run. It would require a trusted network, not an open one, unless that can get solved somehow.
- sgsjchs 4mo agoMake multiple nodes do the same job, compare results.
- deleted 4mo ago[deleted]
- deleted 4mo ago[deleted]
- rustcleaner 4mo agoCould it be done by making a sparse MoE of thousands, or tens of thousands, of smaller experts in very niche domains? Maybe a tree-like structure of experts which can delegate from relatively general but inaccurate to extremely niche but accurate? Also these experts might be plug-and-play, easily swap out an inferior expert with a stronger one in the future without having to redo the whole pile?
- Zetaphor 4mo agoThat's not really how the experts in an MoE work. They activate on token probabilities and are activated on every token. You don't necessarily have a discrete math expert and a discrete physics expert. And if it were you would still need a router that is trained on all of those domains.
- yorwba 4mo agoMoE models are typically designed for datacenter deployment, where per-token load-balancing is more important, but it's also possible to use a different training objective that encourages domain-specialization of experts: https://allenai.org/blog/emo https://allenai.org/blog/emo But yes, this isn't really useful for distributed training as such because of the router.
- trenchgun 4mo ago>But when people think of decentralized training, they don’t first think of gigantic datacenters, owned by the same company, training models across large distances. Instead, they imagine thousands of small datacenters, or individual consumers, pooling their spare compute over the internet to orchestrate a training run larger than any single actor could manage alone. Many companies are pursuing this vision: Pluralis Research, Prime Intellect and Nous Research have already successfully decentrally trained models at scale. But in practice, training decentrally over the internet has lagged far behind more centralized training. Even their largest models (Pluralis’ 8B Protocol Model, Prime Intellect’s INTELLECT-1, and Nous’ Consilience 40B) have been trained with 1,000x less compute than today’s frontier models (such as xAI’s Grok 4). https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale https://epoch.ai/gradient-updates/how-far-can-decentralized-...
- killerstorm 4mo agoI think it's fundamentally not useful as long as there are other open source model releases. E.g. suppose you make SotA model at a particular size via decentralized training. Amazing. In a month Qwen/Deepseek/etc release a new model which is better. So why would you use the "decentralized one"? Models have limited shelf live while things are improving rapidly, and decentralized training is just more wasteful. However, things might change if we get to what Karpathy calls "cognitive core" - a stable model backbone which can be extended via skills/adapters/etc. Development of extensions to the core can be a lot more decentralized. But for now these decentralized training attempts function largely as a deterrent to anti-open-source collusion
- sho 4mo agoAs I replied to a child comment - this is a nice idea that just isn't tenable in reality. AI hardware isn't just hilariously faster than consumer GPUs, it's also hilariously more power-efficient and has hilariously better connectivity. Every one of these dimensions kills the idea. The far, FAR superior power efficiency means that even if you did harness every public GPU or GPU-like device on earth, you'd end up consuming so much excess electricity it would be cheaper on net to simply take the money that would have gone to the power bill and spend it on your own datacenter. And even if electricity was free, having those GPUs spread over the world with internet-level latency will slow everything down by factors of thousands to millions - if it's feasible at all. Regardless, you're not getting fable-oss this decade, maybe even not this century. It would be better for governments to buy and own their own datacenters, maybe as a coalition, and dedicate their operation to the public good. I believe that is what we actually have to do.
- ux266478 4mo agoAI hardware is for inference, not training. Training uses normal HPC crap. Superpods aren't really power efficient, it's kind of a meme, and it stems from limiting the power draw of other components by having less of them. It's more of a rounding error. > you'd end up consuming so much excess electricity it would be cheaper on net to simply take the money that would have gone to the power bill and spend it on your own datacenter. Costs spread over a large population, it really doesn't matter. You're not getting hundreds of thousands of people to pitch half their monthly electric bill to pay for someone else's datacenter. They will pay the electricity themselves quite happily though, if all they need to do is give you compute. This isn't new. Interconnect is the bottleneck for distributed training, nothing else really.
- pksebben 4mo agoBit of a doozie though, that one. I recall getting really excited over hinton's FF foray, right before he bailed on AI as a societal direction (which, if anyone ever had the right, I suppose he does). If one squints, one can see a backprop-free base being much easier to train on geographically distributed and heterogenous hardware.
- 4mo ago
- slashdave 4mo agoWell, I suppose it is understandable why you want to attack the most obvious problem with such a scheme: obtaining sufficient compute. That does mean you are actually neglecting the more difficult issues.
- labbett 4mo agoSounds like SETI@home but for AGI... SAGI@home?
- DonHopkins 4mo agoSince SAGI can't be practically distributed, and it puts so many people out of work, how about moving all of the unhoused people into the nice warm data centers, and call it home@SAGI. Or is that too close to the plot of The Matrix?
- whiplash451 4mo agoThis could be of interest to you: https://thealliance.ai/projects/tapestry https://thealliance.ai/projects/tapestry
- procflora 4mo agoMan, that project is such bait for my particular sensibilities but just looking at the copy about not sharing your data and only sharing weights has me feeling very disappointed in the project already. I would want a project like this to not elide fact that sharing your weight updates probably effectively means sharing your data too.
- cpdomina 4mo agothere was a project trying to achieve some of those goals a few years ago using p2p: petals https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals their bloom model was also a collaborative effort https://huggingface.co/docs/transformers/en/model_doc/bloom https://huggingface.co/docs/transformers/en/model_doc/bloom
- androiddrew 4mo agoI was wondering what happened to this
- andai 4mo ago>The communication speeds are untenable. Can it be parallelized or not? If you take a model, make two copies, and fine-tune each one on different data, what happens when you merge them? Does it work if you freeze different layers? I think this works if the steps are small enough. And the transfer should become tenable if the steps are big enough. Where's the cutoff?
- mike_hearn 4mo agoYes it can be parallelized, it already is in real AI datacenters and no it doesn't help you. Like everyone else is saying, an AI datacenter is not just a bunch of gaming GPUs connected via normal ethernet and hasn't been for years. At most a decentralized effort could contribute a little bit to some bigger centralized effort by doing inference and sandboxed CPU work. Modern model training isn't just backprop, it's got a huge and growing CPU and inferencing component too, which doesn't require intense inter-node communication. For instance, doing RL rollouts for agentic coding requires a lot of plain old inferencing and sandboxed containers for the models to practice in. The final results are just a set of rollouts and scores that can be uploaded back to a central datacenter for GRPO to adjust the weights (relatively cheap). But then, of course, you'd have to stick to models small enough to fit on people's computers so it'd never be competitive.
- andai 4mo agoKinda sounds like we just need better computers.
- dangerlego5 4mo ago[flagged]
- whateverboat 4mo agoThe biggest problem is accuracy and integrity of the actors in the project.
- deleted 4mo ago[deleted]
- WithinReason 4mo agoThe gradient info can be compressed 10000x with the right tricks, I think it is achievable. Nous claims they did it already: https://github.com/NousResearch/DisTrO https://github.com/NousResearch/DisTrO There are other gradient compression papers from the past reporting large compression rates
- deleted 4mo ago[deleted]
- monkeydust 4mo agoDon't know but could BOINC setup which has been around for ages and mature plus has some incentive mechanism (Gridcoin) be used for this?
- logicchains 4mo ago>I've been contemplating a decentralized model training system for some time using volunteer machines that we all contribute. But, it is astronomically difficult. The communication speeds are untenable. It is already possible: https://arxiv.org/abs/2603.08163 https://arxiv.org/abs/2603.08163 . You don't need to sync so frequently, so it can be done over normal internet, it's just less efficient (takes longer to converge).
- dominotw 4mo agowe will be better of doing political activism for govt to provide open researchers and builders access to gpu in govt built dataceter
- mycall 4mo agoMaybe the training approaches taken to date are wrong for decentralized systems. Setup a virtual subnet you can trust and do training on that. Create a AI model island in a trusted/federated model system -- definitely slower than the typical 'one big model' approach, but scalable to world size modeling. Also, it wouldn't be able to use a transformer architecture. For inspiration, take a look at Google Maps and how it a much more efficient A* divide/conquer hill-climbing architecture. Think minimized matrix math.
- bradfa 4mo agoOther comments also hint at this idea, a distributed training solution is currently an open research problem. Solving it is not easy, yet. But 10 years ago what we have today for LLMs would have looked similarly impossible, so have hope, and apply yourself to the problem if you find it interesting!
- incognito124 4mo agohttps://learning-at-home.github.io/ https://learning-at-home.github.io/
- 0xpgm 4mo agoThere are some attempts at this problem, like Bittensor, Akash Network etc
- Andrew_sooter 4mo agoHave you checked out [petals](https://petals.dev/ https://petals.dev/) It’s doing the same thing, however the project is written in python and there can be some optimizations to make it much more faster.
- hajile 4mo agoAI with blockchain. Maybe we can mix in IoT and VR for the ultimate in buzzword synergy.
- AtlasBarfed 4mo agoI just read something or saw something about in document recalculation being a completely wasted step in every single training run. Is that true
- taylorhou 4mo agoLet's collab. I'm one guy too but I built distributed inference network (teale.com) banging away for about a month with opus/gpt