6 ms·
Think seti at home. Instead of wasting all the compute on bitcoin we pretrain fully open models which can run on people's hardware. A 120b ternary model is the
by ctrw 3y ago
Think seti at home.
Instead of wasting all the compute on bitcoin we pretrain fully open models which can run on people's hardware. A 120b ternary model is the most interesting thing in the world. No one can train one now because you need a billion dollar super computer.
- __loam 3y agoI would expect most of the big tech firms have the capital to build a gpu cluster like that.
- kleinsch 3y agoPeople are estimating that Meta is buying $8-10B in GPUs this year alone. https://www.cnbc.com/amp/2024/01/18/mark-zuckerberg-indicates-meta-is-spending-billions-on-nvidia-ai-chips.html https://www.cnbc.com/amp/2024/01/18/mark-zuckerberg-indicate...
- ImprobableTruth 3y agoTransmission speeds aren't fast enough for this, unless you crank up the batch size ridiculously high.
- FeepingCreature 3y agoLoRA training/merging basically is "crank up the batch size ridiculously high" in a nutshell, right? What actually breaks when you do that?
- brrrrrm 3y agoCranking up the batch size kills convergence.
- FeepingCreature 3y agoWonder if that can be avoided by modifying the training approach. Ideas offhand: group by topic, train a subset of weights per node; figure out which layers have the most divergence and reduce lr on those only.
- brrrrrm 3y agoA provable way to recover convergence is to calculate the hessian. It’s computationally expensive but there are approximation methods.
- mirekrusin 3y agoSETI made sense because there is a lot of data where you download chunk and do expensive computation and return thin result. Model training is unlike that. It's large state that is constantly updated and updates require full, up to date state. This means you cannot distribute it efficiently over slow network with many smaller workers. That's why NVIDIA is providing scalable clusters with specialized connectivity so they have ultra low latency and massive throughput. Even in those setups it takes ie. a month to train base model. Converted to distributed setup this same task would take billions of years - ie. it's not feasible. There aren't any known ways of contributing computation without access to the full state. This would require completely different architecture, not only "different than transformers" but "different than gradient descent", which would be basically creating new branch in machine learning and starting from zero. Safe bet is on "ain't going to happen" - better to focus on current state of art and keep advancing it until it builds itself and anything else we can dream of to reach this "mission fucking accomplished".
- sigmoid10 3y agoThat's wrong. What you described is data parallelism and it would indeed be very tricky to e.g. sync gradients across machines. But this is not the only method of training neural nets (transformers or any other kind) in parallel. If we'd like to train, say, a human brain complexity level model with 10^15 parameters, we'd need a model parallelism approach anyways. It introduces a bit of complexity since you need to make sure that each distributed part of the model can run individually with roughly the same amount of compute, but you no longer need to worry about syncing anything (or have the entire state of anything on one machine). The real questions is if you can find enough people to run this who will never be able to run it themselves in the end, because inference alone will still require a supercluster. If you have access to that, you might as well train something on it today.
- mirekrusin 3y agoLack of data parallelism is implied by computation that is performed. You gradient descend on your state. Each step needs to work on up to date state otherwise you're computing gradient descend from state that doesn't exist anymore and your computed gradient descent delta is nonsensical if applied to the most recent state (it was calculated on old one, direction that your computation calculated is now wrong). You also can't calculate it without having access to the whole state. You have to do full forward and backward pass and mutate weights. There aren't any ways of slicing and distributing that make sense in terms of efficiency. The reason is that too much data at too high frequency needs to be mutated and then made readable. That's also the reason why nvidia is focusing so much on hyper efficient interconnects - because that's the bottleneck. Computation itself is way ahead of in/out data transfer. Data transfer is the main problem and going in the direction of architecture that dramatically reduces it by several orders of magnitude is just not the way to go. If somebody solves this problem it'll mean they solved much more interesting problem – because it'll mean you can locally uptrain model and inject this knowledge into bigger one arbitrarily.
- monkeydust 3y agoReminded me of Petals offering distributed ML inference anyone tried this? https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals
- cchance 3y agoyep was shockingly fast actually
- Der_Einzige 3y agoFor the record, this already exists in the open source world for Stable Diffusion. https://stablehorde.net/ https://stablehorde.net/ You can host a local horde and do "Seti at home" for stable diffusion.