5 ms·
Open source public models trained on kosher data are substantially derisking the AI hype. It makes a lot of sense to push this approach as far as it can get. It
by openrisk 2y ago
Open source public models trained on kosher data are substantially derisking the AI hype. It makes a lot of sense to push this approach as far as it can get. Its similar to SETI at home etc. but potentially with far more impact.
- grumbelbart2 2y agoThey could, would and should. But: Training a state of the art LLM costs millions in GPU, electricity alone. There is no "open" organization at this point that can cover this. Current "open source public models" are shared by big players like Meta to undermine the competition. And they only publish their weights, not the training data, training protocols, training code; meaning it's not reproducible, and questionable if the training data is kosher.
- sebmellen 2y agoDoesn’t Deepseek somewhat counter this narrative?
- htrp 2y agoDon't they have something like 10k plus current gen GPUs?
- lz400 2y agoI understood SETI style meaning crowdsourced. Instead of mining bitcoin you mine LLMs. It's a nice idea I think. Not sure about technical details, bandwidth limitations, performance, etc.
- Mountain_Skies 2y agoSETI had a clear purpose that donors of computer resources could get behind. The LLM corps early on decided to drink the steering poison that will keep there from ever being a united community for making open LLMs. At best you'll get a fractured world of different projects, each with its own steering directives.
- Aerroon 2y agoThe internet is for ____. That could be a factor that unites enough people to donate their compute time to build diffusion models. At least if it was easy enough to set up.
- CaptainFever 2y agoRelated: people donating computing power to run diffusion and text models, which is definitely largely used for porn. https://stablehorde.net/ https://stablehorde.net/ Or the large amounts of community efforts (not exactly crowd sourced though) for diffusion fine-tunes and tools! Pony XL, and other uncensored models, for example. I haven't kept up with the rest, because there's just too much.
- lostmsu 2y agoYou don't have to donate, we will pay you for idle time of your gaming GPU: https://borg.games/setup https://borg.games/setup
- HPsquared 2y agoUnfortunately, LLM training is not as computationally easy (embarrassingly parallel) as mining bitcoins.
- lz400 2y agodamn it! but nice research area
- ghxst 2y agoIf that were to be solved (if at all possible, and feasible / competitive) I can definitely see "LLM mining" be a historic milestone. Also much closer to the spirit of F@H in some sense, depending how you look at it. Would there be a financial incentive? And how would it be distributed? Could you receive a stake in the LLM proportional to the contribution you did? Would that be similar in some sense to purchasing stock in an AI company, or mining tokens for a crypto currency? Potentially a lot of opportunity here.
- HPsquared 2y agoThe models are too large to fit on a desktop GPU's VRAM. Progress would either require smaller models (MoE might help here? not sure) or bigger VRAM. For example training a 70 billion parameter model would require at least 140GB of VRAM in each system, whereas a large desktop GPU (4090) has only 24GB. You need enough memory to run the unquantized model for training, then stream the training data through - that part is what is done in parallel, farming out different bits of training data to each machine.
- mr_toad 2y agoData parallel training is not the only approach. Sometimes the model itself needs to be distributed across multiple GPU. https://www.microsoft.com/en-us/research/blog/zero-deepspeed-new-system-optimizations-enable-training-models-with-over-100-billion-parameters/ https://www.microsoft.com/en-us/research/blog/zero-deepspeed... The communications overhead of doing this over the internet might be unworkable though.
- myrmidon 2y agoI think it's important to remember that we know neural networks can be trained to a very useful state from scratch for 24 GJ: This is 25 W for 30 years (or 7000 kWh, or a good half ton of diesel fuel), which is what a human brain consumes until adulthood. Even though our artificial training efficiency is worse now, likely to stay worse because we want to trade efficiency for faster training, and because we want to cram more knowledge into our training data than a human would be exposed to, it still seems likely to me that we'll get within orders of magnitude of this sooner or later. Even if our training efficiency topped out at a hundred times worse than a biological system, that would be the energy equivalent of <100 tons of diesel fuel. Compared to raising and educating a human (and also considering this training can the be utilized for billions of queries before it becomes obsolete) that strikes me as a very reasonable cost (especially compared to the amounts of energy we wasted on cryptocurrency mining without blinking an eye...)
- dbspin 2y agoThis misses that evolution has been pre-training the human cognitive architecture - brain, limbic system, sympathetic and parasympathetic nervous systems, coevolved viral and bacterial ecosystems - for millions of years. We're not a tabula rasa training at birth to perfectly fit whatever set of training data we're presented. Far from it. Human learning is more akin to RAG, or test time training - specialising a heavily pre-trained model. It's not that we're born with very much knowledge, it's more that we're heavily specialised to acquire and retain certain kinds of knowledge and behaviour that are adaptive in the EEA (environment of evolutionary adaptedness). If the environment then doesn't provide the correct triggers at the correct times for activation of various learning mechanisms - best known being the critical period for language acquisition, we don't unfold into fully trained creatures. Bear in mind also that the social environment is vital both for human learning and functioning - we learn in the emotional, cognitive and resource provision context of other humans. And what we learn are behaviours that are effective in that context. Even in adulthood, the quickest way to make our cognitive architecture break down is to deny us social contact (hence the high rates of 'mental illness' in solitary confinement).
- ben_w 2y ago
- TZubiri 2y agoAsking to share the training data is a bit too much, it's petabytes of data, probably has privacy implications. You can study and reproduce with your own training data right?
- agentultra 2y agoProbably legals ones too. Such aa evidence of copyright infringement.
- deleted 2y ago[deleted]
- musha68k 2y agoIt's been in the air; and bittensor, hyperbolic & co have been on some of the angles for a while. Models, by the people and for the people, Wikipedia + "SETI at home" style. Eventually and with pre/training ofc, this will include inference too.
- medion 2y agoAren’t there a ton of blockchain projects trying to do this kind of distributed compute / LLM with tokenisation rewards etc? Theta project comes to mind
- ben_w 2y agoI don't see how it really helps? We could pass laws requiring models to demonstrate their training sets irregardless of how the training is distributed; and conversely if this is a community-led project, those also have copyright issues to deal with (wikipedia for example). I suspect there's also a problem in that, e.g. ten million student essays about different pages of Harry Potter can each in isolation be justified by the right to quote small fragments for critical purposes, but the collection together isn't because it quotes an entire book series.
- brookst 2y agoI think you’re doing away with the fairness exception for criticism? Copyright is intended to reward investment in creative works by giving sole license to distribute. It is not intended to create a monopoly on knowledge about the work. If I can ask an LLM (or person!) “what’s the first sentence in Harry Potter?” And then “what’s the second sentence?” and so on, that does not mean they are distributing the work in competition with the rights holders. We have gone way overboard with IP protections. The purpose of copyright is served when Rowling buys her 10th mansion. We do not need to further expand copyright to make it illegal to learn from a work or to remember it after reading.
- ben_w 2y ago> I think you’re doing away with the fairness exception for criticism? Perhaps, but it's more an example of the problem: something can be fine at small scale, but cause issues when everyone does it. Tragedy of the commons, but with words. (From an even more extreme point of view, consider that an image generating AI trained on nothing but photographs taken from drones flying and androids walking all over the place would be able to create photo-realistic images of anything, irregardless of if even one single human artist's works end up in the training set, which in turn means the current concerns about "did the human artists agree to this use" will be quickly made irrelevant because there were none in the training set in the first place). "Quantity has a quality all its own", whoever really said it first. > Copyright is intended to reward investment in creative works by giving sole license to distribute. It is not intended to create a monopoly on knowledge about the work. Sure, but laws change depending on economics. I can easily believe AI will lead to either much stronger or much weaker copyright laws. Depends who is wielding the power when the change comes. > If I can ask an LLM (or person!) “what’s the first sentence in Harry Potter?” And then “what’s the second sentence?” and so on, that does not mean they are distributing the work in competition with the rights holders. Isn't that a description of how BitTorrent works? And The Pirate Bay is kinda infamous for "distributing the work in competition with the rights holders". > We have gone way overboard with IP protections. The purpose of copyright is served when Rowling buys her 10th mansion. We do not need to further expand copyright to make it illegal to learn from a work or to remember it after reading. I agree, and was already in favour of radical changes to copyright rules well before LLMs. (That said, it's more complex because of how hit-driven lots of things are, which means that while nobody needs to defend Rowling's second billion, having looked at the distribution of book sales in the best-seller lists… most of them will/would have need/ed a second source of income to keep publishing).
- netdevphoenix 2y agowhat's "kosher data"? Never heard of that before