9 ms·
What's the point of downloading it when it'd just stagnate? This isn't like regular software where people can easily put in hard work and sweat to improve it.
by Vetch 4y ago
What's the point of downloading it when it'd just stagnate? This isn't like regular software where people can easily put in hard work and sweat to improve it.
LLMs have the unfortunate limitation of being both powerful and lending themselves to centralized control choke-points due to how resource intensive they are to train. Under this paradigm, I fear commercial entities will be able to easily navigate the legal landmines and continually improve while open efforts perpetually lag far behind.
There are many vested interests who want this control for various reasons they justify as: protection from x-risk, keeping it out of the hands of abusers and bullies, economic advantage. Their reasons for want of control are either well intended but wrong-headed or profit-motivated and disingenuous.
Rather than challenging the likes of GPT-3 and Copilot enabling freedom, I fear folks will be forced to send all their videos, pictures, text and code to the servers of Microsoft, Amazon and Google or lose access to advantages as LLMs continue to improve at a rapid clip.
- visarga 4y ago> LLMs have the unfortunate limitation of being both powerful and lending themselves to centralized control choke-points It was hard to accomplish, but you can finetune SD on your computer. They are working on instruction-tuning LLMs as well. In general ML models are not closed boxes inaccessible to us - they can be finetuned, reprompted, you can even average two versions to get a mix of two models. In the last 2 years lots of papers were written on finetuning and prompting, all of them geared towards low resource AI adaptation to new tasks.
- dividedbyzero 4y agoBut you can't selectively re-train them, can you? As in, don't use elements from this part of the training data anymore, but use elements from this body of work that wasn't part of the training data? If I understand correctly you'd still need a full re-training for that.
- visarga 4y agoWhat you can do is - lexical filtering by applying a blacklist of artist names on the original prompt - perceptual filtering - drop all generated images that look too close to copyrighted images in your training set - re-captioning based filtering - use a model to generate captions for an image and apply filters on the captions; you can also filter by visual style - CLIP based filtering where you use embeddings to find nearest neighbours, and if they are copyrighted then you can drop the image - or train a copyright violation detection model that takes generated images and compares them to images from the original authors Copyright enforcement struggles are going to be interesting to watch in this decade. But I think it will slowly become irrelevant, because anything can be generate again slightly different until they finally pass the filters.
- dividedbyzero 4y agoI was aiming more at the centralized-control angle (though I didn't make that very clear), i.e. are open-source models actually viable long-term? If only orgs with absurd amounts of compute can do updates because those imply a full re-training, wouldn't that effectively centralize control over any such model? Is there the option to to an incremental, limited re-training?
- ShamelessC 4y agoMuch of modern deep learning is actually premised on the discovery that training on a large, noisy dataset _first_, and then fine tuning (starting training on new data with the same weights) is generally quicker to converge, and also more accurate. This is part of the motivation for “foundation models”. There’s another paradigm called student/teacher models where a randomly initialized model updates it’s weights according to another pretrained model. This could (maybe?) be used to achieve the desired effect of a model that learned in a “clean room”.
- plutonorm 4y agoyou can retrain on completely separate data - I am currently doing this.
- Vetch 4y agoThat's fine until the next brand new model based on a better architecture where the above hacks won't suffice. My concerns here are long term, like 1 or 2 years out in AI-years.
- frognumber 4y ago> What's the point of downloading it when it'd just stagnate? Because it's already good enough to have made it's way into many of my workflows. I do feel that many companies will, ironically, use "ethical" as a pretext to not be open.
- lvncelot 4y ago> I do feel that many companies will, ironically, use "ethical" as a pretext to not be open. I mean this isn't even speculative anymore after what happened with - hilariously named - OpenAI
- Vetch 4y agoWhat about future models with fewer artifacts that are much easier to communicate with and better at generation? Opportunity costs might favor just sending your data to and paying corps with compliance guarantees than spend time fiddling with 2022 era diffusion models. And don't forget this affects 3D, video snippets and music going forward. > that many companies will, ironically, use "ethical" as a pretext to not be open. Yes, weaponized ethics as sleight of hand for control is a common historical pattern.
- pmoriarty 4y ago"Opportunity costs might favor just sending your data to and paying corps with compliance guarantees than spend time fiddling with 2022 era diffusion models." This is exactly why I pay $30 per month for MidJourney. The output is just phenomenally better than most of the images coming out of SD, and the UI is much better as well. It's just not worth my time fiddling with SD if the results are so bad in comparison. If/when SD catches up, I'd jump ship to using it in a heartbeat.
- danuker 4y ago> when it'd just stagnate? While it'd be difficult to improve upon the model, it might be easy enough to finetune it if needed, and it's certainly worth it to USE it as is. There is a limited number of models costing 6 digits in dollars in train time and are freely available. There is certainly value in preserving them, in a world of artificial scarcity.
- deepserket 4y ago> I fear commercial entities will be able to easily navigate the legal landmines and continually improve while open efforts perpetually lag far behind Is it possible to crowdsource AI training with something that looks similar to folding@home?
- pmoriarty 4y agoIt's not just processing power that smaller open projects lack in comparison to large corporations, but data. AI thrives and depends on large amounts of clean, well labeled data. Large corporations understand this and have hoarded data for a long time now. Some of them have also managed to label this data by millions of people through things like Recaptcha, or just by hiring lots of people to do it. Open datasets tend to be much smaller and dirtier than small, open projects have access to. I suppose it would be possible to, over time, collect lots of data and crowd-source some project to clean it up and label it well enough to be useful, then crowd-source the AI model training itself, but it would probably take a long time and by then corporate-owned AI models will already dominate (as they do now with MidJourney, for example, being way better in my experience than Stable Diffusion, but with time the difference will only get starker). I'd also be concerned with such ostensibly open projects eventually going closed and commercial as IMDB did after getting lots of work by volunteers freely giving their time to writing reviews.
- dougabug 4y agoData can be crowd sourced, too. Wikipedia demonstrated that crowdsourced data can be pretty competitive. More recently the open LAION data sets have become widely used by both tech giants and independent researchers.
- rfoo 4y ago> Wikipedia demonstrated that crowdsourced data can be pretty competitive. The problem is DL is really sensitive to dirty data, disproportionately so. At $DAYJOB once we cleaned the dataset, removed a few mislabeled identity/face pairs (very few, about 1 in 1e4) and the metrics goes up a lot.
- petercooper 4y agoWhat's the point of downloading it when it'd just stagnate? The quality of the output you can get with the models right now have perpetual utility IMO. If you use it to create patterns, backgrounds, or even just for inspiration creations right now, it might be a shame if it didn't progress (depending on your position) but it's fine as-is if you put in the work to compose and refine the raw output.
- gauravvij137 4y agoThe only way to get rid of centralized choke points is to actually go decentralized. At Q Blocks, we're working on making this solution a reality for a lot of the ML devs constrained by the computing costs on cloud.
- RobotToaster 4y ago>LLMs have the unfortunate limitation of being both powerful and lending themselves to centralized control choke-points due to how resource intensive they are to train. I wonder if that will continue. My understanding is that's partially because it currently relies on GPUs, which until relatively recently there was a limited demand for, and the market is basically controlled by a single company. Will we see cheaper special purpose AI accelerators? Like happened with crypto mining ASICs.