5 ms·
Others have answered your question, but I'll add that the market for high quality AI models is not similar to the software marketplace, where there is always an
by icapybara 3y ago
Others have answered your question, but I'll add that the market for high quality AI models is not similar to the software marketplace, where there is always an open source alternative (and where open source is often the state of the art).
LLMs take so much engineering effort, research, and compute that it's unlikely there will be good open source alternatives in the near future. Right now your only real option is OpenAI (or maybe Anthropic) and that seems unlikely to change anytime soon.
The only reason we have LLAMA is because Meta threw us a bone. They might not do that again.
- rjzzleep 3y ago> LLMs take so much engineering effort, research, and compute that it's unlikely there will be good open source alternatives in the near future. Right now your only real option is OpenAI (or maybe Anthropic) and that seems unlikely to change anytime soon. does it though? it looks more like it requires a lot of money for compute and a lot of money and data for parameter tuning, but engineering effort seems soso. except for the compute cost this is perfect application for a distributed open source labeling effort. Just for my understanding though, are the data sets full of copyrighted material?
- muyuu 3y agoSome are, some aren't. See Koala for instance. The problem with Koala is that it fine-tunes on open sourced data, but makes no claims about the data for the base LLaMA models. https://bair.berkeley.edu/blog/2023/04/03/koala/ https://bair.berkeley.edu/blog/2023/04/03/koala/ The irony is that openAI and Meta themselves might be in flaky ground for having trained models on other people data with dubious rights to do so in many instances, and then using it to produce output commercially. But this is a new frontier and enforcement might be effectively not possible unless new legislation requires reproducibility and audits on the data sets or something like that. But without that, how do you know exactly how did they arrive at a given set of weights with Montecarlo algorithms and arbitrary fine tuning? You basically don't know what was there and you cannot prove they didn't achieve those results with perfectly clean data. PS: https://medium.com/geekculture/list-of-open-sourced-fine-tuned-large-language-models-llm-8d95a2e0dc76 https://medium.com/geekculture/list-of-open-sourced-fine-tun...
- Arelius 3y ago> You basically don't know what was there and you cannot prove they didn't achieve those results with perfectly clean data. I mean you totally do though, right? You just need one instance of the LLM reproducing information that would only have been able to by violating copyright. I mean, it's theoretically possible that it could have reproduced it from scratch, infinite monkies on typewriters sort of thing, but statistically we can rule that out on pretty short notice. Adding on to this, I don't think the argument that OpenAI, Google and others are ultimately making will be that they don't violate copyright, but instead will ultimately be that their violation is sufficiently transformative such that it constitutes fair-use.
- muyuu 3y agonot only it's theoretically possible, it happens and it can already be observed on clean lab experiments with normally used parameters the probability that LLMs produce copyrighted information is no proof that it was trained with it exactly, esp. when parameters are set so they don't repeat outputs
- Arelius 3y agoI think you'll find that it is in fact proof by all practical standards we use outside of formal mathematics.
- muyuu 3y agothe moment you cannot in any practical way tell if the data set was corrupted with copyrighted material, nobody will convict you for any accidental violations that may occur, even in the astronomically low probability that they do with standard parameters
- Arelius 3y agoMy point is being able to reliably reproduce copyright works will function as a very practical way to tell if the dataset was corrupted with copyrighted material. In that way it’ll be a lot easier to prove that a dataset was corrupted, then proving the negative.
- kkielhofner 3y agoI try not to predict the future but similar things were said about Open Source in the 90s. Then IBM threw their weight behind it (they were still pretty relevant), RedHat was and is a success, etc. I remember when the scales completely tipped on the Linux kernel and the top X contributors were from Intel, etc as opposed to individual hobbyist devs. Nvidia is an obvious one here - they already do a ton of large model/research work because good models sell a lot of hardware. I would not be surprised at all if in they're already working internally on this (they're due for a new large model/arch release anyway). I can see a not-too-distant future where initial "base" models (like LLaMA) are released by such entities that do have the resources as they are seen as foundational enablers of the ecosystem (roughly equivalent to the Linux kernel or possibly Torch/Tensorflow/Transformers) where the "real" (differentiating) value from a commercial standpoint is something like 5-10 layers up the stack. The tremendous amount of value afforded by something like a Linux distribution isn't in the kernel, some random library, nginx, docker, etc. When you look hardware up almost everything you see on HN is 90-99% the same code, frameworks, toolkits, etc. Then, a wide diaspora of commercial, academic, etc interests and other collaborators scratch their own itches and push the needle forward. Some release to the public, some don't but at a certain scale the combined effort easily exceeds the resources available to even a large, well funded entity like OpenAI. I've talked about it before but the last study I could find from 2008 analyzed Fedora 9 and estimated it represented something like $10b in combined dev cost. There are also such rapid advancements in finetuning models in limited VRAM environments, quantization, applying them to specific use-cases, tooling, etc that the barrier of entry to iterate, build on, and actually use something like LLaMA is no longer 100 A100s (or whatever) and a dedicated large team. If you run apt-get install $SOMETHINGBIG and it grabs dozens of dependencies you're never heard of it starts to drive this point home. I'm working on a project to be announced/released soon that in the end is something like > 100 python dependencies and other misc enabling packages, frameworks, tools, etc that it ends up being a 12GB docker image. Our "magic", meanwhile, is something like 1k LoC. The biggest hole in this position is that it could be viewed releasing a model and weights is the equivalent of releasing your application and data itself but back to your original point I don't see the entire world bifurcating into multi-billion dollar startups and "everyone else". Or maybe I'm just being optimistic :).
- lhl 3y ago> Right now your only real option is OpenAI (or maybe Anthropic) > The only reason we have LLAMA is because Meta threw us a bone IMO, this is pretty inaccurate, you can look at my other post in the thread to see how many other recent and ongoing projects there are. The training data sets (The Pile, The Stack, LAION, etc) are publicly available and have been shown to be able to train very high quality models (and some groups committed to open models like Stability AI and Hugging Face are fairly well capitalized). Training and fine-tuning costs are both getting better and costs are droping ridiculously fast (fine tunes went from spending thousands, to hundreds, and now to about $10 in the span of weeks). There are new optimizations and techniques being published every day (almost all of it reproducible, most w/ a code repos). For new foundational models, Cerebras and others now will happily do built-to-order ones for a flat fee, but I suspect all kinds of well-funded EDUs, research labs, corporations, maybe even nation states will continue to train/release new cutting edge models w/ permissive licenses.
- mejutoco 3y ago> LLMs take so much engineering effort, research, and compute that it's unlikely there will be good open source alternatives in the near future. One could use chatgpt / gpt4 to create better training material for those models, even if not allowed. In that sense there is an advantage to being second here.