4 ms·
I'm just a lowly outsider to the AI space, but calling these open source models seems kind of like calling a compiled binary open source. If you don't have a w
by quasse 2y ago
I'm just a lowly outsider to the AI space, but calling these open source models seems kind of like calling a compiled binary open source.
If you don't have a way to replicate what they did to create the model, it seems more like freeware than open source.
- advael 2y agoAs an ML researcher, I agree. Meta doesn't include adequate information to replicate the models, and from the perspective of fundamental research, the interest that big tech companies have taken in this field has been a significant impediment to independent researchers, despite the fact that they are undeniably producing groundbreaking results in many respects, due to this fundamental lack of openness This should also make everyone very skeptical of any claim they are making, from benchmark results to the legalities involved in their training process to the prospect of future progress on these models. Without being able to vet their results against the same datasets they're using, there is no way to verify what they're saying, and the credulity that otherwise smart people have been exhibiting in this space has been baffling to me As a developer, if you have a working Llama model, including the source code and weights, and it's crucial for something you're building or have already built, it's still fundamentally a good thing that Meta isn't gating it behind an API and if they went away tomorrow, you could still use, self-host, retrain, and study the models
- warkdarrior 2y agoThe model is public, so you can at least verify their benchmark claims.
- advael 2y agoGenerally speaking, no. An important part of a lot of benchmarks in ML research is generalization. What this means is that it's often a lot easier to get a machine learning model to memorize the test cases in a benchmark than it is to train it to perform a general capability the benchmark is trying to test for. For that reason, the dataset is important, as if it includes the benchmark test cases in some way, it invalidates the test When AI research was still mostly academic, I'm sure a lot of people still cheated, but there was somewhat less incentive to, and norms like publishing datasets made it easier to verify claims made in research papers. In a world where people don't, and there's significant financial incentive to lie, I just kind of assume they're lying
- celdon25 2y agoWhich option would be better? A) Release the data, and if it ends up causing a privacy scandal, at least you can actually call it open this time. B) Neuter the dataset, and the model All I ever see in these threads is a lot of whining and no viable alternative solutions (I’m fine with the idea of it being a hard problem, but when I see this attitude from “researchers” it makes me less optimistic about the future) > and the credulity that otherwise smart people have been exhibiting in this space has been baffling to me Remove the “otherwise” and you’re halfway to understanding your error.
- wanderingbort 2y ago> Release the data, and if it ends up causing a privacy scandal... We can't prove that a model like llama will never produce a segment of its training data set verbatim. Any potential privacy scandal is already in motion. My cynical assumption is that Meta knows that competitors like OpenAI have PR-bombs in their trained model and therefore would never opensource the weights.
- advael 2y agoThis isn't a dilemma at all. If Facebook can't release data it trains on because it would compromise user privacy, it is already a significant privacy violation that should be a scandal, and if it would prompt some regulatory or legislative remedies against Facebook for them to release the data, it should do the same for releasing the trained model, even through an API. The only reason people don't think about it this way is that public awareness of how these technologies work isn't pervasive enough for the general public to think it through, and it's hard to prove definitively. Basically, if this is Facebook's position, it's saying that the release of the model already constitutes a violation of user privacy, but they're betting no one will catch them If the company wants to help research, it should full-throatedly endorse the position that it doesn't consider it a violation of privacy to train on the data it does, and release it so that it can be useful for research. If the company thinks it's safeguarding user privacy, it shouldn't be training models on data it considers private and then using them in public-facing ways at all As it stands, Facebook seems to take the position that it wants to help the development of software built on models like Llama, but not really the fundamental research that goes into building those models in the same way
- celdon25 2y ago> it seems more like freeware than open source. What would you have them do instead? Specifically?
- ori_b 2y agoRelease the training set and the code that was used to train the model, or stop calling it open source. If you can't fork it and take the project in your own direction, it's not open source.
- wongarsu 2y ago> If you don't have a way to replicate what they did to create the model, it seems more like freeware Isn't that a bit like arguing that a linux kernel driver isn't open source if I just give you a bunch of GPL-licensed source code that speaks to my device, but no documentation how my device works? If you take away the source code you have no way to recreate it. But so far that never caused anyone to call the code not open-source. The closest is the whole GPL3 Tivoization debate and that was very divisive. The heart of the issue is that open source is kind of hard to define for anything that isn't software. As a proxy we could look at Stallman's free software definition. Free software shares a common history with open source and in most open source software is free/libre, and the other way around, so this might be a useful proxy. So checking the four software freedoms: - The freedom to run the program as you wish, for any purpose: For most purposes. There's that 700M user restriction, also Meta forbids breaking the law and requires you to follow their acceptable use policy. - The freedom to study how the program works, and change it so it does your computing as you wish: yes. You can change it by fine tuning it, and the weights allow you to figure out how it works. At least as well as anyone knows how any large neural network works, but it's not like Meta is keeping something from you here - The freedom to redistribute copies so you can help your neighbor: Allowed, no real asterisks - The freedom to distribute copies of your modified versions to others: Yes So is it Free Software™? Not really, but it is pretty close.
- advael 2y agoThe model is "open-source" for the purpose of software engineering, and it's "closed data" for the purpose of AI research. These are separate issues and it's not necessary to conflate them under one term