4 ms·
Mistral 7B [1] and many models stemming from it are released under permissive Apache license. Some might argue that a "pure" open-source would require the data
by ebalit 3y ago
Mistral 7B [1] and many models stemming from it are released under permissive Apache license.
Some might argue that a "pure" open-source would require the dataset and the training "recipe" as it would be needed to reproduce the training, but it would be so expensive that most people wouldn't be able to do much with it.
IMO, a release with open weights without the "source" is much better than the opposite, a release with open source and no trained weights.
And it's not like there was no progress on the open dataset front:
- Together just released RedPajama V2, with enough tokens to train a very sizeable base model.
- Tsinghua released UltraFeedback which allowed more people to align models using RLHF methods (like the Zephyr models from Hugging Face)
- and many many others
[1] https://mistral.ai/news/announcing-mistral-7b/ https://mistral.ai/news/announcing-mistral-7b/
[2] https://github.com/togethercomputer/RedPajama-Data https://github.com/togethercomputer/RedPajama-Data
- graphGL 3y ago> IMO, a release with open weights without the "source" is much better than the opposite, a release with open source and no trained weights. Why not both?
- HPsquared 3y agoEven if they don't release the (copyrighted) training data itself, do they provide a "recipe" to reproduce it? Like, a list of sources used and how they were harvested?
- ebalit 3y agoIt really depends on the release. It's not the case for the base Mistral 7B model. But many finetuned models are released with some information about the dataset used.
- DonHopkins 3y agoHow about "Open Binary"?
- andy99 3y ago> Some might argue that a "pure" open-source would require the dataset and the training "recipe" as it would be needed to reproduce the training, but it would be so expensive that most people wouldn't be able to do much with it. I've actually argued the opposite. https://www.marble.onl/posts/considerations_for_copyrighting_AI.html https://www.marble.onl/posts/considerations_for_copyrighting... When you look at the freedoms underlying open source, you can exercise them without the training data. Most of what disqualifies the licenses I mentioned above from being classically open source is use restrictions.
- KRAKRISMOTT 3y agoYes, I think most of the misunderstanding comes from people who are not in ML. A lot of the times you can't release the training pipeline without getting into grey areas of training data copyright. A ML model is a high dimensional distribution that can be sampled from infinitely, given enough time and money you can train another model (possibly even with a different architecture) that closely models the original model in generation capability.