4 ms·
The repo source code is Apache 2.0 licensed, but the weights are not. Just in case anybody else is excited then misled by their tagline "Building the Next Gene
by orra 3y ago
The repo source code is Apache 2.0 licensed, but the weights are not.
Just in case anybody else is excited then misled by their tagline "Building the Next Generation of Open-Source and Bilingual LLMs".
- est31 3y agoMore reading on the weight license: https://news.ycombinator.com/item?id=38159862 https://news.ycombinator.com/item?id=38159862
- mmastrac 3y agoThe model license, excerpts: https://github.com/01-ai/Yi/blob/main/MODEL_LICENSE_AGREEMENT.txt https://github.com/01-ai/Yi/blob/main/MODEL_LICENSE_AGREEMEN... 1) Your use of the Yi Series Models must comply with the Laws and Regulations as well as applicable legal requirements of other countries/regions, and respect social ethics and moral standards, including but not limited to, not using the Yi Series Models for purposes prohibited by Laws and Regulations as well as applicable legal requirements of other countries/regions, such as harming national security, promoting terrorism, extremism, inciting ethnic or racial hatred, discrimination, violence, or pornography, and spreading false harmful information. 2) You shall not, for military or unlawful purposes or in ways not allowed by Laws and Regulations as well as applicable legal requirements of other countries/regions, a) use, copy or Distribute the Yi Series Models, or b) create complete or partial Derivatives of the Yi Series Models. “Laws and Regulations” refers to the laws and administrative regulations of the mainland of the People's Republic of China (for the purposes of this Agreement only, excluding Hong Kong, Macau, and Taiwan).
- heroprotagonist 3y agoNot really very open. Though it makes me wonder what the model might say about Tiananmen Square, Uyghurs, reeducation camps, or any other thing you're not really supposed to talk about in China. Is there a benchmark for bias in model outputs? No doubt China has one, somewhere, except it's not skewed towards prevention.
- echelon 3y agoWeights are trained on copyrighted data. I think that ethically, weights should be public domain unless all of the data [1] is owned or licensed by the training entity. I'm hopeful that this is where copyright law lands. It seems like this might be the disposition of the regulators, but we'll have to wait and see. In the meantime, maybe you should build your product in this way anyway and fight for the law when you succeed. I don't think a Chinese tech company is going to find success in battling a US startup in court. (I would also treat domestic companies with model licenses the same way, though the outcome could be more of a toss up.) "Break the rules." "Fake it until you make it." Both idioms seem highly applicable here. [1] I think this should be a viral condition. Finetuning on a foundational model that incorporates vast copyrighted data should mean downstream training also becomes public domain.