7 ms·
Show HN: A fully open-source (Apache 2.0)implementation of llama
We believe that AI should be fully open source and part of the collective knowledge.
The original LLaMA code is GPL licensed which means any project using it must also be released under GPL.
This "taints" any other code and prevents meaningful academic and commercial use.
Lit-LLaMA solves that for good.
- theaniketmaurya 4y agoI am in love with this implementation considering the ability to run on 8 GB VRAM and Apache 2.0 license.
- theaniketmaurya 4y agoI am curious though how would the model weights work out?
- rasbt 4y agoI guess that means time to fire up a few GPUs later today and get some weights! We should have a weight exchange platform for that maybe, haha.
- A4ET8a8uTh0 4y agoYou mean like a blockchain? I jest, but only a little.
- yewnork 4y agoI see this as a win for the AI community. The key for LLMs is to enable people to train collaboratively and innovate more quickly in this space. Are there any examples or demos available that showcase the capabilities of "lit-llama"?
- deleted 4y ago[deleted]
- 2Gkashmiri 4y agoBs. Prevents meaningful academic..... How the hell does agpl prevent academic use? Commercial use sure because agpl follows 4 freedoms and commercial often wants to take someone else's work, slap their brand without acknowledging the original work. That and the downstream is often closed source for "business reasons" which causes their users to not enjoy the fruits of the first party's licensing. Where does academia come into it? Are researchers now keeping everything under wraps for "shareholders interests"? Isn't academia supposed to be open culture from the start without any restrictions so what am I missing or are they mixing two unrelated things? Also, I think I might be wrong but isn't it merely converting llama into their version? Uh ...
- ftxbro 4y agoI'm not saying this is how it should be, but a lot of the author lists of published papers on scaling properties of large language models have been employees in research divisions within big tech companies or academics holding dual positions with those companies and with their university. > Where does academia come into it? Are researchers now keeping everything under wraps for "shareholders interests"? Isn't academia supposed to be open culture from the start without any restrictions so what am I missing or are they mixing two unrelated things? Yeah academia was never perfect, but it's becoming more and more like you describe. It's been happening for a while and that's a whole other thing.
- AmuVarma 4y agoLlama by FB is under a non-commercial license not a GPL license, so I assume you are using a different base model, what model is that?
- sp332 4y agoThis isn't a new model, it's just new code.
- charcircuit 4y agoThe reference inference code is GPL. https://github.com/facebookresearch/llama/blob/main/LICENSE https://github.com/facebookresearch/llama/blob/main/LICENSE
- querez 4y agoIANAL, but this seems very fishy to me: 1) I don't understand how this isn't a derivative work of the original code, as I very highly doubt you've done a clean room implementation. I doubt this would hold up in court. 2) Doesn't the original FB license also apply to the weights? Just re-implementing the code would not change the license on the weights. So while THE CODE may now be re-licensed, the weights would still fall under the original license. I'd love if someone with more legal understanding could shed some light on this.
- MacsHeadroom 4y ago>I don't understand how this isn't a derivative work of the original code The original code is Apache 2 licensed. Derivatives are fine and allowed. This retains the same Apache 2 license as Facebook's code. It's only the model that isn't covered by that permissive Apache 2 license. A model produced by a derivative of the permissively licensed code, or even by the original code itself, is not a derivative or the original non-permissively licensed model produced by the original code and is non-infringing even if it is a bit-perfect replica. > Doesn't the original FB license also apply to the weights? Again, there are different licenses for the code and the model and neither license actually applies to the weights within the model only the actual exact model. If this project produced a bit-for-bit replica of Facebook's model it would still not infringe on that model's license. But it doesn't produce a bit-for-bit replica. Even if Facebook were to re-run their same training code on their same hardware would they could not produce the exact same weights as before since massively parallel matrix multiplications are not deterministic. Benign environmental noise like microscopic fluctuations in temperature make a difference in the outcome.
- kristjansson 4y ago> Apache 2 Isn't the original GPLv3[0]? [0]: https://github.com/facebookresearch/llama/blob/main/LICENSE https://github.com/facebookresearch/llama/blob/main/LICENSE
- lantiga 4y agoCorrect, the original is GPL 3. To produce this implementation from the LLaMA paper we started from github.com/karpathy/nanoGPT, the LLaMA architecture is really similar to GPT. For instance we added rotary positional encoding starting from the original RoPE repo published with the paper. We finally ran the original model to make sure the two models were numerically.
- barefeg 4y agoBut aren’t the weights still not for commercial use?
- PostOnce 4y agoSo all the copyrighted stuff they trained it on is fair game, but the weights are not? This is legally uncharted territory I think. How much work would it be to create a "transformative" work from the llama weights, so that facebook has no claim?
- Ciantic 4y agoThat's what I thought too, the source code was not an issue so much as that. What we need is some sort of "Large Language Model at Home" (like SETI@home was) that could crowdsource the creation of the model which would be free to use.
- barefeg 4y agoRight, so sort of like https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals but for the training phase. I suppose different training runs could be proposed via a RFC type of procedure. Then it’s not only the open source model maintainers that put the effort, but also supporters of the project can “donate” their hardware resources.
- lantiga 4y agoSome form of that is very likely the future
- nynx 4y agoThere are already a million ways to run LLaMA. This doesn't change the issue at all, which is that the weights aren't commercially licensed.
- rasbt 4y agoI think some businesses and people are worried about using GPL code in their code bases because that's incompatible with their own licenses.
- theaniketmaurya 4y agoYes, agree that the weights aren't commercially licensed (yet)! The other ways to run LLaMA are using GPL license which makes it difficult for commercial use even if someone trains and upload the weights publicly. This could be a step in for the change :)
- oneshtein 4y agoPlease, train your model on some texts about copyright, licenses, open source licenses, GPL, because it produces nonsense. 1) GPL license is not an engine. It cannot run anything. 2) A product, produced by a code or machine or mechanically, is not copyrightable at all, or it has license of it original source. You can create GPL products on MS Windows using MS compiler. You can create proprietary products on Linux using GNU compiler. A binary output of a program, produced by a compiler such as GCC, is the derivative from it source code, so it has the same license as the source code. The compiler just transforms the text source file into the binary output file. For AI, the situation is the same: AI engine transforms a source data set into a binary file with weights, so binary weights is the derivative from source data set, thus its license is the same as in the original source set. When AI engine runs weights, it transforms an input query into output using also data from the source data set, thus creating a product with mixed content, which must obey licenses from both sources.
- alexb_ 4y ago>GPL...prevents meaningful academic and commercial use WTF are you talking about?
- theaniketmaurya 4y agoGPL is a copyleft license which requires you to share anything that you build using the original software. This makes it difficult for commercial use.
- n3t 4y ago> GPL is a copyleft license which requires you to share anything that you build using the original software. That's not true. > This makes it difficult for commercial use. Yeah, too bad it's so difficult for companies to use Linux commercially. /s
- homarp 4y agoRed Hat 30th anniversary - https://news.ycombinator.com/item?id=35337146 https://news.ycombinator.com/item?id=35337146
- adeon 4y agoI think implying that GPL is not "fully open source" is a hot take. It's specifically designed to ensure you and anyone you distribute your code gets the same freedoms. Maybe you don't agree that it's a good license but that is its intention. GPL vs BSD-type licenses I guess is decades long argument by now. Maybe I'm a naive idealist but IMO the GPL-family of licenses are underrated. You can use them to make sure you don't work for free for someone who won't share their improvements. I liked the choice of AGPL for AUTOMATIC1111 Stable Diffusion web UI. (https://github.com/AUTOMATIC1111/stable-diffusion-webui https://github.com/AUTOMATIC1111/stable-diffusion-webui) Commercial interests are very allergic to AGPL which ensures the project stays community-run and new features and fixes will prioritize the most ordinary user doing things for fun.
- cuuupid 4y agoI think OP mischaracterized the issue with the license, its more that the weights don’t fall under the same scope. They’re research use only, no commercial use allowed.
- rasbt 4y agoNot sure, but I think the point was that if you have something in GPL license (like the code in this case) it's open source, but that doesn't mean you can use that for your business application. That's because GPL requires you open sourcing all derivative work and most businesses don't want to/can't do that.
- lantiga 4y agoThe AI ecosystem is almost entirely Apache 2/MIT/BSD, and GPL is just incompatible with it. This is a blocker to mixing and matching, a simple Apache 2 rewrite fixes that problem. Weights? It’s another issue but we’ll be looking forward to fixing that too.
- dTal 4y agoHow is it incompatible? You can use code under all of those licenses in a GPLd work.
- ficiek 4y agoIf you hate GPL so much then I assume that you don't run any GPL licensed code on your machines then. I admire your resolve because I would think that is pretty hard!
- homarp 4y agollama.cpp is also MIT https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp previously discussed here https://news.ycombinator.com/item?id=35100086 https://news.ycombinator.com/item?id=35100086 and one of the rust wrapper: https://news.ycombinator.com/item?id=35171527 https://news.ycombinator.com/item?id=35171527 (also MIT)
- ipsum2 4y agoFYI, there's something fishy going on in this thread. Multiple people from the LightningAI team theaniketmaurya (developer advocate for Lightning AI) and rasbt (developer at Lightning AI) are shilling for this post without disclosing their affiliations. The account that submitted this (osurits) also only has two comments, also with the same behavior. Having interacted with the Lightning AI team in the past, this is unsurprising behavior.
- philipkglass 4y agoIf you suspect vote manipulation, email hn@ycombinator.com. Dang is good about replying to email and he has server-side logs available for more investigation.
- nl 4y agoJust noting that HuggingFace has a Llama code implementation[1]. It's also under an Apache 2 license. While this seems to be nice code I don't particularly see any reason to use that over HuggingFace transformers, where you can easily swap out alternative implementations. Also, going to legal restrictions on the Facebook LLama code when there are much stronger restrictions on the use of the model seems an odd thing to do. It's true that in some - not all - jurisdictions it is possible the model might not be copyrightable - but you'd have a bold legal department to rely on those arguments. It's also moderately likely that an instruction-tuned Llama (like Alpaca) would be copyrightable even in those jurisdictions. TL;DR: Use the HuggingFace transformers library. You can experiment with Llama and switch to truly free models like GPT-J or anything new that arrives very easily. [1] https://huggingface.co/docs/transformers/main/model_doc/llama https://huggingface.co/docs/transformers/main/model_doc/llam...
- blendergeek 4y ago> We believe that AI should be fully open source and part of the collective knowledge. As do I. > The original LLaMA code is GPL licensed which means any project using it must also be released under GPL. Yep. This ensures that AI is "fully open source and part of the collective knowledge." > This "taints" any other code and prevents meaningful academic and commercial use. Taints? As in "makes fully open source"? Isn't that the goal? > Lit-LLaMA solves that for good. Lit-LLaMA helps people create proprietary closed-source AI instead of the fully open source AI required by Llama. Okay.
- leke 4y agoI'm still confused about this. Does it require you to have a chatGPT API key for it to work?
- deleted 4y ago[deleted]
- javimh 4y agoNo, the GPL doesn't prevent meaningful academic or commercial use; rather, it seeks to prevent individuals from taking advantage of free software to limit the freedom of other users. It is important to note that if you live in a free country, there are laws that protect the liberties of all citizens and prevent actions that could restrict those freedoms.