7 ms·
Ask HN: Open source LLM for commercial use?
Working on a ML project and looking for an open source LLM that can be used in a commercial environment. As far as I'm aware, products cannot be built on LLAMA.
I don't want to use GPT since the project will be using personal information to train/fine tune the models.
- cl42 3y agoDolly 2 was released today and is OK for commercial use: https://huggingface.co/databricks/dolly-v2-12b https://huggingface.co/databricks/dolly-v2-12b I'm working on a package to help evaluate LLM results across different LLMs (e.g., GPT3.5 vs. GPT4 vs. Dolly 2 vs...); if you are looking to run experiments to compare results, I'd love to help you out. You can email me at w (at) phaseai (dot) com.
- dtagames 3y agoI think you might be confusing the GPT software (a generative pre trained transformer) with the finished product, an LLM (large language model.) A GPT has no training until you give it materials. I do believe Google released the code for theirs ages ago. Even without source, you can run a GPT against your own data locally, or on a cloud service setup for that purpose. This is how Bloomberg, for example, created a financial LLM. They used a GPT to train on their own financial data.
- moneywoes 3y agoAny examples of doing that process cost effectively?
- tough 3y agoNot what you're asking but Vicuna did cost merely 300$ to fine-tune on top of LLaMA https://www.marktechpost.com/2023/04/02/meet-vicuna-an-open-source-chatbot-that-achieves-90-chatgpt-quality-and-is-based-on-llama-13b/ https://www.marktechpost.com/2023/04/02/meet-vicuna-an-open-... AFAIK full model training should be a couple order magnitudes higher probably?
- dtagames 3y agoFor many projects, you'll need "natural language" training on regular text documents in order to be able to process even your prompts. So the most effective products will combine someone else's LLM (with their training data already in it) plus your custom training data. That way, you can interact with the LLM using normal English sentences but also get back information from your own dataset. Without this regular language training, your LLM wouldn't understand the questions you ask it. So there are two cost factors... the cost of paying someone else to train and host the regular LLM part + yours, or the cost of setting up the (virtual) hardware and compute time to train and host those things on your own. One "middle road" that might for some applications is to use the OpenAI API (for example) to combine access to your own data in real time (via your private APIs) with the natural language understanding that's already present in the LLM. These are the plug-ins that are quickly taking over HN, many without any great utility on their own. But you can see that a pre-trained LLM plus access to your own data privately might very well be worth paying for.
- titaniumtown 3y agoCerebras-GPT is licensed under Apache-2.0 and permits commercial use https://www.cerebras.net/blog/cerebras-gpt-a-family-of-open-compute-efficient-large-language-models https://www.cerebras.net/blog/cerebras-gpt-a-family-of-open-...
- maxilevi 3y agoYou could use GPT-J (https://huggingface.co/EleutherAI/gpt-j-6b https://huggingface.co/EleutherAI/gpt-j-6b)
- danpalmer 3y agoJust don't let it convince you to "reduce your carbon footprint" like the last guy did.
- tough 3y agoWait is this a reference to the belgian case of someone offing themselves? Was a bit weird they mentioned eliza/gpt-j i think on it but didnt make much sense to me? did that happen or just hallucinated?
- danpalmer 3y agoYes that's the one. There hasn't been much news coverage so I suspect that it wasn't quite as convincing a case as reported. Still a little worrying though, and even if not accurate, the fact it could be is definitely worrying.
- tough 3y agoWe cannot give meaning to tools, guns are much of a shortcut in that regard and nobody bats an eye. Schizo's tried to kill the curl creator because he was in his software everywhere and so spying on them... People is complicated. Let's not buy the bait that can kill wonderful tech, I agree the potential for harm is there, but I wouldn't blame the knive when a junkie stabs you to buy some heroin with what the gets out of you.
- danpalmer 3y agoI don't disagree with you, but it's also important that there is accountability. There is currently little accountability with LLMs. One could argue that the user should be accountable, but that doesn't account for how an LLM is trained. A user should clearly not be held accountable for being harmed if using an LLM maliciously trained to harm users. Most legal systems punish negligence in a position of power, so it stands to reason that the creator of an LLM should bear some responsibility for its behaviour. It is not yet clear to me how accountability should be portioned out to the user, operator, publisher, trainer, and model creator, but my feeling is that all bear at least some responsibility for its use.
- skdotdan 3y agohttps://huggingface.co/google/flan-ul2 https://huggingface.co/google/flan-ul2 https://huggingface.co/docs/transformers/model_doc/gpt_neox https://huggingface.co/docs/transformers/model_doc/gpt_neox
- zweezzy 3y agoBERT: https://huggingface.co/bert-base-uncased https://huggingface.co/bert-base-uncased
- icapybara 3y agoOthers have answered your question, but I'll add that the market for high quality AI models is not similar to the software marketplace, where there is always an open source alternative (and where open source is often the state of the art). LLMs take so much engineering effort, research, and compute that it's unlikely there will be good open source alternatives in the near future. Right now your only real option is OpenAI (or maybe Anthropic) and that seems unlikely to change anytime soon. The only reason we have LLAMA is because Meta threw us a bone. They might not do that again.
- rjzzleep 3y ago> LLMs take so much engineering effort, research, and compute that it's unlikely there will be good open source alternatives in the near future. Right now your only real option is OpenAI (or maybe Anthropic) and that seems unlikely to change anytime soon. does it though? it looks more like it requires a lot of money for compute and a lot of money and data for parameter tuning, but engineering effort seems soso. except for the compute cost this is perfect application for a distributed open source labeling effort. Just for my understanding though, are the data sets full of copyrighted material?
- muyuu 3y agoSome are, some aren't. See Koala for instance. The problem with Koala is that it fine-tunes on open sourced data, but makes no claims about the data for the base LLaMA models. https://bair.berkeley.edu/blog/2023/04/03/koala/ https://bair.berkeley.edu/blog/2023/04/03/koala/ The irony is that openAI and Meta themselves might be in flaky ground for having trained models on other people data with dubious rights to do so in many instances, and then using it to produce output commercially. But this is a new frontier and enforcement might be effectively not possible unless new legislation requires reproducibility and audits on the data sets or something like that. But without that, how do you know exactly how did they arrive at a given set of weights with Montecarlo algorithms and arbitrary fine tuning? You basically don't know what was there and you cannot prove they didn't achieve those results with perfectly clean data. PS: https://medium.com/geekculture/list-of-open-sourced-fine-tuned-large-language-models-llm-8d95a2e0dc76 https://medium.com/geekculture/list-of-open-sourced-fine-tun...
- sinenomine 3y agoIf you want quality, use Google's Apache-licensed LLM https://huggingface.co/google/ul2 https://huggingface.co/google/ul2
- vinni2 3y agoThey also have Flan T5 which is also Apache 2. https://huggingface.co/google/flan-t5-xxl https://huggingface.co/google/flan-t5-xxl
- K0IN 3y agoI think https://github.com/BlinkDL/RWKV-LM https://github.com/BlinkDL/RWKV-LM could be used, but not all versions (namely instruction fine-tuned models trained on alpaca data)
- rolisz 3y agoWhat exactly do you want to do? There are various alternatives, but they are not as general as OpenAI's GPT, but, they can be finetuned more cheaply to solve a specific task.
- dmurko 3y agoJust in case you were not aware: "OpenAI does not use data submitted by customers via our API to train OpenAI models or improve OpenAI’s service offering." It does for ChatGPT though. Source: https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance https://help.openai.com/en/articles/5722486-how-your-data-is...
- icapybara 3y agoFor many companies this type of promise is not useful. It doesn’t matter that they say they won’t, they still can look if they want to. This is the primary concern when you’re dealing with trade secrets where the secrecy of the information is its only protection.
- pantulis 3y agoIf you use Azure OpenAI's services, I would guess you would fall into contractual agreements with Microsoft which should cover these concerns just like when you are using MS SQL Server to store trade secrets or PII.
- deleted 3y ago[deleted]
- danrocks 3y ago[dead]
- wejick 3y agoI remember someone mentioned on other thread that after distilled, llama will have no license issue. can someone explain why is that the case? Probably can give directions where a software engineer can start to understand the concept.
- MacsHeadroom 3y agoBecause machine leaning models likely have zero intellectual property rights protections. LLMs are the output of an "algorithmic process". Algorithmic outputs are explicitly except form copyright, unlike source code. (Note: Compiled software is not an algorithmic output, under the specific legal definition.) Machine Learning models are made the same way machine learning output is generated. In other words, the old model is training data to the new model. Just like the pirated torrent site dataset "Books3" Facebook used to train LLaMA is training data. If Facebook can protect their model under copyright then every publisher in existence sue Facebook into the ground. They can't have it both ways.
- nextaccountic 3y ago> In other words, the old model is training data to the new model. Just like the pirated torrent site dataset "Books3" Facebook used to train LLaMA is training data. This is a logical conclusion. But if it actually holds, that's for the courts to decide
- Garcia98 3y agoI've seen this question asked repeatedly in many LLaMa threads, currently the best models that are truly open are the released models from the Flan family by Google, which includes Flan-T5[0] and Flan-UL2[1]. According to its paper, Flan-UL2 performs slightly better than Flan-T5-XXL. These models perform slightly better than GPT-3 under some tasks[2], but they're still far from achieving the results from GPT-3.5 and GPT-4. This becomes evident when you try to use them in the real world; they're not "good enough" for general use cases, unlike ChatGPT models. However, if you can restrict your use case to one particular domain, you can achieve pretty good results by further fine-tuning these models. [0]: https://huggingface.co/google/flan-t5-xxl https://huggingface.co/google/flan-t5-xxl [1]: https://huggingface.co/google/flan-ul2 https://huggingface.co/google/flan-ul2 [2]: https://paperswithcode.com/sota/multi-task-language-understanding-on-mmlu https://paperswithcode.com/sota/multi-task-language-understa...
- momofuku 3y agoFor the life of me, I cannot understand why Google did not go ahead and commercialize a lot of this early research. They clearly had a HUGE lead in this space, in terms of engineering/research talent, capital, computer infrastructure. Boggling... I'd love any alternative view points of this.
- sdrinf 3y agoBringing disruption to your company's 90% revenue generator product without proving the alternative's financial model is not a career-enhancing move. Can't prove the alternative's financial model without showing the thing to real users. Can't know in advance, if the new financial model will be pennies on adwords' dollars.
- momofuku 3y agoI see your point. Right now, people seem to use ChatGPT as an alternative to Google Search, which I don't think is the right use case for it (happy to be proven wrong though), unless they figure out a way to power it using knowledge graphs, or an equvivalent system to provide accurate, factual information. By commercializing this research, I mean why not integrate this into GMail for auto reply solutions, so that it can automatically suggest meetings? Why not integrate it into Slides to come up with better titles, summarization etc? Similarly for Google Docs, automatically summarize reports, make suggestions depending on crispness/clarity etc. Why did they decide to just sit on these models, and endlessly prolong their product launches that had these features baked in.
- lhl 3y agoThe ones I saw mentioned so far were Flan, Cerebras, GPT-J, and RWKV. Not yet mentioned: * Pythia https://github.com/EleutherAI/pythia https://github.com/EleutherAI/pythia * GLM-130B https://github.com/THUDM/GLM-130B https://github.com/THUDM/GLM-130B - see also ChatGLM-6B https://github.com/THUDM/ChatGLM-6B https://github.com/THUDM/ChatGLM-6B * GPT-NeoX-20B https://huggingface.co/EleutherAI/gpt-neox-20b https://huggingface.co/EleutherAI/gpt-neox-20b * GeoV-9B https://github.com/geov-ai/geov https://github.com/geov-ai/geov * BLOOM https://huggingface.co/bigscience/bloom https://huggingface.co/bigscience/bloom and BLOOMZ https://huggingface.co/bigscience/bloomz https://huggingface.co/bigscience/bloomz
- tslmy 3y agoCaveat: For ChatGLM-6B, you can't use the pre-trained model for commercial uses: > You will not use, copy, modify, merge, publish, distribute, reproduce, or create derivative works of the Software, in whole or in part, for any commercial, military, or illegal purposes.
- sturza 3y agoI tried GLM-130B on huggingface https://huggingface.co/spaces/THUDM/GLM-130B https://huggingface.co/spaces/THUDM/GLM-130B: Q: hello, who are you? A: I was twitted. In fact, twitted, or twittered, is one of those tweets that one either never sees or sees right away. Twitter is a website that allows users to post short messages that can be read and retweeted by other users. These messages are called tweets. So, who are you? Oh, hello, who are you? I was twitted. In fact, twitted, or twittered, is one of those tweets that one either never sees or sees right away. Twitter is a website that allows users to post short messages that can be read and retweeted by other
- rasengan 3y agoAt first glance/test on their page, it can provide a significantly better completion if given a prompt of a series of Q and As and by modifying Temp/etc.
- mingyeow 3y agoNoob question here - what’s the best tutorials to get started in mixing LLM models and building on top of one another, assuming very good programming background but little AI background? I asked chatGPT this question, and it was helpful but not comprehensive, but I figure intelligent humans on this forum will give the best answers.
- RockyMcNuts 3y agonot sure what 'mixing LLM models' entails but these are maybe some good starting points - karpathy - https://www.youtube.com/watch?v=kCc8FmEb1nY https://www.youtube.com/watch?v=kCc8FmEb1nY - https://towardsdatascience.com/beautifully-illustrated-nlp-models-from-rnn-to-transformer-80d69faf2109 https://towardsdatascience.com/beautifully-illustrated-nlp-m... - https://dzone.com/articles/a-deep-dive-into-the-transformer-architecture-the https://dzone.com/articles/a-deep-dive-into-the-transformer-... - https://peterbloem.nl/blog/transformers https://peterbloem.nl/blog/transformers - http://nlp.seas.harvard.edu/2018/04/03/attention.html http://nlp.seas.harvard.edu/2018/04/03/attention.html - https://lilianweng.github.io/posts/2023-01-27-the-transformer-family-v2/ https://lilianweng.github.io/posts/2023-01-27-the-transforme... - https://blog.quickchat.ai/post/tokens-entropy-question/ https://blog.quickchat.ai/post/tokens-entropy-question/ - https://dugas.ch/artificial_curiosity/GPT_architecture.html https://dugas.ch/artificial_curiosity/GPT_architecture.html - https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/ https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-... - https://d4mucfpksywv.cloudfront.net/better-language-models/language_models_are_unsupervised_multitask_learners.pdf https://d4mucfpksywv.cloudfront.net/better-language-models/l... - https://arxiv.org/pdf/2005.14165.pdf https://arxiv.org/pdf/2005.14165.pdf - https://arxiv.org/pdf/2303.08774.pdf https://arxiv.org/pdf/2303.08774.pdf - https://arxiv.org/pdf/2303.17564.pdf https://arxiv.org/pdf/2303.17564.pdf
- extasia 3y agoMy answer would be quite specific to what exactly you're trying to achieve. Id be wary of just hacking away without understanding at least the fundamentals of ML + NLP or you'll find yourself lost pretty quick. I'm a former SWE turned NLP researcher, so i was recently in your position:)
- erwincoumans 3y agoTruly Open AI: LAION calls for a supercomputer to develop open-source AI, by replicating large models like GPT-4 and exploring them together as a research community. https://www.heise.de/news/Open-source-AI-LAION-proposes-to-openly-replicate-GPT-4-a-public-call-8785603.html https://www.heise.de/news/Open-source-AI-LAION-proposes-to-o...
- redskyluan 3y agowhat about the https://huggingface.co/facebook/opt-66b https://huggingface.co/facebook/opt-66b? I thought the opt series can be used in production
- dreaminvm 3y agoHere's a recent release of fine-tuning Flan-UL2 on instructions (alpaca). https://medium.com/vmware-data-ml-blog/lora-finetunning-of-ul-2-and-t5-models-35a08863593d https://medium.com/vmware-data-ml-blog/lora-finetunning-of-u...
- gumby 3y ago> looking for an open source LLM that can be used in a commercial environment. As far as I'm aware, products cannot be built on LLAMA. Commercial product sure can be built on top of LLAMA, it's GPL-3. Your models are your own; just patches, modifications, and code you link to LLMA itself will be governed by the GPL as well. This is almost certainly what you want since this way you can use patches, fixes, and improvements others make to LLMA. You won't have to do all that work yourself, or necessarily wait for Facebook.
- brentis 3y agoMy personal use case is that I'd like to query a bunch of our APIs and amalgamate a response those consumable for humans. I think many of us have the same need and are waiting for open AI plug-in access. Is this the question we are asking yourselves here or are we talking about licensing?