17 ms·
Launch Lamini: The LLM Engine for Rapidly Customizing Models as Good as ChatGPT
- sharonzhou 3y agoHi HN! I’m super excited to announce Lamini, the LLM engine that gives every developer the superpowers that took the world from GPT-3 to ChatGPT! I’ve seen a lot of developers get stuck after prompt-tuning for a couple days or after fine-tuning an LLM and it just gets worse—there’s no good way to debug it. I have a PhD in AI from Stanford, and don’t think anyone should need one to build an LLM as good as ChatGPT. A world full of LLMs as different & diverse as people would be even more creative, productive, and inspiring. That’s why I’m building Lamini, the LLM engine for developers to rapidly customize models from amazing foundation models from a ton of institutions: OpenAI, EleutherAI, Cerebras, Databricks, HuggingFace, Meta, and more. Here’s our blog announcing us and a few special open-source features! https://lamini.ai/blog/introducing-lamini https://lamini.ai/blog/introducing-lamini Here’s what Lamini does for you: Your LLM outperforms general-purpose models on your specific use case You own the model, weights and all, not us (if foundation model allows it, of course!) Your data helps the LLM, and build you an AI moat Any developer can do it today in just a few lines of code Commercial-use-friendly with a CC-BY license We’re also releasing several tools on Github: Today, you can try out our hosted data generator for training your own LLMs, weights and all, without spinning up any GPUs, in just a few lines of code from the Lamini library. https://github.com/lamini-ai/lamini/ https://github.com/lamini-ai/lamini/ You can play with an open-source LLM, trained on generated data using Lamini. https://huggingface.co/spaces/lamini/instruct-playground https://huggingface.co/spaces/lamini/instruct-playground Sign up for early access to the training module that took the generated data and trained it into this LLM, including enterprise features like virtual private cloud (VPC) deployments. https://lamini.ai/contact https://lamini.ai/contact
- jasonjmcghee 3y agoYou’re building some seriously exciting stuff! Looking forward to diving in.
- jeffybefffy519 3y agoIm confused, what are you actually offering? Does my fine tuning data get shared with your platform’? Does the model get fine tuned on your end or my own system? Do you host the model?
- gdiamos 3y agoNoting that the Github repo includes a data pipeline for instruction fine tunining. What's the difference between this and other data pipelines like Alpaca?
- verdverm 3y agoAren't you Greg Diamos, the founder, why are you asking this instead of answering?
- acapybara 3y agoForgot to switch to sock puppet account.
- gdiamos 3y agoThis was a frequently asked question among my friends. I’m really curious to see how someone who hasn’t been staring at the docs for weeks would explain it.
- verdverm 3y agoTo HN, this looks like faking engagement, which is against the posting guidelines. This is a question you should instead ask in a user interview, or at a minimum qualify when asking here as one of the people involved in the project.
- eschluntz 3y agoVery exciting! Glad to finally be able to get beyond prompt engineering. What's the pricing model like?
- gdiamos 3y agoFree open source libraries. Paid LLM hosting. 50% cheaper than OpenAI, pay per compute needed to run & create the LLM. Export the weights anytime you want. Enterprise VPC deployments.
- deleted 3y ago[deleted]
- atulika612 3y agoIf I want to export the model and run it myself, can I do that?
- primordialsoup 3y agoCongrats! I went to your demo and asked for words that end in agi. This is what I got: -- agi, agi, agi, agi, agi, agi, agi These are some of the words that end in agi. You can also use the word agi in a sentence. For example, "I am going to the grocery store to get some agi." These are some of words that end in agi. These are some words that end in agi. maximize, maximize, maximize, maximize, maximize, maximize, maximize, maximize These are some words that ends in agi -- So I think this needs more work to get to "as good as ChatGPT". But having said that, congrats on the landing
- avereveard 3y agoyeah as usual these model can barely sustain a conversation and fall apart the moment actual instructions are given. typical prompt they fail to udnerstand: "what is pistacchio? explain the question, not the answer." all these toy llm: "pistacchio is..." gpt is the only one that consistently understand these instructions: "The question "what is pistachio?" is asking for an explanation or description of the food item..." this makes these llm basically useless for obtaining anything but hallucinated data.
- vidarh 3y agoIt only makes them useless.of you insist on asking them in ways you already know will provide bad results instead of adapting your prompts. This is a bit like complaining that your compiler refuses to produce the right outputs for code you've already determined is incorrect.
- avereveard 3y agoAsking LLM from things they learned in training mostly result in hallucinations and in general makes you unable to detect by which amount they are hallucinating: these models are unable to reflect on their output, and average output token probability is a lousy proxy for confidence scoring their results. On the other hand, no amount of prompt engineering seems to make these LLM able to do question and answer over source documents which is the only realistic way by which factual information can be retrieved You're welcome to bring examples of it tho if you're so confident.
- furyofantares 3y agoThe actual post doesn't say "as Good as ChatGPT", why does the HN title? I don't really care to click on something I know is obviously lying to me.
- ec109685 3y agoThis headline is totally editorializing. Stick with the source one. “Introducing Lamini, the LLM Engine for Rapidly Customizing Models” So much click bait in the LLM space.
- deleted 3y ago[deleted]
- mkl 3y agoIs it still editorialising when OP is the CEO of the company?
- baobabKoodaa 3y agoYes
- iguana 3y agoTrivial examples show that this isn't nearly as good as ChatGPT. The headline should be changed.
- mensetmanusman 3y agoIt took Open.ai 10 years of fine tuning, can’t expect things to work as well in day 1.
- gdiamos 3y agoI like blog post title. Introducing Lamini, the LLM Engine for Rapidly Customizing Models Obviously it still takes a huge amount of work to customize a model to be as good as GPT4 or ChatGPT, that’s exactly why we are building Lamini. To give developers tools to make it easier. Hopefully it is clear that it will take more work than 1 day.
- cultofmetatron 3y agoI hope this turns out be as good as chatgpt and not "we have chatgpt at home"
- acapybara 3y agoTry SuperCOT 30B.
- 8thcross 3y agolooks great...looking forward to trying it out
- batch12 3y agoI've been playing a bit with stacking transformer adapters to add knowledge to models and so far it has met my needs. It doesn't have the same illusion of intelligence, but so far it's just as good as a multitasking intern, so I am still having fun with it. I wonder if this is basically doing the same thing.
- leobg 3y agoInteresting. Do you know if this can be done with Sentence Transformers, too? Picking a good performing one from HF. Then training an adapter for the domain (unsupervised). Then adding another one using actual training triplets (base, similar, non-similar)?
- batch12 3y agoI haven't done this with sentence transformers but I imagine it's possible since they can be loaded as regular transformers. Check out https://github.com/huggingface/peft https://github.com/huggingface/peft -- they've packaged it up nicely- and read up on LoRA (https://arxiv.org/pdf/2106.09685.pdf https://arxiv.org/pdf/2106.09685.pdf) That should get you started.
- leobg 3y agoThank you. Peft and adapters seem to be two different things though, no? AFAIK there are other libraries for adapters (forgot the name). Is peft what you were talking about when you said adapters in your original comment?
- batch12 3y agoI was of the understanding that Lora was one flavor of adapters, but I am still learning so I may be wrong. I yet gotten too deep into other transformer adapters yet (still reading).
- mise_en_place 3y agoWhy wouldn’t we use something like DeepSpeed? It’s a one-click on Azure. What’s the value add?
- digitcatphd 3y agoGPT at this point is more than an LLM, it is a baseline layer of logic using the underlying transformer technology. This will be challenging to replicate without the same size of data sets