6 ms·
There is a more detailed explanation at https://unsloth.ai/introducing https://unsloth.ai/introducing
by TheGeminon 3y ago
There is a more detailed explanation at https://unsloth.ai/introducing https://unsloth.ai/introducing
- apsec112 3y agoThat... doesn't really explain how they can get such a high number? Standard FLOP efficiency on fine-tuning big models is like 30-40%. How can you get 750%?
- danielhanchen 3y agoHey! Great question! That's what I'm confused about as well! So in GPUs the goal is to saturate the GPU with matrix multiplies instead of data movement. I'll write a more detailed blog but approximately: 1. Flash Attention v2 reduces the time taken by 17% or so 2. RoPE Triton kernels: -7.1% 3. RMS Layernorm in Triton: -3.1% 4. Cross Entropy in Triton: -1% 5. Manual autograd for MLP: -4% 6. Manual QKV autograd: -2% 7. Manual O autograd: -2% 8. Smart cache evictions and reduced data duplications etc: -30% 9. And other tricks in the Max and Pro versions makes it 30x faster You can see it's just tricks in each step, which accumulate together to make to go faster. I'll write up a blog post to detail it all in the future!!!
- demosthanos 3y ago> And other tricks in the Max and Pro versions makes it 30x faster This feels like the collecting underpants meme. Phase 1: Get to the same performance as other methods. Phase 2: ???. Phase 3: Now you're at 750%! You may or may not actually have succeeded at what you claim to, but you're not being very persuasive. I realize that you're trying to turn these tricks into a profit and revealing them would destroy that possibility, but you're going to have a really hard time persuading people to pay for a product that does something that enormous teams of PhDs at BigTech haven't been able to pull off on the basis of "trust me".
- danielhanchen 3y agoI agree fully - what do you suggest then? OSS the entire code base and using AGPL3? I tried that with https://github.com/danielhanchen/hyperlearn https://github.com/danielhanchen/hyperlearn to no avail - we couldn't even monetize it at all, so I just OSSed everything. I listed all the research articles and methods in Hyperlearn which in the end were gobbled up by other packages. We still have to cover life expenses and stuff sadly as a startup. Do you have any suggestions how we could go about this? We thought maybe an actual training / inference platform, and not even OSSing any code, but we decided against this, so we OSSed some code. Any suggestions are welcome!
- wsxiaoys 3y agoWow, this is a great topic. I don't really have specific suggestions, but I'd like to contribute some thoughts on the matter. Monetizing anything isn't inherently problematic; the challenge lies in defining what should be paid for and what should be offered for free. In the realm of open-source products and SaaS, the common practice is to provide free self-hosting options while charging for cloud hosting or enterprise-specific features, such as access control and authentication integrations. However, the landscape becomes significantly more challenging for LLMOps (assuming you are still focusing on training as a major aspect of your business, which can be categorized as LLMOps). Historically, there haven't been many success stories in this area (with exceptions like wand.ai, which focusing on tracking experiments). I believe this difficulty arises from the largely ad-hoc nature of training and fine-tuning processes, making standardization a challenge, coupled with the infrequency of these tasks. That being said, training/finetuning is a valuable technique. However, transforming it into a company that offers products is really challenging. Successful examples in this realm typically depend heavily on solution customization or consulting-oriented business models.
- danielhanchen 3y agoThanks for the points! I agree monetization in the LLM Ops space is hard and complex. Agreed fully on customizing solutions or consulting. Yep self hosting solutions like Redhat, or DBs like MongoDB or Gitlab's dashboard style approach could work - the issue is now as you mentioned we offer training and finetuning. We do plan to offer inference as well, plus the data gathering process, and the final prompt engineering side - but we thought why not have a shot? It's possible best to make a training and inference platform - maybe some sort of personal ChatGPT training for the public - everyone can train their own personal ChatGPT not via ChatGPT's in context learning or RAG, but coupled with actual fast 30x finetuning, a personal bot can truly be possible. Thaks for the suggestions!
- deleted 3y ago[deleted]