27 ms·
Qlora won't work well to add knowledge via private data. Parameter efficient methods are not useful for these cases at the 8b scale without a more complex trai
by bradfox2 2y ago
Qlora won't work well to add knowledge via private data.
Parameter efficient methods are not useful for these cases at the 8b scale without a more complex training procedure that periodically merges back adapters. Maybe at the 70B scale.
- tpurves 2y agoWhat scale of company do you need to be to actually be able afford and get return on investment on retraining base models with your own proprietary knowledge and docs? Considering also the implications of continually retraining?
- sdesol 2y agoI was under the impression that you wouldn't. If you want access to proprietary knowledge, you would use RAG + LLM.
- littlestymaar 2y agoI don't think anyone has the answer to this question yet.
- bradfox2 2y agoThe only experience I have is first hand, what my company is doing for our client base. We are doing continuous pretraining and the rest of the alignment stack training on about 10B private tokens + private customer data to produce private custom models for companies in the 500 to 3000 employee range. We built and operate a single rack cluster that cost mid 6 figures in order to be able to do this. These models get combined with rag for highly specific technical doc authoring and other uses.
- tpurves 2y agoThis is very helpful context on what works right now, thanks for sharing.
- anonymousDan 2y agoCan you point to any literature on this by any chance? I would be really interested to see some in depth analysis.
- bradfox2 2y agoI don't have the arxiv link bookmarked, but there was a paper written on pretraining with Lora. It involved merging adapters back every n steps with good results.