4 ms·
I don't have the arxiv link bookmarked, but there was a paper written on pretraining with Lora. It involved merging adapters back every n steps with good result
by bradfox2 2y ago
I don't have the arxiv link bookmarked, but there was a paper written on pretraining with Lora. It involved merging adapters back every n steps with good results.