3 ms·
FYI, vLLM also just added experimental multi-lora support: https://github.com/vllm-project/vllm/releases/tag/v0.3.0 https://github.com/vllm-project/vllm/release
by semmulder 3y ago
FYI, vLLM also just added experimental multi-lora support: https://github.com/vllm-project/vllm/releases/tag/v0.3.0 https://github.com/vllm-project/vllm/releases/tag/v0.3.0
Also check out the new prefix caching, I see huge potential for batch processing purposes there!
- brucethemoose2 3y agoMissed this, thanks. Everything is moving so fast!
- tgaddair 3y agoYes, we (LoRAX devs) saw that (we know the author pretty well). It's a useful addition, though quite a bit simpler than our level of support for multi-LoRA inference. We're planning on doing a more comprehensive comparison soon, now that it's officially out. I will say that if you want to explore the forefront of this multi-LoRA inference, definitely worth giving LoRAX a look. We just added support for per-request model merging (https://predibase.github.io/lorax/guides/merging_adapters/ https://predibase.github.io/lorax/guides/merging_adapters/) as an example, and are planning on continuing to double down on this idea of combining adapters in some pretty unique ways.