3 ms·
Not directly related to K2, but why do a lot of the newly released models basically say day zero day support in vllm, slang but often not llama.cpp? Llama.cpp
by SillyUsername 1mo ago
Not directly related to K2, but why do a lot of the newly released models basically say day zero day support in vllm, slang but often not llama.cpp?
Llama.cpp is then often a few days behind, which given it's the only inference engine supporting older architectures is quite frustrating.
- walrus01 1mo agoDevelopers with lots of VC money to burn are working on things like B100/B200/B300 which are well supported in VLLM, everything else in terms of supporting more mundane GPUs or other platforms is ancillary to the main task of getting the thing trained and aligned.