3 ms·
Issue is that llama.cpp is the best way to run models on hardware that isn't nvidias.
by noosphr 1mo ago
Issue is that llama.cpp is the best way to run models on hardware that isn't nvidias.
- alightsoul 1mo agoExcept when they have less than 16 gb of ram?
- dannyw 1mo agoA lot of llama.cpp contributions come from the community and ecosystem, like Unsloth. If something goes awry, I fully expect lots of forks.
- MrDrMcCoy 29d agoThere already are a lot of forks for things they decline to implement. TurboQuant, ROCmFPX, and more. I need to set up an agent that will loop on merging them.