5 ms·
If you are doing nothing but consuming models via llama.cpp, is the AMD chip an obstacle? Or is that more a problem for research/training where every CUDA featu
by 3eb7988a1663 1y ago
If you are doing nothing but consuming models via llama.cpp, is the AMD chip an obstacle? Or is that more a problem for research/training where every CUDA feature needs to be present?
- _lvbh 1y agoLlama.cpp works well on AMD, even for really outdated GPUs. Ollama refuses to work with my RX 570 from 2019 but llama.cpp supports it via Vulkan.
- Havoc 1y ago>Ollama refuses to work with my RX 570 from 2019 but llama.cpp supports it via Vulkan. That's a bit odd given ollama utilizing llama to do the inference...
- LorenDB 1y agoSee recent discussion about this very topic: https://news.ycombinator.com/item?id=42886680 https://news.ycombinator.com/item?id=42886680
- DiabloD3 1y agoIt isn't odd at all. Ollama uses an ancient version of llama.cpp, and was originally meant to just be a GUI frontend. They forked, and then never resynchronized... and now lack the willpower and technical skill to achieve that. Ollama is essentially a dead, yet semi-popular, project with a really good PR team. If you really want to do it right, you use llama.cpp.
- washadjeffmad 1y agoDon't you dare say anything unpositive about Ollama this close to whatever it is they're planning to distinguish themselves from llama.cpp. They've been out hustling, handshaking, dealmaking, and big businessing their butts off, whether or not they clearly indicate the shoulders of the titans like Georgi Gerganov they're wrapping, and you are NO ONE to stand in their way. Do NOT blow this for them. Understand? They've scooted under the radar successfully this far, and they will absolutely lose their shit if one more peon shrugs at how little they contribute upstream for what they've taken that could have gone to supporting their originator. Ollama supports its own implementation of ggml, btw. gglm is a mysterious format that no one knows the origins of, which is all the more reason to support Ollama, imo.
- DiabloD3 1y agoMan, best /s text I've seen on here in awhile. I hope other people appreciate it.
- DiabloD3 1y agoI don't bother with Nvidia products anymore. In a lot of ways, they're too little too late. Nvidia products generally perform worse per dollar, perform worse per watt. In a single GPU situation, my 7900XTX has gotten me farther than a 4080 would have, and matches the performance I expect from a 4090 for $600 less, and also 50-100w less. Now, if you're buying used hardware, yeah, go buy used, not new high-VRAM Nvidia models, the ones with 80+GB. You can't buy those used from AMD customers yet, as they're happily holding onto them; they perform so well, the need to upgrade isn't happening yet.
- mdp2021 1y ago> my 7900XTX has gotten me farther than a 4080 would have But is the absence of CUDA a constraint? Do neural networks work "out of the box"? How much of a hassle (if at all) is it to make things work? Do you meet incompatible software?
- DiabloD3 1y agollama.cpp is the SOTA inference engine that everyone in the know uses, and has a Vulkan backend. Most software in the world is Vulkan, not CUDA, and CUDA only works on a minority of hardware. Not only that, AMD has a compatibility layer for CUDA, called HIP, part of the ROCm suite of legacy compatibility APIs, that isn't the most optimal in the world but gets me most of the performance I would expect from a similar Nvidia product. Most software in the world (not just machine learning related stuff) is written in an API that is cross-compatible (OpenGL, OpenCL, Vulkan, Direct family APIs). Nvidia continually sending a message of "use CUDA" really means "we suck at standards compliance, and we're not good at the APIs most software is written in"; since everyone has realized the emperor wears no clothes, they've been backing off on that, and are slowly improving their standards compliance for other APIs; eventually, you won't need the crutch of CUDA, and you shouldn't be writing software today in it. Nvidia has a bad habit of just dropping things without warning when they're done with them, don't be an Nvidia victim. Even if you buy their hardware, buying new hardware is easy: rewriting away from CUDA isn't (although, certainly doable, especially with AMD's HIP to help you). Just don't write CUDA today, and you're golden.