3 ms·
How's the Strix Halo? I'd really like to get a local inference machine so that I don't have to use quantized versions of local models.
by dimgl 8mo ago
How's the Strix Halo? I'd really like to get a local inference machine so that I don't have to use quantized versions of local models.
- Tepix 8mo agoPrompt preprocessing is slow, the rest is pretty great.
- evilduck 8mo agoWorks great for these type of MOE models. The ability to have large amounts of VRAM let you run different models in parallel easily, or to have actually useful context sizes. Dense models can get sluggish though. AMD's ROCm support has been a little rough for Stable Diffusion stuff (memory issues leading to application stability problems) but it's worked well with LLMs, as does Vulkan. I wish AMD would get around to adding NPU support in Linux for it though, it has more potential that could be unlocked.