3 ms·
I've been using MiniMax H3 on my M5 Pro 64GB MacBook Pro through ComfyUI. It works extremely well. I had to modify the default ComfyUI workflows to use a GGUF
by Meleagris 2mo ago
I've been using MiniMax H3 on my M5 Pro 64GB MacBook Pro through ComfyUI. It works extremely well.
I had to modify the default ComfyUI workflows to use a GGUF quant (city96's ComfyUI-GGUF custom node, UnetLoaderGGUF in place of the stock loader) [0].
I use the model labeled Q5_K_M. There is Q8_0 available as well, which is 34GB and fits fine in 64GB unified memory if you keep resolution modest.
The main issue is speed, a ~9-second 480x864 clip at 20 steps takes me a bit over an hour. So this will be cool to try for the speed up alone.
There's a lot of great information and workflows available to follow on the r/StableDiffusion subreddit.
[0] https://huggingface.co/Abiray/MiniMax-H3-GGUF/tree/main/unet https://huggingface.co/Abiray/MiniMax-H3-GGUF/tree/main/unet
- alexgoodhart 2mo agoI wonder how much faster your m5 pro is compared to my M1 Max @ 64gb
- yieldcrv 2mo agowait till the M7 bro you’re almost there, rumor has it that Apple is skipping the M6 but it still might be 2028
- ignoramous 2mo agoDiscussion: https://news.ycombinator.com/item?id=48676795 https://news.ycombinator.com/item?id=48676795
- jonplackett 2mo agoWhat is the quality of the output like compared to something like Veo?
- Art9681 2mo agoIt's better than all private video models like Veo. Yes, I too am incredulous they released this open weights. It's a VERY disruptive model in all the good ways.
- antirez 2mo agoThis implementation is much faster on my M5 Max, like a few minutes for the same video, but on an M5 Max with 128GB, didn't test on M5 Pro. About memory, could be executed on 64GB with a few changes.
- Manfrednotfunny 2mo agoMemory bandwidtih between pro and max is double. 300gb/s vs. 600gb/s btw.
- antirez 2mo agoDoes not matter much in this case. GPU bound.
- Manfrednotfunny 2mo agoSeems to be true, but also seems hard t obenchmark with max having more GPU cores too.
- dragonwriter 2mo agoMy understanding is that that tends to be more critical with LLMs than image/video gen models, which are relatively more compute vs. memory transfer intensive than LLMs
- zozbot234 2mo agoPerformance might still end up being bounded by data transfer speed if SSD streaming is heavily used to make up for limited RAM. By comparison, it doesn't take many parallel-batched sessions to make LLM decode compute-bound on typical hardware (hence seeing very limited gains from even wider batching), but this just doesn't apply when streaming weights from disk, the setting is completely different.
- Myzura 2mo agoHow much free space do you have left after running this llm model? Have you tried to develop your own model with the M5?
- thousand_nights 2mo ago> a ~9-second 480x864 clip at 20 steps takes me a bit over an hour that's rough. for comparison, i tried the exact same parameters on my 5090 RTX and it took 2 minutes to generate. i believe diffusion models are primarily compute bound so the macs aren't really the ideal hardware for this kind of stuff
- vimto 2mo agoGGUF is outdated in the latest versions of Comfy-UI. If you want a good balance of size, speed and quality you should use the int8_convrot model from the official Comfy Org Repo https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffus...
- Meleagris 2mo agoSo I did test this, and it doesn't work because the quantized layers need torch._int_mm, which PyTorch's MPS backend doesn't implement. It just throws NotImplementedError.
- SV_BubbleTime 2mo agoThis is good advice if you have nvidia, but for Mac does not apply currently.
- dragonwriter 2mo agoGGUF is unsupported by ComfyUI’s memory management system that enables running models much larger than fit in VRAM with tolerable efficiency via weight streaming, but for unified memory systems that system is less relevant (unless using models too big to run in unified memory AND having fast enough mass storage to benefit from direct-from-disk weight streaming.)
- deleted 2mo ago[deleted]