4 ms·
This implementation is much faster on my M5 Max, like a few minutes for the same video, but on an M5 Max with 128GB, didn't test on M5 Pro. About memory, could
by antirez 2mo ago
This implementation is much faster on my M5 Max, like a few minutes for the same video, but on an M5 Max with 128GB, didn't test on M5 Pro. About memory, could be executed on 64GB with a few changes.
- Manfrednotfunny 2mo agoMemory bandwidtih between pro and max is double. 300gb/s vs. 600gb/s btw.
- antirez 2mo agoDoes not matter much in this case. GPU bound.
- Manfrednotfunny 2mo agoSeems to be true, but also seems hard t obenchmark with max having more GPU cores too.
- dragonwriter 2mo agoMy understanding is that that tends to be more critical with LLMs than image/video gen models, which are relatively more compute vs. memory transfer intensive than LLMs
- zozbot234 2mo agoPerformance might still end up being bounded by data transfer speed if SSD streaming is heavily used to make up for limited RAM. By comparison, it doesn't take many parallel-batched sessions to make LLM decode compute-bound on typical hardware (hence seeing very limited gains from even wider batching), but this just doesn't apply when streaming weights from disk, the setting is completely different.