4 ms·
The VRAM usage is closer to a 47B model - although only 2 experts are used at a time for inference, all experts are needed to complete it.
by sanjiwatsuki 3y ago
The VRAM usage is closer to a 47B model - although only 2 experts are used at a time for inference, all experts are needed to complete it.
- discordance 3y agoConfirmed. Currently running Mixtral 8x7B gguf (Q8_0) on a Macbook Pro M1 Max w 64GB ram, and RAM usage is sitting at 48.8 GB.
- karolist 3y agoHow many t/s?
- discordance 3y agoAround 15 - 20 t/s
- karolist 3y agoThank you, got the same build M1 Max during Christmas B&H sale and can confirm it's amazing for running local LLMs.