3 ms·
For those adventurous souls who have gobs of memory, a good GPU, and plenty of disk space, some of the 8x22b models are now up. Use `ollama run wizardlm2:8x22b-
by Patrick_Devine 2y ago
For those adventurous souls who have gobs of memory, a good GPU, and plenty of disk space, some of the 8x22b models are now up. Use `ollama run wizardlm2:8x22b-q4_0`, `ollama run wizardlm2:8x22b-q8_0`, or `ollama run wizardlm2:8x22b-fp16`. The 4 bit quantized version is about 80GB, the 8 bit quantized version is about 150GB, and the non-quantized version is 281GB (!).
- Me1000 2y agoThe 4bit quant doesn't seem to work for me, I keep getting: Error: exception create_tensor: tensor 'blk.0.ffn_gate.0.weight' not found I've tried downloading it twice now.
- Patrick_Devine 2y agoWhich version of Ollama are you using? You can check with `ollama -v`. Also, using `ollama list | grep wizardlm2` the 8x22b version should have ID `abda6e58fd1d`.
- Me1000 2y agoI ended up having to download the latest version directly from GitHub, and that fixed it. Looks like the 0.1.32 mac release hasn't been posted to your website yet.
- blixt 2y agoI've been playing with the 8x22B version for several coding tasks today, and compared to DBRX and Command-R+, two other large and capable local models recently released, this WizardLM-2 is just a lot better. It follows the instructions better, and writes higher quality code, with little details that even ChatGPT 4 gets wrong. Here's one example of the same prompt against the three models: https://twitter.com/blixt/status/1780191143747101072 https://twitter.com/blixt/status/1780191143747101072 (Running on an M3 Max 128GB it's quite fast – the times shown are including initial loading time which does take a while)