4 ms·
The 7B model is available on ollama if you want to try it: `ollama run wizardlm2` or `ollama run wizardlm2:7b`. We're still crunching the 8x22B model to get it
by Patrick_Devine 2y ago
The 7B model is available on ollama if you want to try it: `ollama run wizardlm2` or `ollama run wizardlm2:7b`.
We're still crunching the 8x22B model to get it ready, and the 70B model isn't yet available.
- syntaxing 2y agoIf you can computationally afford it, 7b-q5_K_M is a way better choice. Default :7B goes to q4_0 which might give you subpar results.
- xyc 2y agoTried 7b q5 with some RAG tasks. Seems quite impressive.
- dimask 2y agoAnd quantised gguf files by Bartowski if somebody wants to download and run through llama.cpp directly https://huggingface.co/bartowski/WizardLM-2-7B-GGUF https://huggingface.co/bartowski/WizardLM-2-7B-GGUF
- xyc 2y agoAlso gguf files by abroxis: https://huggingface.co/ABX-AI/WizardLM-2-7B-GGUF-IQ-Imatrix https://huggingface.co/ABX-AI/WizardLM-2-7B-GGUF-IQ-Imatrix
- Patrick_Devine 2y agoFor those adventurous souls who have gobs of memory, a good GPU, and plenty of disk space, some of the 8x22b models are now up. Use `ollama run wizardlm2:8x22b-q4_0`, `ollama run wizardlm2:8x22b-q8_0`, or `ollama run wizardlm2:8x22b-fp16`. The 4 bit quantized version is about 80GB, the 8 bit quantized version is about 150GB, and the non-quantized version is 281GB (!).
- Me1000 2y agoThe 4bit quant doesn't seem to work for me, I keep getting: Error: exception create_tensor: tensor 'blk.0.ffn_gate.0.weight' not found I've tried downloading it twice now.
- Patrick_Devine 2y agoWhich version of Ollama are you using? You can check with `ollama -v`. Also, using `ollama list | grep wizardlm2` the 8x22b version should have ID `abda6e58fd1d`.
- Me1000 2y agoI ended up having to download the latest version directly from GitHub, and that fixed it. Looks like the 0.1.32 mac release hasn't been posted to your website yet.
- blixt 2y agoI've been playing with the 8x22B version for several coding tasks today, and compared to DBRX and Command-R+, two other large and capable local models recently released, this WizardLM-2 is just a lot better. It follows the instructions better, and writes higher quality code, with little details that even ChatGPT 4 gets wrong. Here's one example of the same prompt against the three models: https://twitter.com/blixt/status/1780191143747101072 https://twitter.com/blixt/status/1780191143747101072 (Running on an M3 Max 128GB it's quite fast – the times shown are including initial loading time which does take a while)