2 ms·
I suspect this might be due to MTP mispedictions. Either because the 3.8 quantized model weights do not include MTP heads, or as what happened with Ornith-1.5 r
by woadwarrior01 1mo ago
I suspect this might be due to MTP mispedictions. Either because the 3.8 quantized model weights do not include MTP heads, or as what happened with Ornith-1.5 recently, corrupted MTP heads, or due to a software issue in Ollama.
- rbanffy 1mo agoI remember an article a couple weeks back where someone used an HPE server with two older Xeon E-series CPUs and got reasonable performance by compiling llama with optimal switches for the architecture (memory alignment, page sizes, etc). I was impressed because those Xeons only had AVX2. If you are willing to spend a little more, you can get a slightly newer Xeon with AVX-512 and 4 sockets. I can't find the post though. Edit: a quick search from a server refurbisher nearby gives me a dual Xeon Gold 6330 28-core machine with 512GB and a 16GB V100 GPU. The memory is spread across 32 slots, for about £6,470.