5 ms·
Meta is rocking AI. As of last week I have been using their excellent muse coding harness with their model Muse Spark 1.2. Starting this morning I am running t
by mark_l_watson 2mo ago
Meta is rocking AI. As of last week I have been using their excellent muse coding harness with their model Muse Spark 1.2.
Starting this morning I am running their new local 30B model muse-glimmer on my old MacMini 32G using Ollama (remember to increase the context size!) and pi coding harness. I am getting good results with muse-glimmer running locally, with the caveat that everything runs slowly (e.g., give it a task and then go walk outside or do Qi Gong exercises for a while).
- spaceywilly 2mo agoNewb question but I’m curious what would help it to run faster? Would it need more vRAM or just system memory?
- spmurrayzzz 2mo agoThe biggest gain you'll get is faster memory, provided you have enough capacity to load all the weight into vram. The DGX sparks and Apple silicon memory bandwidth (and also memory access latency) drag down the decode speed quite a bit. I have two GPU rigs both with 2x RTX Pro 6000, can get ~250 tk/s decode with deepseek-v4-flash in native mixed precision. For context, in antirez's dwarfstar project he only gets ~20-40 tk/s on the same model @ 2bpw on M5 Max. The latter is for sure usable if it's your only option, but it's really hard for me to personally go back to speeds like that when I've experienced the former. (Also worth noting dwarfstar only has experimental support for dspark spec dec, when that lands it will definitely give a big boost at higher acceptance rates)
- wincy 2mo agoIt runs very quickly on my RTX 5090 fwiw. Whole thing is loading entirely into vRAM with a ~130k context size (the max) fitting as well.
- codazoda 2mo agoThat's a $5k 32GB card for anyone who doesn't know all these off the top of their heads (like myself).
- khimaros 2mo agoseems to underperform on Terminal Bench compared with qwen3.6-27b: 51.7 vs 60.7
- mark_l_watson 2mo agoTo be honest, I never give benchmarks a look. I just use the models for whatever I need to work on, so I can't really make comparisons that are useful for other people.
- cube00 2mo agoFriends Don't Let Friends Use Ollama https://news.ycombinator.com/item?id=47788385 https://news.ycombinator.com/item?id=47788385
- lenerdenator 2mo agoWhat do you use instead?
- reilly3000 2mo agotry oMLX or vMLX - both great projects that offer some amazing performance optimizations for Apple Silicon that utilize UMA and NVME caching efficiently. https://omlx.ai https://omlx.ai https://vmlx.net https://vmlx.net That said, it's been a few weeks since I've looked so maybe llama.cpp has those features now... they really do move that quickly.
- computershit 2mo agoI use llama.cpp w/ llama-swap https://github.com/ggml-org/llama.cpp https://github.com/ggml-org/llama.cpp https://github.com/mostlygeek/llama-swap https://github.com/mostlygeek/llama-swap
- LeBit 2mo agoThanks for the link to llama-swap. Didn’t know about it and will definitely install it.
- rancor 2mo agoFYI, llama-server can now be run in router mode so llama-swap is probably only needed for more exotic scenarios.
- kingo55 2mo agoI'm running it in router mode, but people on Reddit were recommending people use llama-swap instead. Am I missing something by using router mode?
- deleted 2mo ago[deleted]