4 ms·
Nemotron-3-Nano-30B-A3B[0][1] is a very impressive local model. It is good with tool calling and works great with llama.cpp/Visual Studio Code/Roo Code for loca
by breput 8mo ago
Nemotron-3-Nano-30B-A3B[0][1] is a very impressive local model. It is good with tool calling and works great with llama.cpp/Visual Studio Code/Roo Code for local development.
It doesn't get a ton of attention on /r/LocalLLaMA but it is worth trying out, even if you have a relatively modest machine.
[0] https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B...
[1] https://huggingface.co/unsloth/Nemotron-3-Nano-30B-A3B-GGUF https://huggingface.co/unsloth/Nemotron-3-Nano-30B-A3B-GGUF
- jychang 8mo agoIt was good for like, one month. Qwen3 30b dominated for half a year before that, and GLM-4.7 Flash 30b took over the crown soon after Nemotron 3 Nano came out. There was basically no time period for it to shine.
- ThrowawayTestr 8mo agoGenuinely exciting to be around for this. Reminds me of the time when computers were said to be obsolete by the time you drove them home.
- breput 8mo agoIt is still good, even if not the new hotness. But I understand your point. It isn't as though GLM-4.7 Flash is significantly better, and honestly, I have had poor experiences with it (and yes, always the latest llama.cpp and the updated GGUFs).
- binary132 8mo agoI recently tried GLM-4.7 Flash 30b and didn’t have a good experience with it at all.
- breput 8mo agoIt feels like GLM has either a bit of a fan club or maybe some paid supporters...
- bhadass 8mo agoSome of NVIDIA's models also tend to have interesting architectures. For example, usage of the MAMBA architecture instead of purely transformers: https://developer.nvidia.com/blog/inside-nvidia-nemotron-3-techniques-tools-and-data-that-make-it-efficient-and-accurate/ https://developer.nvidia.com/blog/inside-nvidia-nemotron-3-t...
- nextos 8mo agoDeep SSMs, including the entire S4 to Mamba saga, are a very interesting alternative to transformers. In some of my genomics use cases, Mamba has been easier to train and scale over large context windows, compared to transformers.
- binary132 8mo agoI find the Q8 runs a bit more than twice as fast as gpt-120b since I don’t have to offload as many MoE layers, but is just about as capable if not better.
- superjan 8mo agoOh those ghastly model names. https://www.smbc-comics.com/comic/version https://www.smbc-comics.com/comic/version
- deskamess 8mo agoDo they have a good multilingual embedding model? Ideally, with a decent context size like 16/32K. I think Qwen has one at 32K. Even the Gemma contexts are pretty small (8K).