3 ms·
It mostly trades some potential performance loss for speed, especially at longer contexts. Nemotron 3 Super doesn't perform quite as well on benchmarks as the
by lambda 6mo ago
It mostly trades some potential performance loss for speed, especially at longer contexts.
Nemotron 3 Super doesn't perform quite as well on benchmarks as the similarly sized Qwen3.5 122B A10B model, but it goes faster and is cheaper to run.
https://artificialanalysis.ai/?models=gpt-oss-120b%2Cmistral-small-4%2Cnvidia-nemotron-3-super-120b-a12b%2Cqwen3-5-122b-a10b https://artificialanalysis.ai/?models=gpt-oss-120b%2Cmistral...
Now, you're not exactly comparing apples to apples there, since the training process (mix of data for pre-training, and the fine tuning stages of instruction turning, RLVR, etc) could have as much or more impact on how well it does as the architecture itself. Nemotron 3 Super does get better scores on performance than GPT-OSS 120B and Mistral Small 4, both also similarly sized open weights models.