3 ms·
The most salient thing about these models is that they're non-reasoning models. This makes then very token efficient and particularly well suited for local infe
by woadwarrior01 5mo ago
The most salient thing about these models is that they're non-reasoning models. This makes then very token efficient and particularly well suited for local inference where decoding is usually slower than with datacenter GPUs.
Link to HF collection:
https://huggingface.co/collections/ibm-granite/granite-41-language-models https://huggingface.co/collections/ibm-granite/granite-41-la...
- lostmsu 5mo agoProbably worse than Gemma 4 or Qwen 3.6 with thinking off.