8 ms·
Granite 4.1: IBM's 8B Model Matching 32B MoE
https://research.ibm.com/blog/granite-4-1-ai-foundation-models https://research.ibm.com/blog/granite-4-1-ai-foundation-mode...
- mdp2021 5mo agoWish they also released an embedding model, in the line of their previous: compact (while good)...
- steveharing1 5mo ago[dead]
- ibgeek 5mo agoThey did: https://huggingface.co/collections/ibm-granite/granite-embedding https://huggingface.co/collections/ibm-granite/granite-embed... 311M and 97M versions.
- steveharing1 5mo agoThanks for letting me know
- mdp2021 5mo agoThank you! Yes, after yours I found out that they have produced many NNs in their "Granite 4.1" category: Granite Vision 4.1; Granite Speech 4.1; Granite Guardian 4.1; Granite Embedding Multilingual R2 - with, of course, the "Small Language Models" https://research.ibm.com/blog/granite-4-1-ai-foundation-models https://research.ibm.com/blog/granite-4-1-ai-foundation-mode...
- RugnirViking 5mo agosounds interesting. Here's hoping they release a 32B model, thats a pretty good sweet spot for feasibility of home setups. edit: I just realised they do actually have a 30b release alongside this. Haven't tried it yet.
- 2ndorderthought 5mo agoTry qwen 3.6. it will knock your socks off
- cindyllm 5mo ago[dead]
- 2ndorderthought 5mo agoI test drove it yesterday. It's pretty impressive at 8b. Runs on commodity hardware quickly. Qwen3.6 35b a3b is still my local champion but I may use this for auto complete and small tasks. Granite has recent training data which is nice. If the other small models got fine tuned on recent data I don't know if I would use this at all, but that alone makes it pretty decent. The 4b they released was not good for my needs but could probably handle tool calls or something
- steveharing1 5mo agoYea, No doubt Qwen 3.6 open weights are far more strong
- rnadomvirlabe 5mo agoWhy no doubt?
- steveharing1 5mo agoBecause Qwen 3.6 pushes way above its weight. Granite 8B is impressive, but Qwen still wins on raw capability, especially for coding.
- actionfromafar 5mo agoWay above its weights.
- drittich 5mo agoNanobanana for scale.
- rnadomvirlabe 5mo agoYou just asserted the same thing again. Why do you say this is the case?
- Havoc 5mo agoInteresting to see a pivot away from MoE by both IBM and mistral while the larger classes of SOTA of models all seem to be sticking to it. Quick vibe check of it- 8B @ Q6 - seems promising. Bit of a clinical tone, but can see that being useful for data processing and similar. You don't really want a LLM that spams you with emojis sometimes...
- embedding-shape 5mo agoMakes sense, dense for small models, dense or MoE for larger ones, end up fitting various hardware setups pretty neatly, no need for MoE at smaller scale and dense too heavy at large scale.
- npodbielski 5mo agoI never want LLM to span me with emojis. What is the use case for that? I find it highly annoying.
- Havoc 5mo agoThink it can be a plus in moderation. eg in openclaw it can add some character But yea dislike that style where each heading and bullet point gets an emoji
- 2ndorderthought 5mo agoShh people are paying for each token. Don't get them asking too many questions
- 100ms 5mo ago> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left
- cbg0 5mo agoSo are we saying it's fine that the article is written by an LLM as long as it doesn't have the tell-tale signs of LLMs?
- ramon156 5mo agoIt's more about curating the things you're publishing. Why would I bother reading what you couldn't bother to read?
- alienbaby 5mo agoThey could easily have read it, and thought , that communicates the information that it needs to. No point creating busywork for yourself just shuffling words around when the information is there, no? I guess it depends on what you want out of the article. Substance, or style?
- lelanthran 5mo ago> They could easily have read it, and thought , that communicates the information that it needs to. I'd they aren't self-aware enough or smart enough to determine that what they wrote is indistinguishable from text generation, how probable is it that they have something of value to add to any thought?
- 100ms 5mo agoI don't really see reason to complain about tool use, so long as the result is cohesive, accurate and that ultimately means a human has at least read their own output before publishing. It's a bit like receiving a supposedly personal letter that starts "Dear [INSERT_FIRST_NAME_FIELD]," are you really going to read such a thing?
- 5mo ago
- cbg0 5mo agoThe real "sleeper" might be https://huggingface.co/ibm-granite/granite-vision-4.1-4b https://huggingface.co/ibm-granite/granite-vision-4.1-4b if the benchmarks hold up for such a small model against frontier models for table & semantic k:v extraction.
- uf00lme 5mo agoWoah, is this part of the future of models? Basically little models you can use as tools.
- 2ndorderthought 5mo agoIt's looking like running your own mini ecosystem is the way of the future to me. No data centers, just a decent GPU 16-24gb of VRAM, CPU, and 32gb of RAM.
- cyanydeez 5mo agoI'm pretty sure there's someone somewhere who'll create a proper harness that's equivalent to one giant model. The difficulty is mostly local hardware has lot of memory constraints. Targeting 128GB would seem to be the current sweet spot. If we could get out of the corporate market movers of buying up all the memory, we could maybe have more. Regardless, the people in the 80s capable of pruning programs to fit on small devices is likely happening now. I'd bet most of the Chinese firms are doing it because of the US's silly GPU games among other constraints.
- tosh 5mo agoIBM announcement: https://research.ibm.com/blog/granite-4-1-ai-foundation-models https://research.ibm.com/blog/granite-4-1-ai-foundation-mode...
- agunapal 5mo agoIf you really think about why MoE came into existence, its to save significant cost during training, I don't think there was any concrete evidence of performance gains for comparable MoE vs dense models. Over the years, I believe all the new techniques being employed in post training have made the models better.
- zozbot234 5mo agoMoE models will have far more world knowledge than dense models with the same amount of active parameters. MoE is a no-brainer if your inference setup is ultimately limited by compute or memory throughput - not total memory footprint - or alternately if it has fast, high-bandwidth access to lower-tier storage to fetch cold model weights from on demand.
- regularfry 5mo agoYes, this. I can run the 122B Qwen3.5 MoE usably on one 4090 + 64GB RAM. That's a monster of a model, comparatively speaking.
- aitchnyu 5mo agoTangential. I'm a newb, can you name the concept of partitioning weights so we dont need to load whole thing?
- agunapal 5mo agoDo you mean model sharding?
- vessenes 5mo agoI think you mean inference compute? I believe all expert weights are updated in each backward pass during MoE training. The first benefit was getting a sort of structured pruning of weights through the mechanism of expert selection so that the model didn’t need to go through ‘unnecessary’ parts of the model for a given token. This then let inference use memory more efficiently in memory constrained environments, where non-hot or less common experts could be put into slow RAM, or sometimes even streamed off storage. But I don’t think it necessarily saved training cost; if it did, I’d be interested to learn how!
- whalesalad 5mo ago[flagged]
- dissahc 5mo agoqwen3.5 9b outperforms granite 4.1 30b by a huge amount (32 vs 15 on artificialanalysis benchmark)... i have no idea what made the writer of this article say so many demonstrably incorrect things
- tokenhub_dev 5mo ago[flagged]
- robotmaxtron 5mo ago"open source" show me.
- jasonlotito 5mo agoApache 2.0 License. Did you not click the link to the project? They even list it in the article. > Apache 2.0 across the board, so commercial use is clean. Did you just stop when you saw open source and come post this here because you couldn't be bothered to... look at the project and see it's cleanly and clearly listed. Edit: Like. I get it. It's fine to question open source. But this isn't hidden. It's repeated and made clear multiple times. They even link to the license: https://www.apache.org/licenses/LICENSE-2.0 https://www.apache.org/licenses/LICENSE-2.0 It wasn't hidden, it wasn't in some weird, out-of-the-way place. In fact, I found it so easily that I genuinely questioned whether it was real because of your comment. Like, why would anyone post what you posted if it was this easy to find? NOPE! It was right there.
- speedgoose 5mo agoIf I give you an amd64 elf binary under Apache2 license, is it open source?
- EagnaIonat 5mo agoCan you clarify what you mean? If you check HF you will see its Apache2 and the datasets were also permissive. It's one of the few models on the market where the creator indemnifies it against copyright claims. https://research.ibm.com/blog/granite-ethical-ai https://research.ibm.com/blog/granite-ethical-ai
- speedgoose 5mo agoOh sorry. Do we have the sources like Nvidia's Nemotron?
- EagnaIonat 5mo ago
- dash2 5mo agoNah, I ain't reading that. If they can't be bothered to get a human to write it, it can't be that important. I'm glad for them though. Or sorry that happened.
- altmanaltman 5mo ago[dead]
- deleted 5mo ago[deleted]
- osener 5mo agoThis is the official announcement: https://research.ibm.com/blog/granite-4-1-ai-foundation-models https://research.ibm.com/blog/granite-4-1-ai-foundation-mode... It is not the researchers' fault that some slop got posted here instead.
- deleted 5mo ago[deleted]
- theblazehen 5mo ago> models are judged by GPT-4 An interesting choice
- m3at 5mo agohttps://research.ibm.com/blog/granite-4-1-ai-foundation-models https://research.ibm.com/blog/granite-4-1-ai-foundation-mode... Original article on IBM research Hugging face weights: https://huggingface.co/collections/ibm-granite/granite-41-language-models https://huggingface.co/collections/ibm-granite/granite-41-la...
- pjmalandrino 5mo agoVery impressive series of SLM by IBM here. I have been using it with their Chunkless RAG concept and it is fitting very well! (for curious https://github.com/scub-france/Docling-Studio https://github.com/scub-france/Docling-Studio) I convinced that SLM are a real parto of solution for true integrated AI in process...
- 0xbadcafebee 5mo agoPeople complain a lot about LLM-written articles, but the human comments here on HN are far worse. Mostly a bunch of people extremely proud of themselves for not reading an LLM-written article, and then a bunch of people who take it at face value and make the model seem almost useful, and one comment that actually looked at other benchmarks. Good 'ol humanity, good at.. being emotional... and not doing analysis..... The article makes some good points about model design (how different size models within a family can get similar results, how to filter out hallucination, math result reinforcement), so that's worth understanding. It's analyzing a paper, which only discussed 3 sizes of the same model family. But what the article doesn't say is, compared to other model families, Granite 4.1 8B sucks. The only benchmark it does well at compared to other models is non-hallucination and instruction following. Qwen 3.5 4B (among other models) easily outclass it on every other metric. This article teaches a valuable lesson about reading articles in general. You can take useful information away from them (yes, despite being written by LLM). But you should also use critical thinking skills and be proactive to see if the article missed anything you might find relevant.
- phkahler 5mo ago>> The only benchmark it does well at compared to other models is non-hallucination and instruction following. I think instruction following is going to be the most useful thing these models do. Add a voice interface and access to a bunch of simple, straight-forward devices or APIs and you have a mildly useful assistant. If that can be done in 8B parameters it will soon run on edge devices. That's solid usefulness.
- encrux 5mo agoAnything that beats alexa-level intelligence on an edge-device is what I'd call useful as well, which shouldn't be too hard. It's mind-boggling how bad current voice assistants sometimes are when you prompt them some fairly easy questions.
- steveharing1 5mo ago[dead]
- cubefox 5mo agoIt's strange that they don't include reasoning training (RLVR). Their justification doesn't sound convincing: > While reasoning models have grown in popularity in recent years, their abilities aren’t always the most efficient way to get a result. In enterprise settings, token costs and speed are often as important as performance. That is why turning to less expensive, non-reasoning models with similar benchmark performance for select tasks like instruction following and tool calling makes sense for enterprise users. I guess they currently don't have the ability to do proper RLVR.
- mdp2021 5mo agoI may have misunderstood: is not reasoning training (RLVR) independent from the use of the "<think>" tags - is it not a method that improves results in reasoning? How do we know that it was not carried out? Incidentally: I am trying to spend some time researching in the progresses in the area (the jump from parroting, to inconsistent apparent reasoning, to reliable reasoning).
- dimitrismrtzs 5mo agoThe 8B class closing the gap with 32B is the real story of 2026 for anyone running models locally. I've been using smaller models for agent tool-use and the progress this year is real. The gap that still matters most isn't intelligence — it's consistency on structured output. When you chain 5+ tool calls in sequence, even a small per-call reliability difference compounds fast. Would love to see Granite 4.1 benchmarked specifically on multi-step function calling rather than just general benchmarks.
- woadwarrior01 5mo agoThe most salient thing about these models is that they're non-reasoning models. This makes then very token efficient and particularly well suited for local inference where decoding is usually slower than with datacenter GPUs. Link to HF collection: https://huggingface.co/collections/ibm-granite/granite-41-language-models https://huggingface.co/collections/ibm-granite/granite-41-la...
- lostmsu 5mo agoProbably worse than Gemma 4 or Qwen 3.6 with thinking off.
- smj-edison 5mo agoOn the topic of local models, is there a good equivalent to something like Claude's chat interface? I've recently started transitioning to open models after getting fed up with Claude's usage limits (I'm not in a position to drop $200/month), and for coding tasks Kimi 2.6 has been about the same as Sonnet in my experience. The only thing I've found myself missing is a nice interface to ask it questions and have it help me with my math assignments.
- steveharing1 5mo agoYou can try Open WebUI. Its genuinely useful when it comes to running open models locally with a clean interface
- RationPhantoms 5mo agoYep, couple Open WebUI for general chats and OpenCode for software-specific tasks and it feels close to Claude Desktop and Claude Code.
- camdv 5mo agoOllama does this, as does llama-server from llama.cpp
- rangerelf 5mo agollama-server from the llama.cpp package has a local web interface.
- steveharing1 5mo agoyes. I've used it a lot. its very simple and good
- simonw 5mo agoI've been mostly using LM Studio for this recently. Ollama has an OK chat UI now too. 'brew install llama.cpp' gets you 'llama-server' which provides quite a good web UI.
- 5mo ago
- simonw 5mo agoThe Granite 4.1 3B model is only 2GB from Unsloth: https://huggingface.co/unsloth/granite-4.1-3b-GGUF https://huggingface.co/unsloth/granite-4.1-3b-GGUF I ran it in LM Studio and got a pleasingly abstract pelican on a bicycle (genuinely not bad for a tiny 3B model - it can at least output valid SVG): https://gist.github.com/simonw/5f2df6093885a04c9573cf5756d34f59#file-granite-4-1-3b-svg https://gist.github.com/simonw/5f2df6093885a04c9573cf5756d34...
- tredre3 5mo agoDo you have any reasons to believe that granite is more immune to the effects of quantization than other tiny models? Otherwise it seems odd to judge a tiny model true capabilities by using its 4bit quant.
- simonw 5mo agoThis model is small enough that it might be sensible to try the same prompts against all of the quant sizes to try and spot any differences.
- samagragune 5mo ago[dead]
- SwellJoe 5mo agoI wish AI slop articles were somehow automatically flagged and deaded. They're all flowery verbose piles of crap. Yeah, the model is interesting, but the article is trash. I can't believe real humans are willing to sign their name to this stuff.
- subscribed 5mo agoPeople posting this kind of "articles" stuff probably bank on AI-led recruitment that will improve their score during the process based on the "contribution" (lol).
- mdp2021 5mo agoI read that IBM pioneered the concept of "shifting through "mid-training" from "guessing the next token" to "guessing the next logical step"". I am wondering how far is the research from "enhancing apparent reasoning" to "achieving solid, reliable reasoning". If techniques existed to shift from "guess the next highly probable" token to "guess the best next logical step", as some interpreted said research, should not that be the foremost objective?
- RandyOrion 5mo agoAlthough the performance claim of 8b dense matching 32b moe is somewhat questionable, thank you granite team for releasing small dense LLMs.
- nickpsecurity 5mo agoOn changing the training mix, H20 did that with Danube in 2024: https://arxiv.org/pdf/2401.16818 https://arxiv.org/pdf/2401.16818 With those results, I would've already done that in any models I got to train. There's also the principle that the LLM's are often better at what they saw last in their training set. That also justifies putting more logic, code, and math in at the end for an analytical or coding model. So, a few precedents for that technique already.
- nielsbot 5mo agoVery much an aside, but I'm struck by IBM's consistent iconic design language. For me it harkens all the way back to the futuristic design in 2001: A Space Odyssey from 1968. But you can also see it in their old mainframe hardware designs and other places.
- rasz 5mo agome: "Bits 5-4 indicate the video mode", how do you derive VIDEO_MASK_00b = 0x03C0 from this? Granite 4.1 8B: "I apologize for the confusion earlier. Let's clarify the derivation of the VIDEO__MASK_00b based on the description "Bits 5-6 indicate the video mode" from the equipment_list section. Given that bits 5 and 6 are used for video mode, we need to create a mask that isolates these two bits: Bit 5 is the least significant bit of the group (2^5 = 32 decimal, 0b00100000 in 8-bit form). Bit 6 is the next bit up (2^6 = 64 decimal, 0b01000000 in 8-bit form). Combining these bits: ... Therefore, the correct mask for VIDEO_MASK_00b should be 0x00C0" Errors on top of errors when converting description into binary numbers. Its hopeless for basic task like parsing/generating headers :(
- latentframe 5mo agoThe limit is changing from scaling parameters to scaling datas quality however compute is still the big constraint
- sexylinux 5mo agoIs this a model that will create reliable output or will it also produce errors?
- peter_d_sherman 5mo ago>"Stage two was RLHF training on general chat prompts using a reward model to improve helpfulness. This worked. AlpacaEval scores jumped around 18.9 points on average compared to the fine-tuned checkpoints. Then something broke. The RLHF stage, while improving chat quality, caused math benchmark scores to drop. GSM8K and DeepMind-Math both regressed." Observation: Math (which when fully decomposed, results in Logic) is at the core of how computers (traditional/older, non-LLM, programming languages work. If an LLM gets Math training wrong at any stage for any reason, then, in my opinion, that should be viewed as something that needs to be fixed at a lower level, not a higher one; not a later training level... I think it would be interesting exercise to train an LLM that only deals in simple Math, simple English, and only the ability to compute simple equations (+,-,x,/)... like, what's the absolute minimum in terms of text and layers necessary to train a model like that? I think some interesting understandings could be potentially be had by experimentation like that... I myself would love a pure (simplest, smallest possible) Text-to-Math only LLM (TTMLLM, TTMSLM?) , along with all of the necessary corpuses (which would ideally be as small as possible) and instructions necessary to train such an LLM...