6 ms·
Wow Google is becoming the new pre Llama 4 Meta when it comes to releasing open weights models.
by nickandbro 4mo ago
Wow Google is becoming the new pre Llama 4 Meta when it comes to releasing open weights models.
- embedding-shape 4mo agoI dunno, feels a bit unfair to companies that actually do FOSS releases (Gemma 4 being released under Apache 2.0 license) to compare them to a company that never done any FOSS releases, and mostly done proprietary "available to download" releases.
- seba_dos1 4mo agoNote that a binary released under Apache 2.0 license does not yet make it FOSS.
- embedding-shape 4mo agoAgreed, miles ahead though from "proprietary" which is what Meta been using for most model releases. Ideally companies would share the fucking datasets and training code already, but no, no one wants to talk about the source of those or even share the ones they have as then who knows what comes out of Pandora's box...
- jimmy76615 4mo agoNVIDIA does a pretty good job on that front.
- redman25 4mo agoIDK this model release is a bit disappointing considering the community has been chomping at the bit for the 124ba4b model. There was some leaked info about it but people suspect it was not released because it was too close to gemini flash in performance.
- brianwawok 4mo agoEvery other Google model I have tried felt very weak compared to qwen models. I dont have a ton of use case for multimodal though, so its very possible this is a fantastic multimodal model.
- wongarsu 4mo agoGemma 4 27b and 32b feel pretty capable for text and visionn. Comparable with qwen, maybe a bit better on tool calling heavy tasks I am not overly impressed with the smaller gemma models. And gemma 3 was a bit of a mixed bag, great at some things, bad at most others
- verdverm 4mo agoqwen3.6 was my favorite, then I tried the deepseek-v4-{flash,pro} still making my way through deep dives on the chinese open weights, they are all pretty good and way more cost / resource effective
- thot_experiment 4mo agoHard disagree, Qwen multimodal is way better than google's, but Gemma 31b runs laps around Qwen 27B in complex engineering tasks. Maybe Qwen is better at slopcoding web framework CRUD, but for embedded dev there's no comparison.
- avadodin 4mo agoE4B is decent at instruction following. It managed to produce a deliverable on par with the lowest tier of paid models. Even higher tiers often just ignore all rules when they feel like it. I wish it was an 8BA1B MoE model with the newer acceleration 1B or maybe even a tailor-made sub-1B slapped on top. That would make it an awesome local model for the average laptop.