4 ms·
It’s a 27B model, I highly doubt that.
by impure 2y ago
It’s a 27B model, I highly doubt that.
- xnx 2y agoWhich part do you doubt? That it is the most powerful? That it runs on a single GPU?
- warkdarrior 2y agoWhat's a better model to run on one GPU?
- int_19h 2y agoQwQ-32
- snovv_crash 2y agoIf I ask qwq32 anything that is even slightly complicated it will ramble until it exceeds the context window, then forget my question. Q4k which is all that fits (with context) on a 3090. Gemma3 27B gives me a rapid 1shot response, and actually works really well for the type of rubber duck brainstorming partner I often need.
- hy4000days 2y agoI’ve had this same exact experience.
- int_19h 2y agoTry giving it a river crossing puzzle with substitutions. QwQ can take a lot of time but it will solve it. Gemma will just confidently give you a wrong answer, and will keep giving you wrong answers if you point out the mistakes. Now, yes, QwQ will take a lot of tokens to get there (in one case it took it over 5 minutes running on Mac Studio M1 Ultra). Nevertheless, at least it can solve it.
- snovv_crash 2y agoYeah, but how many river crossing puzzles and murder mystery games was it trained on, and how many times do I actually need to solve a river crossing puzzle?
- thot_experiment 2y agoIt's unequivocally not. What usecase do you have that QwQ-32 is outperforming Gemma3? In real world uses I didn't even prefer it to Gemma 2.
- int_19h 2y agoAnything that requires reasoning rather than regurgitating. For a simple example, try the classic river crossing puzzle with non-trivial substitutions. Gemma can't solve it even if you keep pointing out where it fucks up. To be fair, this also goes for all non-CoT 70B models, so it's not surprising. But QwQ can solve it through sheer persistence because it can actually fairly reliably detect when it's wrong, and it just keeps hammering until it gets it done.
- thot_experiment 2y agoYes ok but when does this come up in the real world?
- int_19h 2y agoIt's an example of a problem that requires actual reasoning to solve, and also an example of a "looks similar therefore must use similar solution" trap that LLMs are so prone to. Translating this to code, for example, it means that Gemma is that much more likely to pretend to solve a more complicated problem that you give it by "simplifying" it to something it already knows how to solve.
- hnisoss 2y agoQWQ 32B q4 OR deepseek-r1-qwen-32b for reasoning, Qwen2.5 Coder 32B q4 (pair with QWQ 32B) for coding
- simonw 2y agoWhat's a better model that can run on a single GPU?
- idonotknowwhy 2y agoIt depends what you're trying to do. Coding - Mistral-Small-2503 or Qwen2.5-32b-Coder Reasoning - QwQ-32b Writing - Gemma-3-27b is good at this. etc
- simonw 2y agoRight, this thread is about which models are better than Gemma-3-27B. I'm a fan of Mistral Small 3 personally but I've not spent enough time with it, Gemma and the new Mistral Small 3.1 to have an opinion of which of those is the "best" model. The best indicator of model quality I can find right now is still https://lmarena.ai/?leaderboard= https://lmarena.ai/?leaderboard= Gemma 3 27B holds an impressive 8th place right now, second highest non-proprietary model after DeekSeek R1 (at 6th). QwQ-32B is 12th. Weirdly I couldn't find either of the Mistral Small 3 models on there.