5 ms·
April has been a crazy month for open weights models. I've been using Claude Code for work and Kimi 2.6 for personal projects and Kimi has been very good. Glm-5
by rubslopes 5mo ago
April has been a crazy month for open weights models. I've been using Claude Code for work and Kimi 2.6 for personal projects and Kimi has been very good. Glm-5.1 is also great. Qwen, Mimo and Deepseek I need to test some more, but they all have been producing good results. I have the impression that they are all are at the same level, or close to, Sonnet 4.6.
- bombcar 5mo agoWhat are you running them on?
- wswope 5mo agoNot OP, but having explored the field a good bit, Openrouter + pi harness in a devcontainer work great as a sane starting point. Highly recommend as a clean way to try out the upstart models.
- rubslopes 5mo agoHarness: opencode Subscription: opencode go I also use a claw agent[1] via Telegram, which uses pi.dev under the hood with my opencode go subscription. [1] I forked one of those Claw projects (bareclaw) and made many changes to it.
- abustamam 5mo agoWhen you say harness what do you mean? I see the term thrown around a lot and I think it's lost its meaning in some fashion.
- phainopepla2 5mo agoIn this context it means the tool you use the models with. So Claude Code is a harness, OpenAI's codex, Opencode, pi, etc. Those are all cli harnesses
- abustamam 5mo agoGotcha, thanks!
- rubslopes 5mo agoA fellow user replied below, but it refers to the software that uses the LLM (Claude Code, Opencode, pi.dev, etc.). --- Funny you mention that, because I started noticing the word 'harness' being used everywhere about a month ago, even though I hadn’t seen it before (in this context). As I don’t trust my memory, I assumed I had just been overlooking it and added it to my vocabulary. However, a Google Trends search does show increased usage since the end of March: https://trends.google.com.br/trends/explore?date=today%203-m&q=harness https://trends.google.com.br/trends/explore?date=today%203-m...
- gwerbin 5mo agoInteresting timing, because I think it was in March when I had a chat with Gemini about what the heck these things are supposed to be called, and that's where I first heard the term. It's probably just a coincidence. But that would be pretty interesting if we have an example of some kind of memetic phenomenon where one or more popular LLMs makes a claim that people then start to repeat as true, or at least follow up on it and start writing about it, and in so doing the claim becomes true. Even if it didn't happen in this case, I feel like it's only a matter of time.
- slopinthebag 5mo agoThey are close to Opus, not Sonnet.
- 2ndorderthought 5mo agoThe little qwen36 is at sonnet level . Kimi2.6 is about opus. The one can run on a single GPU on your gaming pc. The other you can run way cheaper from a provider. Or if you are really wealthy and have lots of gpus can run it yourself. Not sure where deepseek 4 sits
- ryandrake 5mo agoWould "lots of gpus" even help for huge models? Maybe this is exposing my lack of knowledge but don't you need to keep the whole model and context in a single GPU's VRAM? My understanding is that multiple GPUs help with scaling (can handle N X inference requests simultaneously) but it doesn't help with using large models. If that were the case, I could jam another GPU in my box and double the size of model I can serve.
- Kirby64 5mo ago> Would "lots of gpus" even help for huge models? Maybe this is exposing my lack of knowledge but don't you need to keep the whole model and context in a single GPU's VRAM? How do you think the large providers do inference? No single GPU has 1TB plus of memory on board. It’s a cluster of a bunch of gpus.
- 2ndorderthought 5mo ago1t model instances(opus, gpt,etc) are not running on a single GPU. The catch is how the cards communicate and how the model is broken up. There's a bit that goes into it but the answer is yes the more gpus the bigger the model you can run.
- ryandrake 5mo agoReally cool. I'm very much still learning about this stuff. Sounds like this inter-GPU communication is a feature of special hardware (not consumer GPUs).
- nozzlegear 5mo agoI've been using qwen 3.6 with oMLX on my M1 Mac Studio and it's been awesome. Took a while to get things set up, figure out which of the hundreds of models would be a good fit for my use case, and then get it strapped into opencode's harness, but it works! Its slower than a hosted model, obviously, but I'm tickled pink that I can give it a relatively complex chore, like I would've with my a Claude Pro subscription, and it'll churn away on it with good results and no god damn arbitrary usage limits.