5 ms·
China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for
by jjcm 2mo ago
China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you.
What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.
- icedrift 2mo agoI'm still skeptical of the smaller models after the talent exodus a few months ago.
- jimbo808 2mo agoAt this point I feel like the only factor differentiating SOTA models now is who they’re propagandizing you on behalf of (not considering agentic tooling/state management, etc).
- Zambyte 2mo agoQwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.
- snapplebobapple 2mo agoWorks great with room to spare on my lenovo pgx too
- bitexploder 2mo agoI find 35B A3B viable as well, but your harness and runtime really matters to get tool calling and such dialed in. In fact, I would encourage you to experiment with it some as I find I get more reliable output from 35B A3B, though 27B is still generally smarter. A3B with a review cycle or two from 27B is great for me. One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B model, but it is close enough it can usually figure it out with the right tools.
- jona-f 2mo agoHow do you do the review cycles? Is this some automated feature of your harness? Do you have a generic prompt for this or ask 27B yourself?
- bitexploder 2mo agoUsually something like dsv4 running them. Depends in the horizon. It’s kind of an overnight thing for the project. Use matt pocock skills or similar. Build good spec. Iterate on it with a big model or your brain. Break it down into pieces and features like you would for a human. Tell them to work on tickets. Have other agent review. Repeat. Stuff gets built. You can keep a smarter agent in a loop (think bash) to review and have a simple decision tree: implement, review, mark done, pick next ticket. When no tickets stop. If error or pathology detected, touch a stop file. That is what I run overnight. A3B likes a simple harness as does 27B, mostly unmodified Pi with a proxy to fix model bugs. It takes some investment. It isn’t batteries included. These small models are not very smart. They need a pretty narrow task domain.
- monster_truck 2mo agoSame! The only reason I'm not using it more is because it's summertime. I'm not in any hurry. Setting the memory to "fast timings" is good for 8-12% more tokens/second if you haven't tried yet. I miss the slightly older days of AMD when powerplay tables were unlocked and we could configure the timings and voltages manually, there's another 30% being left on the table ez
- MrDrMcCoy 2mo agoWhat do you mean by 'Setting the memory to "fast timings"'? The only runtime I can get working for my GPUs is llama.cpp, which I haven't seen anything like that in its argument set. My perusal of the options for vllm and sglang didn't suggest anything similar either before failing miserably.
- sonic45132 2mo agoI think they are referring to the AMD drivers on windows, under the overclocking section you can enable fast memory timings. Not sure if this sort of thing is exposed on linux.
- monster_truck 2mo agoYou can edit sys files or use amdmemorytweak on linux.
- doginasuit 2mo agoAnother potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.
- AustinDev 2mo agoThey always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
- michelsedgh 2mo agoWhat an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.
- AussieWog93 2mo agoI'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.
- dullcrisp 2mo agoWhy isn’t it more like building a hardware store using lumber you purchased from a competing hardware store? Or founding a school using an education you obtained at a different school?
- ygjb 2mo agoBecause that doesn't satisfy the narrative of American exceptionalism. It's easier to point at something and say it was stolen or copied than it is to compete, especially with the political climate in the US. This isn't an anti-American sentiment. It is an anti-corporate/regulatory capture/embrace and extinguish sentiment (which probably reads the same to many people these days).
- monster_truck 2mo agoSomething I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome
- aliasxneo 2mo agoAnd I had the opposite experience. It's a really interesting phenomenon that I can't really explain. My co-founder swears by Deepseek and yet just the other day we were conversing and he was telling me about some of the issues with the way the AI was behaving and trying to show off the cool workarounds he came up with to limit it. I was like, "Interesting, yeah, I've literally never had that problem." I suspect that the models are genuinely close and that certain experiences get felt across providers but are inconsistent enough to convince people one is superior to the other. I for one have tried Deepseek on and off since my co-founder is fond of it and I've stopped trying now because I never have a good experience.
- wanderlust123 2mo agoThat’s not surprising. I have been using Deepseek and it consistently produces excellent output given the right context howevwr. It depends on the task as it does have blindspots.
- surgical_fire 2mo agoThat may be true. I switched to DeepSeek entirely once I decided to put 10 bucks on it and I realized that it could do whatever I was throwing at Claude or ChatGPT prior to that. I recommended it to one of my friends, and he was surprised DeepSeek could solve task that Claude got stuck at. I was surprised at it too. I know others that tried and were less impressed too.
- dw_arthur 2mo agoWho is going to break it to the Americans that China is more than a slight favorite to win an existential battle over which country is better at math?