6 ms·
We dropped Claude. It's pretty clear this is a race to the bottom, and we don't want a hard dependency on another multi-billion dollar company just to write sof
by dakiol 6mo ago
We dropped Claude. It's pretty clear this is a race to the bottom, and we don't want a hard dependency on another multi-billion dollar company just to write software
We'll be keeping an eye on open models (of which we already make good use of). I think that's the way forward. Actually it would be great if everybody would put more focus on open models, perhaps we can come up with something like the "linux/postgres/git/http/etc" of the LLMs: something we all can benefit from while it not being monopolized by a single billionarie company. Wouldn't it be nice if we don't need to pay for tokens? Paying for infra (servers, electricity) is already expensive enough
- nate8bit 6mo agoAny recommendations on good open ones? What are you using primarily?
- blahblaher 6mo agoqwen3.5/3.6 (30B) works well,locally, with opencode
- zozbot234 6mo agoMind you, a 30B model (3B active) is not going to be comparable to Opus. There are open models that are near-SOTA but they are ~750B-1T total params. That's going to require substantial infrastructure if you want to use them agentically, scaled up even further if you expect quick real-time response for at least some fraction of that work. (Your only hope of getting reasonable utilization out of local hardware in single-user or few-users scenarios is to always have something useful cranking in the background during downtime.)
- pitched 6mo agoFor a business with ten or more engineers/people-using-ai, it might still make sense to set this up. For an individual though, I can’t imagine you’d make it through to positive ROI before the hardware ages out.
- zozbot234 6mo agoIt's hard to tell for sure because the local inference engines/frameworks we have today are not really that capable. We have barely started exploring the implications of SSD offload, saving KV-caches to storage for reuse, setting up distributed inference in multi-GPU setups or over the network, making use of specialty hardware such as NPUs etc. All of these can reuse fairly ordinary, run-of-the-mill hardware.
- DeathArrow 6mo agoSince you need at least a few of H100 class hardware, I guess you need at least few tens of coders to justify the costs.
- pitched 6mo agoI see the 512GB Mac Studios aren’t for sale anymore but that was a much cheaper path
- wuschel 6mo agoWhat near SOTA open models are you referring to?
- cyberax 6mo agoI'm backing up a big dataset onto tapes, so I wanted to automate it. I have an idle 64Gb VRAM setup in my basement, so I decided to experiment and tasked it with writing an LTFS implementation. LTFS is an open standard for filesystems for tapes, and there's an implementation in C that can be used as the baseline. So far, Qwen 3.6 created a functionally equivalent Golang implementation that works against the flat file backend within the last 2 days. I'm extremely impressed.
- Gareth321 6mo agoIt is surprisingly competent. It's not Opus 4.6 but it works well for well structured tasks.
- pitched 6mo agoI want to bump this more than just a +1 by recommending everyone try out OpenCode. It can still run on a Codex subscription so you aren’t in fully unfamiliar territory but unlocks a lot of options.
- cpursley 6mo agoHow are you running it with opencode, any tips/pointers on the setup?
- jherdman 6mo agoIs this sort of setup tenable on a consumer MBP or similar?
- pitched 6mo agoFor a 30B model, you want at least 20GB of VRAM and a 24GB MBP can’t quite allocate that much of it to VRAM. So you’d want at least a 32GB MBP.
- zozbot234 6mo agoIt's a MoE model so I'd assume a cheaper MBP would simply result in some experts staying on CPU? And those would still have a sizeable fraction of the unified memory bandwidth available.
- pitched 6mo agoI haven’t tried this myself yet but you would still need enough non-vram ram available to the cpu to offload to cpu, right? This is a fully novice question, I have not ever tried it.
- tredre3 6mo agoYou're correct. If you don't have enough RAM for the model, it can still run but most of it will run on the CPU and be continuously reloaded from the SSD (through mmap). A medium MoE like 35B can still achieve usable speeds in that setup, mind you, depending on what you're doing.
- _blk 6mo agoIs there any model that practically compares to Sonnet 4.6 in code and vision and runs on home-grade (12G-24G) cards?
- macwhisperer 6mo agoim currently running a custom Gemma4 26b MoE model on my 24gb m2... super fast and it beat deepseek, chatgpt, and gemini in 3 different puzzles/code challenges I tested it on. the issue now is the low context... I can only do 2048 tokens with my vram... the gap is slowly closing on the frontier models
- equasar 6mo agoThe thing I dislike about OpenCode is the lack of capabilities of their editor, also, resource intensive, for some reason on a VM it chuckles each 30 mins, that I need to discard all sessions, commits, etc. I don't know if it is bun related, but in task manager, is the thing that is almost at the top always on CPU usage, turns out for me, bun is not production ready at all. Wish Zed editor had something like BigPickle which is free to use without limits.
- Jarred 6mo ago> turns out for me, bun is not production ready What issue did you run into?
- cmrdporcupine 6mo agoGLM 5.1 via an infra provider. Running a competent coding capable model yourself isn't viable unless your standards are quite low.
- myaccountonhn 6mo agoWhat infra providers are there?
- elbear 6mo agoThere's DeepInfra. There's also OpenRouter where you can find several providers.
- culi 6mo agoLMArena actually has a nice Pareto distribution of ELO vs price for this model elo $/M --------------------------------------- glm-5.1 1538 2.60 glm-4.7 1440 1.41 minimax-m2.7 1422 0.97 minimax-m2.1-preview 1392 0.78 minimax-m2.5 1386 0.77 deepseek-v3.2-thinking 1369 0.38 mimo-v2-flash (non-thinking) 1337 0.24 https://arena.ai/leaderboard/code?viewBy=plot&license=open-source https://arena.ai/leaderboard/code?viewBy=plot&license=open-s...
- logicprog 6mo agoLMArena isn't very useful as a benchmark, however I can vouch for the fact that GLM 5.1 is astonishingly good. Several people I know who have a $100/mo Claude Code subscription are considering cancelling it and going all in on GLM, because it's finally gotten (for them) comparable to Opus 4.5/6. I don't use Opus myself, but I can definitely say that the jump from the (imvho) previous best open weight model Kimi K2.5 to this is otherworldly — and K2.5 was already a huge jump itself!
- DeathArrow 6mo agoI am using GLM 5.1 and MiniMax 2.7.
- ahartmetz 6mo ago>we don't want a hard dependency on another multi-billion dollar company just to write software One of two main reasons why I'm wary of LLMs. The other is fear of skill atrophy. These two problems compound. Skill atrophy is less bad if the replacement for the previous skill does not depend on a potentially less-than-friendly party.
- tossandthrow 6mo agoYou can argu that you will have skill atrophy by not using LLMs. We have gone multi cloud disaster recovery on our infrastructure. Something I would not have done yet, had we not had LLMs. I am learning at an incredible rate with LLMs.
- mgambati 6mo agoI kind feel the same. I’m learning things and doing things in areas that would just skip due to lack of time or fear. But I’m so much more detached of the code, I don’t feel that ‘deep neural connection’ from actual spending days in locked in a refactor or debugging a really complex issue. I don’t know how a feel about it.
- afzalive 6mo agoAs someone who's switched from mobile to web dev professionally for the last 6 months now. If you care about code quality, you'll develop that neural connection after some time. But if you don't and there's no PR process (side projects), the motivation to form that connection is quite low.
- hombre_fatal 6mo ago> If you care about code quality, you'll develop that neural connection after some time. No, because you can get LLMs to produce high quality code that has gone through an infinite number of refinement/polish cycles and is far more exhaustive than the code you would have written yourself. Once you hit that point, you find yourself in a directional/steering position divorced from the code since no matter what direction you take, you'll get high quality code.
- tossandthrow 6mo agoThe lock in is so incredibly poor. I could switch to whatever provider in minuets. But it requires that one does not do something stupid. Eg. For recurring tasks: keep the task specification in the source code and just ask Claude to execute it. The same with all documentation, etc.
- i_love_retros 6mo ago> we don't want a hard dependency on another multi-billion dollar company just to write software My manager doesn't even want us to use copilot locally. Now we are supposed to only use the GitHub copilot cloud agent. One shot from prompt to PR. With people like that selling vendor lock in for them these companies like GitHub, OpenAI, Anthropic etc don't even need sales and marketing departments!
- tossandthrow 6mo agoYou are aware that using eg. Github copilot is not one shot? It will start an agentic loop.
- dgellow 6mo agoUnnecessary nitpicking
- tossandthrow 6mo agoWhy? One shoting has a very specific meaning, and agentic workflows are not it? What is the implied meaning I should understand from them using one shot? They might refer to the lack of humans in the loop.
- dgellow 6mo agoYou give a prompt, you get a PR. If it is ready to merge with the first attempt, that’s a one shot. The agentic loop is a detail in their context
- deleted 6mo ago[deleted]
- aliljet 6mo agoWhat open models are truly competing with both Claude Code and Opus 4.7 (xhigh) at this stage?
- Someone1234 6mo agoThat's a lame attitude. There are local models that are last year's SOTA, but that's not good enough because this year's SOTA is even better yet still... I've said it before and I'll say it again, local models are "there" in terms of true productive usage for complex coding tasks. Like, for real, there. The issue right now is that buying the compute to run the top end local models is absurdly unaffordable. Both in general but also because you're outbidding LLM companies for limited hardware resources. You have a $10K budget, you can legit run last year's SOTA agentic models locally and do hard things well. But most people don't or won't, nor does it make cost effective sense Vs. currently subsidized API costs.
- gbro3n 6mo agoI completely see your point, but when my / developer time is worth what it is compared to the cost of a frontier model subscription, I'm wary of choosing anything but the best model I can. I would love to be able to say I have X technique for compensating for the model shortfall, but my experience so far has been that bigger, later models out perform older, smaller ones. I genuinely hope this changes through. I understand the investment that it has taken to get us to this point, but intelligence doesn't seem like it's something that should be gated.
- Someone1234 6mo agoRight; but every major generation has had diminishing returns on the last. Two years ago the difference was HUGE between major releases, and now we're discussing Opus 4.6 Vs. 4.7 and people cannot seem to agree if it is an improvement or regression (and even their data in the card shows regressions). So my point is: If you have the attitude that unless it is the bleeding edge, it may have well not exist, then local models are never going to be good enough. But truth is they're now well exceeding what they need to be to be huge productivity tools, and would have been bleeding edge fairly recently.
- boxingdog 6mo ago[dead]
- dewarrn1 6mo agoI'm hopeful that new efficiencies in training (Deepseek et al.), the impressive performance of smaller models enhanced through distillation, and a glut of past-their-prime-but-functioning GPUs all converge make good-enough open/libre models cheap, ubiquitous, and less resource-intensive to train and run.
- dgellow 6mo agoAnother aspect I haven’t seen discussed too much is that if your competitor is 10x more productive with AI, and to stay relevant you also use AI and become 10x more productive. Does the business actually grow enough to justify the extra expense? Or are you pretty much in the same state as you were without AI, but you are both paying an AI tax to stay relevant?
- senordevnyc 6mo agoEither the business grows, or the market participants shed human headcount to find the optimal profit margin. Isn’t that the great unknown: what professions are going to see headcount reduction because demand can’t grow that fast (like we’ve seen in agriculture), and which will actually see headcount stay the same or even expand, because the market has enough demand to keep up with the productivity gains of AI? Increasingly I think software writ large is the latter, but individual segments in software probably are the former.
- xixixao 6mo agoThis is the “ad tax” reasoning, but ultimately I think the answer is greater efficiency. So there is a real value, even if all competitors use the tools. It’s like saying clothing manufacturers are paying the “loom tax” tax when they could have been weaving by hand…
- SlinkyOnStairs 6mo agoSoftware development is not a production line, the relationship between code output and revenue is extremely non-linear. Where producing 2x the t-shirts will get you ~2x the revenue, it's quite unlikely that 10x the code will get you even close to 2x revenue. With how much of this industry operates on 'Vendor Lock-in' there's a very real chance the multiplier ends up 0x. AI doesn't add anything when you can already 10x the prices on the grounds of "Fuck you. What are you gonna do about it?"
- groundzeros2015 6mo agoYep and in a vendor lock in scenario, fixing deep bugs or making additions in surgical ways is where the value is. And Claude helps you do that, by giving you more information, analyzing options, but it doesn’t let you make that decision 10x faster.
- GaryBluto 6mo ago> open models Google just released Gemma 4, perhaps that'd be worth a try?
- SilverElfin 6mo agoIs that why they are racing to release so many products? It feels to me like they want to suck up the profits from every software vertical.
- Bridged7756 6mo agoYeah it seems so. Anthropic has entered the enshittification phase. They got people hooked onto their SOTAs so it's now time to keep releasing marginal performance increase models at 40% higher token price. The problem is that both Anthropic and OpenAI have no other income other than AI. Can't Google just drown them out with cheaper prices over the long run? It seems like an attrition battle to me.
- DeathArrow 6mo ago>perhaps we can come up with something like the "linux/postgres/git/http/etc" of the LLMs: something we all can benefit from while it not being monopolized by a single billionarie company Training and inference costs so we would have to pay for them.
- groundzeros2015 6mo agoDeveloping linux/postgres/git also costs, and so do the computers and electricity they use.
- throwaway613746 6mo ago[dead]
- michaelje 6mo agoOpen models keep closing the eval gap for many tasks, and local inference continues to be increasingly viable. What's missing isn't technical capability, but productized convenience that makes the API path feel like the only realistic option. Frontier labs are incentivized to keep it that way, and they're investing billions to make AI = API the default. But that's a business model, not a technical inevitability.
- trueno 6mo agoim hoping and praying that local inference finds it's way to some sort of baseline that we're all depending on claude for here. that would help shape hardware designs on personal devices probably something in the direction of what apple has been doing. ive had to like tune out of the LLM scene because it's just a huge mess. It feels impossible to actually get benchmarks, it's insanely hard to get a grasp on what everyone is talking about, bots galore championing whatever model, it's just way too much craze and hype and misinformation. what I do know is we can't keep draining lakes with datacenters here and letting companies that are willing to heel turn on a whim basically control the output of all companies. that's not going to work, we collectively have to find a way to make local inference the path forward. everyone's foot is on the gas. all orgs, all execs, all peoples working jobs. there's no putting this stuff down, and it's exhausting but we have to be using claude like _right now_. pretty much every company is already completely locked in to openai/gemini/claude and for some unfortunate ones copilot. this was a utility vendor lock in capture that happened faster than anything ive ever seen in my life & I already am desperate for a way to get my org out of this.
- hakfoo 6mo agoI'm frustrated that there's not "solid" instructional tooling. I either see people just saying "keep trying different prompts and switching models until you get lucky" or building huge cantilevered toolchains that seems incredibly brittle, and even then, how well do they really work? I get choice paralysis when you show me a prompt box-- I don't know what I can reasonably ask for and how to best phrase it, so I just panic. It doesn't help when we see articles saying people are getting better outcomes by adding things like "and no bugs plz owo" I'm sure this is by design-- anything with clear boundaries and best practices would discourage gacha style experimentation. Can you trust anyone who sells you a metered service to give you good guidance on how to use it efficiently?
- Frannky 6mo agoOpencode go with open models is pretty good
- gbgarbeb 6mo ago[dead]
- sergiotapia 6mo agoI can recommend this stack. It works well with the existing Claude skills I had in my code repos: 1. Opencode 2. Fireworks AI: GLM 5.1 And it is SIGNIFICANTLY cheaper than Claude. I'm waiting eagerly for something new from Deepseek. They are going to really show us magic.
- giancarlostoro 6mo ago> I think that's the way forward. Actually it would be great if everybody would put more focus on open models, I'm still surprised top CS schools are not investing in having their students build models, I know some are, but like, when's the last time we talked about a model not made by some company, versus a model made by some college or university, which is maintained by the university and useful for all. It's disgusting that OpenAI still calls itself "Open AI" when they aren't truly open.
- leonidasv 6mo ago>perhaps we can come up with something like the "linux/postgres/git/http/etc" of the LLMs I fear that this may not be feasible in the long term. The open-model free ride is not guaranteed to continue forever; some labs offer them for free for publicity after receiving millions in VC grants now, but that's not a sustainable business model. Models cost millions/billions in infrastructure to train. It's not like open-source software where people can just volunteer their time for free; here we are talking about spending real money upfront, for something that will get obsolete in months. Current AI model "production" is more akin to an industrial endeavor than open-source arrangements we saw in the past. Until we see some breakthrough, I'm bearish on "open models will eventually save us from reliance on big companies".
- falkensmaize 6mo ago"get obsolete in months" If you mean obsolete in the sense of "no longer fit for purpose" I don't think that's true. They may become obsolete in terms of "can't do hottest new thing" but that's true of pretty much any technology. A capable local model that can do X will always be able to do X, it just may not be able to do Y. But if X is good enough to solve your problem, why is a newer better model needed? I think if we were able to achieve ~Opus 4.6 level quality in a local model that would probably be "good enough" for a vast number of tasks. I think it's debatable whether newer models are always better - 4.7 seems to be somewhat of a regression for example.
- somewhereoutth 6mo agoMy understanding is that the major part of the cost of a given model is the training - so open models depend on the training that was done for frontier models? I'm finding hard to imagine (e.g.) RLHF being fundable through a free software type arrangement.
- zozbot234 6mo agoNo, the training between proprietary and open models is completely different. The speculation that open models might be "distilled" from proprietary ones is just that, speculation, and a large portion of it is outright nonsense. It's physically possible to train on chat logs from another model but that's not "distilling" anything, and it's not even eliciting any real fraction of the other model's overall knowledge.
- tehjoker 6mo agoI don't know what to make of it, I am skeptical of OpenAI/Anthropic claims about distillation, but I did notice DeepSeek started sounding a lot like Claude recently.
- OrvalWintermute 6mo agoI'm increasingly thinking the same as our spend on tokens goes up. If you have HPC or Supercompute already, you have much of the expertise on staff already to expand models locally, and between Apple Silicon and Exo there are some amazingly solutions out there. Now, if only the rumors about Exo expanding to Nvidia are true..
- deleted 6mo ago[deleted]
- finghin 6mo agoI’m imagining a (private/restricted) tracker style system where contributors “seed” compute and users “leech”.
- crgk 6mo agoWho’s your “we,” if you don’t mind sharing? I’m curious to learn more about companies/organizations with this perspective.
- sky2224 6mo agoThis is part of the reason why I'm really worried that this is all going to result in a greater economic collapse than I think people are realizing. I think companies that are shelling out the money for these enterprise accounts could honestly just buy some H100 GPUs and host the models themselves on premises. Github CoPilot enterprise charges $40 per user per month (this can vary depending on your plan of course), but at this price for 1000 users that comes out to $480,000 a year. Maybe I'm missing something, but that's roughly what you're going to be spending to get a full fledged hosting setup for LLMs.
- merlinoa 6mo agoMost companies don't want to host it themselves. They want someone to do it for them, and they are happy to pay for it. If it makes their lives easier and does not add complexity, then it has a lot of value.
- subarctic 6mo agoOut of curiosity, how many concurrent users could you get with a hosting setup at that price? If let's say 10% of those 1000 users were using it at the same time would it handle it? What about 30% or 100%?
- sky2224 6mo agoYou made a good point that I didn't think through fully. It's the concurrent user aspect that heavily impacts things. Currently, you'd probably need quite a bit more investment to the point of having a mini data center to do what I'm proposing. However, we've been seeing advancements in compressing context and capabilities of smaller models that I don't think it'd be too far off to see something like what I'm talking about within the next 5 years.
- wahnfrieden 6mo agoor just use codex
- atleastoptimal 6mo agoOpen models are only near SOTA because of distillation from closed models.
- sourya4 6mo agoyep!! had similar thoughts on the the "linux/postgres/git/http/etc" of the LLMs made a HN post of my X article on the lock-in factor and how we should embrace the modular unix philosophy as a way out: https://news.ycombinator.com/item?id=47774312 https://news.ycombinator.com/item?id=47774312