3 ms·
If you're primarily writing code yourself or meticulously reviewing the output from agents, then you're right. However, if you tried to have any of those models
by jrflo 2mo ago
If you're primarily writing code yourself or meticulously reviewing the output from agents, then you're right. However, if you tried to have any of those models one-shot an app or do some highly agentic work, they would certainly fail. That's the future people are looking towards with these valuations: when its no longer economical for humans to write or even understand code, just let the models drive because they are superhuman at it. Not saying we are there today, but that's when you really start to see the benefit of more expensive models. Luna or Deepseek flash would never find any of the mathematical discoveries or security exploits that the larger models can find.
- _kulang 2mo agoClaude is certainly able to make a superhuman mess. All of its efficacy still hinges upon good architecture and programming principles, which do not seem to be instilled in the model by anything other than luck
- Incipient 2mo agoTwo things I've figured out with pretty good certainty 1) on an existing codebase/hand written, even fable absolutely will "make a mess" if you do large changes at once, or don't check the output. 2) with detailed, WELL NUMBERED (Claude is remarkably good at following number references), multi-level design and build documentation Claude absolutely CAN produce moderately complex applications (crud with some lightly branching business workflows). The corollary to (2) however is can you actually get that working code to be production grade, AND maintain it? that far I haven't got yet.
- _kulang 2mo agoI have quite a bit of success with Claude on Jira tickets as long as the tasks are scoped well. Put effort into the ticket, and fire Claude at the ticket with MCP. I essentially review the PR. But this is only possible because I have created a completely new python development workflow with fixed formatting, style, and unification of tooling around uv. We’ve also split our projects into smaller repos, some which function as libraries and some which are applications depending on those libraries. I think the key to success here is to limit what context is required to do development. I’d say it’s slightly more on the extreme side of “modular” than a codebase I would write otherwise
- formvoltron 2mo agoBut.. what is it that anthropic does that cannot be replicated by open models teamed up with open source? Heck open source even has cheap AI to help write the code now.
- SV_BubbleTime 2mo agoThe moat is money. How do you get more money? Point to your moat and ask for more money!
- jrflo 2mo agoMake new mathematical discoveries, and the security capabilities of the closed models have not yet been rivaled by open ones. Also, just because a model is open now doesn't mean it will always be. If/when China or meta catches up to the frontier, they'll instantly go closed source. China, the biggest surveillance state in the world, would love to have all user data pouring in to its servers. It's just a business strategy, they're not doing it out of the kindness of their hearts.
- malux85 2mo agoI'm not convinced that one-shotting things is anything other than a vanity-metric. Maybe in the distant future where quickly building a visualisation to help explain some concept would be valuable to one shot quickly - but "One shotting an app" is ridiculous because app development (or any development) is never "build it and then finish" but is an interative process, testing feedback, user feedback, and even app-creator communication ambiguity means being able to "one shot an app" is pretty worthless
- yogthos 2mo agoI've been using GLM 5.3 for the past week, and I find to does a straight up better job than Opus 5 on my projects. I use both for agentic work, and GLM tends to dig deeper into tasks on the latest version.
- jjav 2mo ago> you tried to have any of those models one-shot an app If you try the same with the highest cost models, it all ends in tears anyway. A completely new app may seem to work ok but as soon as you're doing ongoing development, it needs to be broken down into smaller pieces of work. We're seeing this a lot at work, some people hope to skip the careful planning and review by using the top models, and it just leads to huge bills and a lot of rewriting.