4 ms·
I'm personally 100% convinced (assuming prices stay reasonable) that the Codex approach is here to stay. Having a human in the loop eliminates all the problems
by ghosty141 8mo ago
I'm personally 100% convinced (assuming prices stay reasonable) that the Codex approach is here to stay.
Having a human in the loop eliminates all the problems that LLMs have and continously reviewing small'ish chunks of code works really well from my experience.
It saves so much time having Codex do all the plumbing so you can focus on the actual "core" part of a feature.
LLMs still (and I doubt that changes) can't think and generalize. If I tell Codex to implement 3 features he won't stop and find a general solution that unifies them unless explicitly told to. This makes it kinda pointless for the "full autonomy" approach since effecitly code quality and abstractions completely go down the drain over time. That's fine if it's just prototyping or "throwaway" scripts but for bigger codebases where longevity matters it's a dealbreaker.
- _zoltan_ 8mo agoI'm personally 100% convinced of the opposite, that it's a waste of time to steer them. we know now that agentic loops can converge given the proper framing and self-reflectiveness tools.
- sealeck 8mo agoConverge towards what though... I think the level of testing/verification you need to have an LLM output a non-trivial feature (e.g. Paxos/anything with concurrency, business logic that isn't just "fetch value from spreadsheet, add to another number and save to the database") is pretty high.
- replygirl 8mo agoin the new world, engineers have to actually be good at capturing and interpreting requirements
- halfcat 8mo agoIn this new world, why stop there? It would be even better if engineers were also medical doctors and held multiple doctorate degrees in mathematics and physics and also were rockstar sales people.
- craigdalton 8mo agoAs a doctor, this sounds like an engineers job.
- NamlchakKhandro 8mo agosounds like the kinds of hyperbole someone whose just been forced to set a linter for the first time
- sandbags 8mo agoBut we’ve been here before. The agile movement originated as a response to the multifarious problems of big design up front.
- bcarv 8mo agoDoes the AI agent know what your company is doing right now, what every coworker is working on, how they are doing it, and how your boss will change priorities next month without being told? If it really knows better, then fire everyone and let the agent take charge. lol
- hyldmo 8mo agoNo, but Codex wouldn’t have asked you those questions either
- bcarv 8mo agoFor me, it still asks for confirmation at every decision when using plans. And when multiple unforeseen options appear, it asks again. I don’t think you’ve used Codex in a while.
- IMTDb 8mo agoA significant portion of engineering time is now spent ensuring that yes, the LLM does know about all of that. This context can be surfaced through skills, MCP, connectors, RAG over your tools, etc. Companies are also starting to reshape their entire processes to ensure this information can be properly and accurately surfaced. Most are still far from completing that transformation, but progress tends to happen slowly, then all at once.
- zeroxfe 8mo ago> it's a waste of time to steer them It's not a waste of time, it's a responsibility. All things need steering, even humans -- there's only so much precision that can be extrapolated from prompts, and as the tasks get bigger, small deviations can turn into very large mistakes. There's a balance to strike between micro-management and no steering at all.
- adw 8mo agoThe prompt is decreasingly relevant. The verification environment you have is what actually matters.
- freakynit 8mo agoI think this all comes down to information. Most prompts we give are severely information-deficient. The reason LLMs can still produce acceptable results is because they compensate with their prior training and background knowledge. The same applies to verification: it's fundamentally an information problem. You see this exact dynamic when delegating work to humans. That's why good teams rely on extremely detailed specs. It's all a game of information.
- rapind 8mo agoMaybe some day, but as a claude code user it makes enough pretty serious screw ups, even with a very clearly defined plan, that I review everything it produces. You might be able to get away without the review step for a bit, but eventually (and not long) you will be bitten.
- jaggederest 8mo agoI use that to feed back into my spec development and prompting and CI harnesses, not steering in real time. Every mistake is a chance to fix the system so that mistake is less likely or impossible. I rarely fix anything in real time - you review, see issues, fix them in the spec, reset the branch back to zero and try again. Generally, the spec is the part I develop interactively, and then set it loose to go crazy. This feels, initially, incredibly painful. You're no longer developing software, you're doing therapy for robots. But it delivers enormous compounding gains, and you can use your agent to do significant parts of it for you.
- rapind 8mo agoI assumed you'd build such a massive set of rules (that claude often does not obey) that you'd eat up your context very quickly. I've actually removed all plugins / MCPs because they chewed up way too much context.
- jaggederest 8mo agoIt's as much about what to remove as what to add. Curation is the key. Skills also give you some levers to get the kind of context-sensitive instruction you need, though I haven't delved too deeply into them. My current total instruction set is around ~2500 tokens at the moment
- Terretta 8mo ago> You're no longer developing software, you're doing therapy for robots. Or, really, hacking in "learning", building your knowhow-base. > But it delivers enormous compounding gains, and you can use your agent to do significant parts of it for you. Strong yes to both, so strong that it's curious Claude Code, Codex, Claude Cowork, etc., don't yet bake in an explicit knowledge evolution agent curating and evolving their markdown knowledge base: https://github.com/anthropics/knowledge-work-plugins https://github.com/anthropics/knowledge-work-plugins Unlikely to help with benchmarks. Very likely to improve utility ratings (as rated by outcome improvements over time) from teams using the tools together. For those following along at home: This is the return of the "expert system", now running on a generalized "expert system machine".
- halfcat 8mo ago> given the proper framing This sounds like never. Most businesses are still shuffling paper and couldn’t give you the requirements for a CRUD app if their lives depended on it. You’re right, in theory, but it’s like saying you could predict the future if you could just model the universe in perfect detail. But it’s not possible, even in theory. If you can fully describe what you need to the degree ambiguity is removed, you’ve already built the thing. If you can’t fully describe the thing, like some general “make more profit” or “lower costs”, you’re in paper clip maximizer territory.
- jondwillis 8mo ago> If you can fully describe what you need to the degree ambiguity is removed, you’ve already built the thing. Trying to get my company to realize this right now. Probably the most efficient way to work, would be on a video call including the product person/stakeholder, designer, and me, the one responsible for the actual code, so that we can churn through the now incredibly fast and cheap implementation step together in pure alignment. You could probably do it async but it’s so much faster to not have to keep waiting for one another.
- retinaros 8mo agogood luck.
- _zoltan_ 8mo agoI've been working on very complex problems with this model and the results I have have surprised people over and over again.
- NuclearPM 8mo ago> If I tell Codex to implement 3 features he won't stop and find a general solution that unifies them unless explicitly told to That could easily be automated.
- Skidaddle 8mo agoBut tokens are way cheaper than human labor
- sejje 8mo agoAider was doing this a long time ago
- xXSLAYERXx 8mo agoI've been using codex for one week and I have been the most productive I have ever been. Small prs, tight rules, I get almost exactly what I want. Things tend to go sideways when scope creeps into my request. But I just close the PR instead of fighting with the agent. In one week: 28 prs, 26 merged. Absolutely unreal.
- vidarh 8mo agoI will personally never consider using an agent that can't be easily pushed toward working on its own for long periods (hours) at a time. It's a total waste of time for me to babysit the LLM.