22 ms·
Why the push for Agentic when models can barely follow a simple instruction?
- kykat 1y agoThe replies are all a variation of: "You're using it wrong"
- motorest 1y ago> The replies are all a variation of: "You're using it wrong" I don't know what you are trying to say with your post. I mean, if two persons feed their prompts to an agent and while one is able to reach their goals the other fails to achieve anything, would it be outlandish to suggest one of them is using it right whereas the other is using it wrong? Or do you expect the output to not reflect the input at all?
- ares623 1y agoI expect the $500 billion magic machine to be magic. Especially after all the explicit threats to me and my friends livelihoods.
- motorest 1y ago> I expect the $500 billion magic machine to be magic. Especially after all the explicit threats to me and my friends livelihoods. That's a problem you are creating for yourself by believing in magical nonsense. Meanwhile, the rest of the world is gradually learning how to use the tool to simplify their work, being it helping onboard onto projects, doing ad-hoc code reviews, serving as sparring partners, helping with design work, and yes even creating complete projects from scratch.
- theshrike79 1y agoIt's giving "I paid $200k for this RV, the cruise control should keep the car on the road while I go in the back to make coffee".
- kykat 1y agoOf course the output reflects the input, that's why it's a bad idea to let the LLM run in a loop without constraints, it's simple maths, if something is 99% accurate, after 5 times is 95% accurate, after 10 steps it's about 90% accurate, after 100 times it's about 36% accurate. For LLMs to be effective, you (or something else) needs to constantly find the errors and fix it.
- ninetyninenine 1y agoI’ve seen LLM catch and fix their own mistakes and literally tell me they were wrong and that they are fixing their self made wrong mistake. This analogy is therefore not accurate as error rate can actually decrease over time.
- kykat 1y agoIf we assume that each action has 99% success rate, and when it fails, it has 20% chance of recovery, and if the math here by gemini 2.5 pro is correct, that means the system will tend towards 95% chance of success. === In equilibrium, the probability of leaving the Success state must equal the probability of entering it. (Probability of being in S) * (Chance of leaving S) = (Probability of being in F) * (Chance of leaving F) Let P(S) be the probability of being in Success and P(F) be the probability of being in Failure. P(S) * 0.01 = P(F) * 0.20 Since P(S) + P(F) = 1, we can say P(F) = 1 - P(S). Substituting that in: P(S) * 0.01 = (1 - P(S)) * 0.20 0.01 * P(S) = 0.20 - 0.20 * P(S) 0.21 * P(S) = 0.20 P(S) = 0.20 / 0.21 ≈ 0.95238
- ninetyninenine 1y agoThat’s math based off of arbitrary initial assumptions. There are numbers that work. All this math is useless. Use your brain. The entire point I’m communicating is that it’s not a given that it must become less accurate. There are multiple open possibilities here and scenarios that can occur. Doing random math here as if you’re dropping the mic is just pointless. It doesn’t do anything. It’s like making up a cosmological constant and saying the universe is collapsing look at my math.
- Alex_L_Wood 1y agoAnd yours is also "you are using it wrong" in the spirit. Are they doing the same thing? Are they trying to achieve the same goals, but fail because one is lacking some skill? One person may be someone who needs a very basic thing like creating a script to batch-rename his files, another one may be trying to do a massive refactoring. And while the former succeeds, the latter fails. Is it only because someone doesn't know how to use agentic AI, or because agentic AI is simply lacking?
- berkes 1y agoAnd some more variations that, in my anecdotal experience make or break the agentic experience: * strictness of the result - a personal blog entry vs a complex migration to reform a production database of a large, critical system * team constraints - style guides, peer review, linting, test requirements, TDD, etc * language, frameworks - quick node-js app vs a java monolyth e.g. * legacy - a 12+ year Django app vs a greenfield rust microservice * context - complex, historical, nonsensical business constraints and flows vs a simple crud action * example body - a simple crud TODO in PHP or JS, done a million times vs a event-sourced, hexagonal architecrtured, cryptographical signing system for govt data.
- throwawayb2025 1y agoI had both good and bad experience. Bad with regex or things involving recursion. Also bad at integration between modules. I have not tried to solve this yet, by giving documentation of both the modules. Also model used impacted. To understand java code it was great. I first ask it to generate the detailed prompt by giving a high level prompt. Then use the detailed prompt to execute the task. Java code size Upto 20k loc is fine. Other wise context becomes big. So you have to do module by module. I believe to have a discussion maybe someone has to take an open source code example and then say it doesn't work. Other people can then discuss and decide. Overall happy with gpt5 and claude code. Edit:updated bad at integration
- motorest 1y ago> Also bad at integration between modules. I have not tried to solve this yet, by giving documentation of both the modules. You should first draft the interface and roll out coverage with automated tests, and then prompt your way into filling in the implementation. If you just post a vague prompt on how you want multiple modules workinh together, odds are the output might not met implicit constraints.
- leptons 1y agoIn my experience it depends on which way the wind is blowing, random chance, and a lot of luck. For example, I was working on the same kind of change across a few dozen files. The prompt input didn't change, the work didn't change, but the "AI" got it wrong as often as it got it right. So was I "using it wrong" or was the "AI" doing it wrong half the time? I tried several "AI" offerings and they all had similar results. Ultimately, the "AI" wasted as much time as it saved me.
- l1ng0 1y ago[dead]
- timschmidt 1y agoI've certainly gotten a lot of value from adapting my development practices to play to LLM's strengths and investing my effort where they have weaknesses. "You're using it wrong" and "It could work better than it does now" can be true at the same time, sometimes for the same reason.
- thundoe 1y agoWhich is true. Like launching a Ferrari at 200mph without steering doesn’t take anyone anywhere, it’s just a very painful waste of money
- ZeWaka 1y agoI find it quite funny that one of the users actually posted a fully AI-generated reply (dramatically different grammar and structure than their other posts).
- gwd 1y agoExactly one of two things is true: 1. The tool is capable of doing more than OP has been able to make it do 2. The tool is not capable of doing more than OP has been able to make it do. If #1 is true, then... he must be using it wrong. OP specifically said: > Please pour in your responses please. I really want to see how many people believe in agentic and are using it successfully So, he's specifically asking people to tell him how to use it "right".
- geldedus 1y agoYes. Because it is the correct answer.
- Julien_r2 1y agoI actually hope to find better answers here than on cursor forum where people seems to be basically saying "it's you fault" instead of answering the actual question which is about trust, process, and real world use of agents.. So far it's just reinforcing my feeling that none of this is actually used at scale.. We use AI as relatively dumb companions, let them go wilder on side projects which have loser constraints, and Agent are pure hype (or for very niche use cases)
- zwnow 1y agoExactly, the actual business value is way smaller people think and its honestly frustrating. Yes they can write boilerplate, yes they sometimes do better than humans in well understood areas. But its negligible considering all the huge issues that come with them. Big tech vendorlocks, data poisoning, unverifiable information, death of authenticity, death of creativity, ignorance of LLM evangelists, power hungriness in a time where humanity should look at how to decrease emissions, theft of original human work, theft of data big tech gets away with since way too long. Its puzzling to me how people actually think this is a net benefit to humanity.
- Zababa 1y agoMost of the issues you listed are moral and not technical. Especially "power hungriness in a time where humanity should look at how to decrease emissions", this may be what you think humanity should do but that is just that, what you think. I derive a lot of business value from them, many of my colleagues do too. Many programmers that were good at writing code by hand are having lots of success with them, for example Thorsten Ball, Simon Willison, Mitchell Hashimoto. A recent example from Mitchell Hashimoto: https://mitchellh.com/writing/non-trivial-vibing https://mitchellh.com/writing/non-trivial-vibing. >Its puzzling to me how people actually think this is a net benefit to humanity. I've used them personally to quickly spin up a microblog where I could post my travel pictures and thoughts. The idea of making the interface like twitter (since that's what I use and know) was from me, not wanting to expose my family and friends to any specific predatory platform like twitter, instagram, etc was also from me, supabase as the backend was from a colleague (helped a lot!), the code was all Claude. The result is that they were able to enjoy my website, including my grandparents that just had to paste an URL on the website. I like to think of it a a perhaps very small but net benefit for a very small part of humanity.
- JCM9 1y agoBecause the hype cycle on the original AI wave was fading so folks needed something new to buzz about to keep the hype momentum going. Seriously, that’s the reason.
- berkes 1y agoDo you have any reasoning or anything else to back this up? Edit: honest question, not a diss or a dismissal. It's an interesting take, one that I believe could be true, but it sounds more like an opinion than a thesis or even fact.
- JCM9 1y agoFolks aren’t seeing measurable returns on AI. Lots written about this. When the bean counters show up, the easiest way to get out of jail is to say “Oh X? Yeah that was last year, don’t worry about it… we’re now focused on Y which is where the impact will come from.” Every hype cycle goes through some variation of this evolution. As much as folks try to say AI is different it’s following the same very predictable hype cycle curve.
- Msurrow 1y agoThe Gartner Hype Cycle [1]. I wonder were AI should be put on the graph, here in the 2025 fall. Just past the peak? Or are we not there yet. [1]: https://en.wikipedia.org/wiki/Gartner_hype_cycle https://en.wikipedia.org/wiki/Gartner_hype_cycle
- eschaton 1y agoWhy do they need reasoning to back it up when the LLMs being promoted don’t actually do any reasoning?
- Havoc 1y agoSame with “context engineering”
- micoti 1y agoI have the exact same question, what is hype all about when models can't do simple things. You prompt the model with generate one unit test for function and it somehow always generate more then one. (Just to start with most simple instruction) I just feel that models are currently not up to speed with experienced engineers where it takes less time to develop something then to instruct model to do it. It is only usefull for boring work. This is not to say that these tools didn't created oportunities to create new stuff, it is just that the hype overestimates the usefullnes of the tools so they can sell them better just like all other things.
- ianpri11 1y agocompleting boring work is still very useful when a large proportion of peoples day jobs are managing CRUD apps
- micoti 1y agoi agree, these tools are usefull. i only oppose agresive marketing that llm is solution for everything. it is just a tool which has its use case, but to me it seems that it is not optimal for use cases that it is advertised. i work on agentic systems and they can be good if agent has a bit-sized chuck of work it needs to do. problme with the coding agents is that for every more complex thing you will need to write a big prompt which is sometimes counter productive and it seems to me that user in cursor thread is pointing in that direction.
- thor-rodrigues 1y agoI think what we should really ask ourselves is: “Why do LLM experiences vary so much among developers?” The simplest explanation would be “You’re using it wrong…”, but I have the impression that this is not the primary reason. (Although, as an AI systems developer myself, you would be surprised by the number of users who simply write “fix this” or “generate the report” and then expect an LLM to correctly produce the complex thing they have in mind.) It is true that there is an “upper management” hype of trying to push AI into everything as a magic solution for all problems. There is certainly an economic incentive from a business valuation or stock price perspective to do so, and I would say that the general, non-developer public is mostly convinced that AI is actually artificial intelligence, rather than a very sophisticated next-word predictor. While claiming that an LLM cannot follow a simple instruction sounds, at best, very unlikely, it remains true that these models cannot reliably deliver complex work.
- tovej 1y agoI would say they can't reliably deliver simple work. They often can, but reliability, to me, means I can expect it to work every time. Or at least as much as any other software tool, with failure rates somewhere in the vicinity of 1 in 10^5, 1 in 10^6. LLMs fail on the order of 1 in 10 times for simple work. And rarely succeed for complex work. That is not reliable, that's the opposite of reliable.
- krisoft 1y agoOne has to look at the alternatives. What would i do if not use the LLM to generate the code? The two answers are “coding myself”, “asking an other dev to code it”. And neither of those approach anywhere a 10^5 failure rate. Not even close.
- logicchains 1y ago>I think what we should really ask ourselves is: “Why do LLM experiences vary so much among developers?” Two of the key skills needed for effective use of LLMs are writing clear specifications (written communication), and management, skills that vary widely among developers.
- falconinthesun 1y ago"You're using it wrong" if a user cannot use a tool intuitively, the tool is not fit for purpose.
- coolfox 1y agoVC mumbo jumbo; you can apply this same logic to literally all of programming
- gwd 1y agoThe most powerful tools are usually renowned to have the most arcane user interfaces. Xkcd's "Uncomfortable Truths Well" said, "You will never find a programming language that frees you from the burden of clarifying your ideas." LLMs don't fundamentally change that dynamic. [1] https://xkcd.com/568/ https://xkcd.com/568/
- geldedus 1y agoThere are hundreds of tools that you must learn how to use to get any result. Totally fit for the purpose. Learn how to properly use that tool and you'll get results.
- mihau 1y agoIt feels to me that the OP on the forum expects this to work: "read this existing function, then read my mind and do stuff" (probably followed by "do better"). It still takes a lot of practice to get good at prompting, though.
- anal_reactor 1y agoLiterally my manager
- fabian2k 1y agoFor me, a big issue is that the performance of the AI tools varies enormously for different tasks. And it's not that predictable when it will fail, which does lead to quite a bit of wasted time. And while having more experience prompting a particular tool is likely to help here, it's still frustrating. There is a bit of overlap for the stuff you use agents and the stuff that AI is good at. Like generating a bunch of boilerplate for a new thing from scratch. That makes the agent mode more convenient for me to interact with AI for the stuff it's useful in my case. But my experience with these tools is still quite limited.
- ehnto 1y agoWhen it works well you both normalise your expectations, and expand your usage, meaning you will hit its limits, and be even more disappointed when it fails at something you've seen it do well before.
- tejtm 1y agoNigh thirty years ago when dabbling in AI I read a quote I will paraphrase as: "when you hear 'intelligent agent'; think 'trainable ant'"
- jstummbillig 1y agoSay more?
- drittich 1y agoLooks like it comes from Scientific American: https://spaf.cerias.purdue.edu/~spaf/Yucks/V5/msg00004.html https://spaf.cerias.purdue.edu/~spaf/Yucks/V5/msg00004.html
- tejtm 1y agogreat job digging up a reference! you must be a bot! :) I read it in a book on AI, unfortunately that aisle in my library is inaccessible due to piles of obsolete crap (I wish I were kidding). But hope springs eternal and if I get back and find it I will return here and add its deets to see if it was published before or after the SI article. got to help future agents from going astray ...
- drittich 1y agoWe all just feed the LLMs now.
- varun_chopra 1y agoMarketing is being done really well in 2025, with brands injecting themselves into conversations on Reddit, LinkedIn, and every other public forum. [1] CEOs, AI "thought leaders," and VCs are advertising LLMs as magic, and tools like v0 and Lovable as the next big thing. Every response from leaders is some variation of https://www.youtube.com/watch?v=w61d-NBqafM https://www.youtube.com/watch?v=w61d-NBqafM On the ground, we know that creating CLAUDE.md or cursorrules basically does nothing. It’s up to the LLM to follow instructions, and it does so based on RNG as far as I can tell. I have very simple, basic rules set up that are never followed. This leads me to believe everyone posting on that thread on Cursor is an amateur. Beyond this, if you’re working on novel code, LLMs are absolutely horrible at doing anything. A lot of assumptions are made, non-existent libraries are used, and agents are just great at using tokens to generate no tangible result whatsoever. I’m at a stage where I use LLMs the same way I would use speech-to-text (code) - telling the LLM exactly what I want, what files it should consider, and it adds _some_ value by thinking of edge cases I might’ve missed, best practices I’m unaware of, and writing better grammar than I do. Edit: [1] To add to this, any time you use search or Perplexity or what have you, the results come from all this marketing garbage being pumped into the internet by marketing teams.
- joshvince 1y ago> if you’re working on novel code, LLMs are absolutely horrible This is spot on. Current state-of-the-art models are, in my experience, very good at writing boilerplate code or very simple architecture especially in projects or frameworks where there are extremely well-known opinionated patterns (MVC especially). What they are genuinely impressive at is parsing through large amounts of information to find something (eg: in a codebase, or in stack traces, or in logs). But this hype machine of 'agents creating entire codebases' is surely just smoke and mirrors - at least for now.
- vallavaraiyan 1y agoWhat is novel code? 1. LLM's would suck at coming up with new algorithms. 2. I wouldn't let an LLM decide how to structure my code. Interfaces, module boundaries etc Other than that, given the right context (the sdk doc for a unique hardware for eg) and a well organised codebase explained using CLAUDE.Md they work pretty well in filling out implementations. Just need to resist the temptation to prompt while the actual typing would take seconds.
- mschuster91 1y agoBecause there is a lot of money tied up in AI now, in a way that doesn't just reek like a bubble waiting to implode but even more stinks like a bunch of what used to be called "wash trading" [1]. And that's just the money side. The "social kool-aid" side is even worse. A lot of very rich and very influential people have bet their career on AI - especially large companies who just outright fired staff to be replaced both by actual AI and "Actually Indians" [2] and are now putting insane pressure on their underlings and vendors to make something that at least looks on the surface like the promised AI dreams of getting rid of humans. Both in combination explains why there is so much half-baked barely tested garbage (or to use the term du jour: slop) being pushed out and force fed to end users, despite clearly not being ready for prime time. And on top of that, the Pareto principle also works for AI - most of what's being pushed is now "good enough" for 80%, and everyone is trying to claim and sell that the missing 20% (that would require a lot of work and probably a fundamentally new architecture other than RNG-based LLMs) don't matter. [1] https://www.bbc.com/news/articles/cz69qy760weo https://www.bbc.com/news/articles/cz69qy760weo [2] https://www.osnews.com/story/142488/ai-coding-chatbot-funded-by-microsoft-were-actually-indians/ https://www.osnews.com/story/142488/ai-coding-chatbot-funded...
- petetnt 1y agoI love how the proposed solution is to essentially gaslighting the model to think that it's an expert programmer and then specify and re-specify the prompt until the solution is essentially inefficient pseudocode. Now we are in a world where amateur coders still cannot code or can't learn from their mistakes while experts are essentially JIRA ticket outsourcing specialists.
- hansmayer 1y agoThe answer is really trivial and really embarrassingly simple, once you remove the engineering/functional/world improvement goggles. The answer is: because the rich folks invested a ton of money and they need it to work. Or at least to make most of the white collar work dependent on it, quality be damned. Hence the ever increasing pushing, nudging, advertising, offering to use the crap-tech everywhere. It seems now it will not win over the engineers. Unfortunately it seems to work with most of the general population. Every lazy recruiter out there is now using chatgpt to generate job summaries and "evaluate" candidates. Every "office worker" of the general type deadweight you meet at every company is happy to use it to produce more powerpoints, slides and documents for you drown in. And I won't even mention the "content" business model of the influencers.
- thaumasiotes 1y agoThis idea is getting a lot of attention right now. e.g. https://www.noahpinion.blog/p/americas-future-could-hinge-on-whether https://www.noahpinion.blog/p/americas-future-could-hinge-on...
- rokkamokka 1y agoI find it funny that the page subheader is "If the economy's single pillar goes down, Trump's presidency will be seen as a disaster". Is it not a disaster already? The fast slide towards autocracy should certainly be viewed as a disaster if nothing else.
- yoyohello13 1y agoThe fact that sending the US military to occupy our own cities is not seen as a failure is... something.
- deleted 1y ago[deleted]
- wkat4242 1y agoAt our place we have two types of users. One that is a deep evangelist, says it revolutionised their office work and has no idea that it might have accuracy problems. I guess those are the people that just create a lot of hot air. The others tried it and ran into the obvious Achilles heels and are now pretty cautious. But use it for a thing or two.
- deleted 1y ago[deleted]
- gwd 1y agoFWIW all my coding with LLMs is very hands-on. What I've ended up doing with LLMs is something like the following: 1. New conversation. Describe at a high level what change I want made. Point out the relevant files for the LLM to have context. Discuss the overall design with the LLM. At the end of that conversation, ask it to write out a summary (including relevant files to read for context next time) in an "epic" document in llm/epics/. This will almost always have several steps, listed in the document. Then I review this and make sure it's in line with what I want. 2. New conversation. We're working on @llm/epics/that_epic.md. Please read the relevant files for context. We're going to start work on step N. Let me know if you have any questions; when you're ready, sketch out a detailed plan of implementation. I may need to answer some questions or help it find more context; then it writes a plan. I review this plan and make sure it's in line with what I want. 3. New conversation. We're working on @llm/epics/that_epic.md. We're going to start implementing step N. Let me know if you have any questions; when you're ready, go ahead and start coding. Monitor it to make sure it doesn't get stuck. Any time it starts to do something stupid or against the pattern of what I'd like -- from style, to hallucinating (or forgetting) a feature of some sub-package -- add something to the context files. Repeat until the epic is done. If this sounds like a lot of work, it is. As xkcd's "Uncomfortable Truths Well" said, "You will never find a programming language that frees you from the burden of clarifying your ideas." LLMs don't fundamentally change that dynamic. But they do often come up with clever solutions to problems; their "stupid questions" often helps me realize how unclear my thinking is; they type a lot faster, and they look up documentation a lot faster too. Sure, they make a bunch of frustrating mistakes when they're new to the project; but if every time they make a patterned mistake, you add that to your context somehow, eventually these will become fewer and fewer.
- deleted 1y ago[deleted]
- AHTERIX5000 1y ago2025 was the year when my fear of being replaced by an AI changed to fear of a big economic disaster caused by AI bubble
- ehnto 1y agoIt's been a rollercoaster, and it's still not clear what's on the other side of the loop.
- thdhhghgbhy 1y agoFirst answer: "you're prompting it wrong." I've heard that a few times now about demented autocomplete.
- yeasku 1y agoYou are using the autocomplete wrong is kind of funny sentence.
- mnky9800n 1y agoAs a scientist there is a ton of boiler plate code that is just slightly different enough for every data set I need to write it myself each time. So coding agents solve a lot of that. At least until you are halfway through something and you realize Claude didn’t listen when you wrote 5 times in capital letters NEVER MAKE UP DATA YOU ARE NOT ALLOWED TO USE np.random IN PLACE OF ACTUAL DATA. It’s all kind of wild because when it works it’s great and when it doesnt there’s no failure state. So if I put on my llm marketing hat I guess the solution is to have an agent that comes behind the coding agent that checks to see if it does its job. We can call it the Performance Improvement Plan Agent (PIPA). PIPAs allow real time monitoring of coding agents to make sure they are working and not slacking off allowing for HR departments and management teams to have full control over their AI employees. Together we will move into the future.
- ghthor 1y agoPIPA scary
- theshrike79 1y agoQuick, don't think of an elephant in a pink tutu! You did, didn't you? As a scientist, you should know that LLMs are pretty bad at understanding negatives because they work on tokens, not words. "NO ELEPHANTS" roughly becomes NO + ELEPHANT. Now "elephant" is in the context and it's going to be "thinking" about it and steering everything towards it. You need to use positive instructions.
- jeswin 1y agoThere are widely divergent views here. It'd be hard to have a good discussion unless people mention what tasks they're attempting and failing at. And we'll also have to ask if those tasks (or categories) are representative of mainstream developer effort. Without mentioning what the LLMs are failing or succeeding at, it's all noise.
- danielbln 1y agoWe'd need: - language/framework - problem space/domain - SRE experience level - LLM (model/version) - agentic harness (claude code, codex, copilot, etc.) - observed failure modes or win states - experience wrangling these systems ("I touched ChatGPT once" vs "I spend 12h/day in Claude Code") And there's more, is the engineer working on a single codebase for 10 years or do they jump around various projects all the time. Is it more greenfield, or legacy maintenance. Is it some frontier never-before-seen research project or CRUD? And so on.
- throw-10-13 1y ago“Youre holding it wrong”
- mihaaly 1y agoWhat I recently experienced on asking for a string manipulation routine that follows very arbitrary logic (for a long existing file format) that it forgots things like UTF string handling (in general, but also its subtle details requiring second round), its own code replacing special characters with escape sequence can be cut in half in limited width fileds (being an input for the function), considers some aspects of the specification document while omiting the others. Needs heavy supervision in the details and constant adjustments. Yet, it makes the bulk of the work. Saves brain energy, that goes into the edge cases then. The overall time is the same, it is just the result could become more robust in the end. Only with good supervision! (which has better chance when we are not worn out with the tedious heavy lifting part) But the one undebatable benefit is that the user can feel the smartest person in the whole wide world having so 'excellent questions', and 'knowing the topic like a pro', or being 'fantastic to spot such subtle details'. Anyone feel inadequate should use an agentic AI to boost self morale! (well, only if the person does not get nauseous from that thick flattering)
- rel0gic 1y agoi just got an aneurysm from reading the comments over there. are people having a stroke?
- intended 1y agoI am going to try and make it a habit to post this request on all LLM Coding questions - Can we please make it a point to share the following information when we talk about experiences with code bots? 1) Language - gives us an idea if the language has a large corpus of examples or not 2) Project - what were you using it for? 3) Level of experience - neophyte coder? Dunning Krueger uncertainty? Experience in managing other coders? Understand project implementation best practices ? From what I can tell/suspect, these 3 features are the likely sources of variation in outcomes. I suspect level of experience is doing significant heavy lifting, because more experienced devs approach projects in a manner that avoids pitfalls from the get go.
- nurettin 1y agoAfter so many months, Gemini pro still shits the bed after failing to update a file several times. I'd expect more from the culmination of human knowledge.
- chrisjj 1y agoWhen I asked Claude "AI" to count the number of text file lines missing a given initial sub-string, it gave an improbably exaggerated result. When I challenged this, it replied "You are right! Let me try again this time without splitting long lines." AI = Absent Intelligence.
- etothet 1y agoI recommend you check out Andrej Karpathy’s 2 YouTube videos on how LLMs work (they are easy to find, but be forewarned they are long!). Once one digs in deeper it becomes clear why a model today might fail at the task you described. Generally speaking, one of the behaviors I see in my day to day work leading engineers is that they often attempt to apply agentic coding tools to problema that don’t really benefit from them.
- chrisjj 1y ago> I recommend you check out Andrej Karpathy’s 2 YouTube videos on how LLMs work I'll reccommend Claude "AI" do that, so it knows to tell the user when it /doesn't/ work.
- rootlocus 1y agoHere's how Claude Code does this: > Find all java files with more than 100 lines. ● I'll search for all Java files with more than 100 lines in the codebase. ● Bash(find . -name "*.java" -type f -exec wc -l {} + | awk '$1 > 100' | sort -nr) ... ● I found 5 Java files with more than 100 lines: 1. File1.java - 315 lines 2. File2.java - 156 lines 3. File3.java - 154 lines 4. File4.java - 130 lines 5. File5.java - 117 lines The largest file is File1.java with 315 lines. Or, if you want to count lines: > How many lines don't start with `import` in File1.java? ● Bash(grep -cv "^import " ./File1.java) ⎿ 287 ● There are 287 lines that don't start with import in File1.java.
- theshrike79 1y agoDon't ask a language model to do math, they're not very good at it. Next time ask it to write a script or a program to do it, it'll most likely one-shot it.
- stanac 1y agoThe problem in this case is that LLMs are bad with golang, I don't write go, I am guessing from my experience with kotlin. I mainly use kotlin (rest apis) and LLMs are often bad at writing it. They e.g. confuse mockk and mockito functions and then agent spiral into a never ending loop of guessing what's wrong and trying to fix it in 5 different ways. Instead I use only chat, validate every output and point out errors they introduce. On the other hand colleagues working with react and next have better experience with agents.
- weitendorf 1y agoBecause golang is so verbose LLMs are still extremely useful at using it without feeling like you're doing data entry or working as a typist. Converting JSON to a properly typed Go struct or doing the 3-5 things required to create+specify+send+deserialize/parse an HTTP request (with explicit error handling) is about 20x faster when I make an LLM do it, and it makes it so I don't dread or avoid those kinds of tasks.
- throw_m239339 1y agoManagement thinks a crutch can effectively replace people massively in sensitive knowledge work. When that crutch starts making errors that cost those businesses millions, or billions, well, hopefully management who implemented all that will get fired... Yes, LLM are useful, but they are even less trustworthy than real humans, and one needs actual people to verify their output, so when agents write 100K lines of code, they'll make mistakes, extremely subtle ones, and not the kind of mistake any human operator would make.
- taherchhabra 1y agoI don't think the models are dumb anymore, codex with gpt5 and claude code can design and build complex systems. The only thing is these models work great on greenfield projects. Legacy projects design evolves over a number of years and LLMs have hard time understanding those unwritten project design decisions
- Netcob 1y agoMy guess is that the reason why AI works bad for some people is the same reason why a lot of people make bad managers / product owners / team leads. Also the same reason why onboarding is atrocious in a lot of companies ("Here's your login, here's a link to the wiki that hasn't been updated since 2019, if you have any questions ask one of your very busy co-workers, they are happy to help"). You have to very good at writing tasks while being fully aware of what the one executing it knows and doesn't know. What agents can infer about a project themselves is even more limited than their context, so it's up to you to provide it. Most of them will have no or very limited "long-term" memory. I've had good experiences with small projects using the latest models. But letting them sift through a company repo that has been worked on by multiple developers for years and has some arcane structures and sparse documentation - good luck with that. There aren't many simple instructions to be made there. The AI can still save you an hour or two of writing unit tests if they are easy to set up and really only need very few source files as context. But just talking to some people makes it clear how difficult the concept of implicit context is. Sometimes it's like listening to a 4 year old telling you about their day. AI may actually be better at comprehending that sort of thing than I am. One criticism I do have of AI in its current state is that it still doesn't ask questions often enough. One time I forgot to fill out the description of a task - but instead of seeing that as a mistake it just inferred what I wanted from the title and some other files and implemented it anyway. Correctly, too. In that sense it was the exact opposite of what OP was complaining about, but personally I'd rather have the AI assume that I'm fallible instead of confidently plowing ahead.
- weitendorf 1y agoI fully agree with this take and think a lot of people at this point are really just being uncharitable to those using AI productively + unwilling to admit their own faults when they fail to see this. How can anybody who has managed or worked with inexperienced engineers, or StackOverflow developers, not see how helpful AI is for delegating the kinds of tasks with that particular flavor of content and scope? And how can anybody who is currently working with those kinds of developers not see how much it's helping them improve the quality of their work? (and yes, it's extremely frustrating to see AI used poorly or for people to submit code for review that they did not even review or even understand themselves. But the fact that that's even possible, that it often times still works, really tells you something... And given the right feedback, most offenders do eventually understand why they ought not to do this, I think) Even for more experienced engineers, for the kind of "unimportant / low priority, uninteresting" work that requires a lot of context and knowledge to get done but isn't really a good use of experienced engineers' time, AI can really lower the barrier to starting and completing those tasks. Let's say my codebase doesn't have any docstrings or unit tests - I can feed it into an LLM and immediately get mediocre versions of all of that and just edit it into being good enough to merge. Or let's say I have an annoying unicode parsing bug, a problem with my regex, or something like that which I can reproduce in tests or a dev environment: a lot of the time I can just give the LLM the part of the code I suspect the bug resides within, tell it what the bug symptoms are and ask it to fix it, and validate the fix. To be honest and charitable to those who do struggle to use AI this way, since it's most likely just a theory of mind issue (they don't understand what the AI does and doesn't know, and what context it needs to understand them and give them what they want), it could very well be influenced by being somewhere on the autism spectrum or just difficulty with social skills. Since AI is essentially a fresh wipe of the same stranger every time you start a conversation with it (unless you use "memory" features designed for consumer chat rather than coding), it never really gets to know you or understand your quirks like most people that regularly interact with those with social difficulties. So I suppose to a certain extent it requires them to "mask" or interact in a way they're unfamiliar with when dealing with computer tools. A lot of people for whatever reason seem also to have decided to become emotionally/personally invested in "AI stupid" to the point that they will just flat out refuse to believe there is value in being able to type some little compiler error or stacktrace into a textbox and 80% of the time get a custom fix in 10% of the time it would have taken to do the same thing on google search+stackoverflow.
- pjmlp 1y agoGot get some of that oil/diamants/...., that's why.
- vigouroustester 1y agoWhile I agree with the sentiment of not just letting it run free on the whole codebase and do what it wants, I still have good experience with letting it do small tasks one at a time, guided by me. Coding ability of models has really improved over the last few months itself and I seem to be clearing less and less AI-generated code mess than I was 5 months ago. It's got a lot to do with problem framing and prompt imo.
- the__alchemist 1y agoLike OP in the link, I'm confused too. And I use LLMs for coding every day! With precise prompts, function signatures provided, only using it for problems I know are solved [by others] etc.
- buzzin__ 1y agoIn machine learning, boosting is a way to combine weak learners into a strong one. Perhaps something similar can be done with language models?
- internet_points 1y agolook up Mixture of Experts, e.g. Mixtral
- ai-christianson 1y agoPeople want predictability from LLMs, but these things are inherently stochastic, not deterministic compilers. What’s working right now isn’t "prompting better," it’s building systems that keep the LLM on track over time: logging, retrying, verifying outputs, giving it context windows that evolve with the repo, etc. That’s why we’ve been investing so much in multi-agent supervision and reproducibility loops at gobii.ai. You can’t just "trust" the model; you need an environment where it’s continuously evaluated, self-corrects, and coordinates with other agents (and humans) around shared state. Once you do that, it stops feeling like RNG and starts looking like an actual engineering workflow, distributed between humans and LLMs.
- dankobgd 1y agobecause they invested billions and now they have to justify it
- deleted 1y ago[deleted]
- timcobb 1y ago> what prompted this post? well just tried to work with gpt5 and gemini pro That's the problem. GPT5, at least in Cursor, doesn't work for coding. It burns tokens and does actually nothing my in experience. Claude 3.5, 4 and 4.5, on the other hand, are pretty solid and make lots of forward progress with minimal instruction. It takes iteration, some skill, some critical thinking, and some hand coding! Yes, LLMs forget things and do random things sometimes, but for me it's a big boost.
- __MatrixMan__ 1y agoOnce you figure out how to get your model to go find the context it needs (for me this usually comes down to really good error messages that feel a bit like a prompt injection attack) and you figure out how to keep the tasks small and uniform-ish such that a passing test for a previous (supervised) task becomes a reason that that output can now be used as context for how to complete the next (unsupervised) task, agents can be pretty darn reliable. Maybe 50% of the problems we solve are repetitive enough for this to make sense, and 50% of those are unpredictable enough that a model in a loop isn't overkill compared to traditional automation, and 50% of those are too small to be worth investing in the necessary scaffolding. But if you're looking at a problem that's in that magical 12.5%, a properly constrained agent is absolutely the way to go.
- alganet 1y agoYou should do a video on that, like a live coding session. Not a tutorial, an organic recording of you being a badass context engineer. That would make it easier to get the message across.
- __MatrixMan__ 1y agoGood suggestion, thanks. I'll likely have to make a "what even is this" video for my coworkers, so maybe the video you're proposing would make a good Part II to that. Might be tricky to convince my company to bless its release but perhaps with some careful editing...
- alganet 1y agoYou could do it on spare time, not using your company's hardware, and work on a public open source repository. This way there's no conflict with potential contracts. Also, LLMs have been around for a while. Maybe you can just search for someone that did it and share one video that you would endorse as representative of what you believe to be good context engineering. It seems that there should be a lot of those around, lots of people are using this tech, aren't they?
- apriljo 1y agoI use git so hallucinogenic AI decisions are easy to revert. Why wouldn't I ask it to clean up years of tech debt while I work on something novel?
- jimkri 1y agoThe push for Agentic models is because people aren't working or they are failing to complete what they are contractually obligated to. I'm working with a client who has APEX code with 0% test coverage. The last consultant company deployed APEX without any test classes. That is a failure on the Salesforce partner, since you need test classes if you want to update any APEX code, and it leaves the client with more work and more costs. I used AI to write 100% test coverage in less than 1 hour. I had to give it direction because the first implementation was not to the Salesforce Dev Test Standards. So I downloaded the SFDC Developer PDF guides on APEX and gave it to AI Studio, and was able to write the test classes correctly. Agents will do what they are programmed and prompted to do, and this is really why corporations are moving in that direction.
- deleted 1y ago[deleted]
- yeasku 1y agoYou worked 1 hour on the testing and is finished? how big is the project and what programming languaje it uses? Has anybody reviewed all the code created for the test?
- jimkri 1y ago1 hour on writing the code, and then 30 minutes on testing. There were 3 tests in the class, the project isn't large, ~200 lines of code, and the point is that the team that deployed it failed to provide this. It's not complicated; it's a failure in process and work that they are required to provide based on their Salesforce Partner agreements, and to meet minimum standards. I already stated the language; it's written in Salesforce APEX. Yes, the code was reviewed and passed all tests within Salesforce as well. Salesforce partners often upload "Dummy" test methods or fail to provide the test classes, as this requires additional time. It's more a reflection on the partner than the individual programmer, since I've been in the industry, I know that devs are thrown onto multiple projects at once. However, with Agents, you can reduce that time. Nevertheless, companies are still failing to do even that. So that is why I think corporations are pushing for more agents; the basics are being skipped because the teams working on it don't have the time.
- nsonha 1y ago[dead]
- DarkNova6 1y agoWhat is the code quality of the average developer? What is the code quality of coding agents? There is your answer why many find AI coding productive while others do not.
- yeasku 1y agoIf you have the money you can get very good developers. If you want to pay 250 a month. You get what you get.
- righthand 1y agoAn agentic chat agent is a worker you can just immediately fire instead of training to be better. Even if you train it better, it’s best to just fire them once the project is done. Executives have invented firing virtual coworkers as a work game.
- sorcercode 1y ago> when we are yet to confidently have any model complete a single simple instruction??? i understand the author might be a little frustrated and employing hyperbole here, but are most folks genuinely having similar problems? at this point I have found LLMs to more often than not follow my instructions. It requires diligent pruning of instructions, effective prompting and planning. but once you get a sense of how to do those three things, it's possible to fly with these coding agents. it does get it wrong occasionally but anecdotally this is like 1/10 in my experience. and interrupting and course-correcting quickly gets me right back on track. I'm just surprised at the skepticism at the usefulness of these tools from the HN comments. there's plenty of reasons to be worried and upset (cost, job transformation and displacement etc) but the effectiveness of coding agents being a common theme, in the comments here as well, is surprising to me.
- true_religion 1y agoI'm not sure what people are considering instructions but it talks about the topics that I tell it to talk about, and when parsing prose it will take specific instruction as to word choice, or tone. This is true across Gemini, ChatGPT, and Qwen.
- bradfa 1y agoMy experience has been that if you take the time to explain what the current state is, what your desired state should be, and to give information on how you want the agent to proceed, that then you can work with the agent to craft a plan, refine the plan, and finally execute the plan. In this mode of operation, the current state of the art is quite impressive. You can't just give it a single sentence and expect it to do something complex correctly. It takes real effort and human time, just like if you were trying to get a smart and capable intern who has no real world experience to do something technical correctly. Just the AI agents work significantly faster than a human intern.
- qazxcvbnmlp 1y ago> My experience has been that if you take the time to explain what the current state is, what your desired state should be, and to give information on how you want the agent to proceed, I have a pet theory. 1. This skill requires a strong theory of mind[1]. 2. Theory of mind is more difficult in those with autism 3. The same autism that makes people really good at coding, and gives them the time to post on online forms like hn, makes it hard to understand how to work with LLMs and how others work with llms. To provide good context to the llm you need to have a good understanding of (1) what it will and will not know, (2) what you know and take for granted(ie a theory of your own mind) (3) what your expectations are. None of this you need to do when you are coding on your own, but are critical on getting a good response from the LLM. See also the black and white thinking that is common in the responses on articles like this.[2] [1]https://en.wikipedia.org/wiki/Theory_of_mind https://en.wikipedia.org/wiki/Theory_of_mind [2]https://www.simplypsychology.org/black-and-white-thinking-in-autism.html https://www.simplypsychology.org/black-and-white-thinking-in...
- Nifty3929 1y agoThis reads as someone who has not learned how, when and why to use a new tool, uses it poorly if at all, and then blames the tool. All the while their peers who have taken the time to understand the new tool and when/how/why to use it are in fact using it very successfully. Instead of thinking "it doesn't work," think "when and how can this work for me?" I had an experience recently where I was working on a python notebook that I didn't create, and which was a bit old. I ran into an error that I was unfamiliar with. There had been added an AI-based "explain this error" button to my notebook interface. I figured it would suck, but I gave it a try anyway. It correctly identified the problem (old library conflicting with newer one), suggested the change (switch from old library to newer fork) and offered to implement it. I clicked "yes" and it made all of the (relatively minor) changes required, including a few spots where some method returns had changed and needed to be unpacked. I scrutinized the changes to ensure I knew what was going on, but it all just worked regardless. Probably 8 or so 1-2 line changes throughout 200 lines of code. This likely saved me roughly an hour, or maybe two, of fiddling around, googling, reading docs, etc. Because I had an open mind and gave it a try, and had reasonable expectations, which were in fact exceeded. I did not start with "refactor my codebase please," and then get mad and give up when it doesn't help much.
- 0xbadcafebee 1y agoLLM is valuable as a research tool. For finding information, answering questions, doing comparisons, etc, it's great. Otoh, an LLM writing code is a dog painting. If you spend enough billions on it, a dog can churn out a Picasso. Doesn't mean it's a good idea.
- interleave 1y agoI'll bite. We're a classic XP shop. To build new features in our brown-field app, we defined about 8 sub-agents such as "red-test-writer", "minimal-green-implementer" and "refactorer". Now all I do in Claude Code is: "Build this feature X using our TDD process and the agents." 30 minutes later the feature is complete, looks better and works better than what I would have built in 30 minutes, is 90% tested and is ready for acceptance testing. Granted it took us years of working XP, pairing, TDD etc. but I keep feeling confused about posts like this. We've been shipping production-grade code written 95% by AI for over a year now. Non-trivial, complex features. There is no secret sauce even, in how we do this. It works. Really, really well for us.
- dearilos 1y agoThis is pretty much what I’m solving. LLMs need a “linter” of sorts that can guide them to write good code.
- hashstring 1y agoThe push seems to work marketing-wise for a lot of folks. I meet a lot of people who interfaced with LLMs for about topics that they are not an expert in, and they now believe that AI is taking over. What a bore.
- jMyles 1y agoA simple, interesting, and too-often-unspoken answer is that LLMs, especially with 100k+ context heaps, seem to be better at tasks that have contextual significance and scenery than they are at following one-off instructions. Or at least, that's very much the experience of my team. Like everyone else, I'm working on an MCP server right now. Mine though, is designed to populate a new context window with memories of Billy Strings shows and other bluegrass phenomena. Why? Well, in addition to that content being interesting to me and my wanting to see what LLMs have to say about it, I also notice that having this context seems to make them much better at solving bluegrass-adjacent problems in software, which is what I'm working on.
- j45 1y agoThe bringing of "agentic" into the mainstream this year was a little odd, crewAI was discovered by a broader group, but many people had quietly been playing with it for a while. The fundamentals of how LLMs are different from what came before it are still being understood. People keep thinking it's like the world they've always known with software and discovering it's not. That could be an issue with LLMs, or how/what they are used for. Seeing different ways takes time. I feel like we're still missing some essential building blocks in widespread use to not only creating agents but making sure they stay stable and consistent across model changes.
- throw-10-13 1y agoshades of “Self driving is coming next year”
- smartbench1 1y agoWith fine tuning becoming cheaper, agents will become very powerful in the next stage. That is an agent with multiple models as intelligence units in their areas of expertise.
- YouAreWRONGtoo 1y agoI think that unless you work on a 100000-1000000 GPU research cluster, you don't know what's currently possible. I wish I could know the kinds of queries that could be answered when there are no economic constraints on existing infrastructure. That would suggest whether they have already hit a scientific wall (and that's the difference between it being a $10T+ industry or a 500B industry). On consumer LLMs, it's still easy to get the LLM to admit queries are beyond its abilities, although many of those questions are also beyond 99.9999% of humanity, to be fair (in that the things I ask don't exist yet anywhere and possibly will never due to their non-trivial engineering nature).
- geldedus 1y agoIt's been months agentic AI writes thousands of lines of code that go to production. Use better tools.