20 ms·
LLMs work best when the user defines their acceptance criteria first
- marginalia_nu 7mo agoI tried to make Claude Code, Sonnet 4.6, write a program that draws a fleur-de-lis. No exaggeration it floundered for an hour before it started to look right. It's really not good at tasks it has not seen before.
- tartoran 7mo agoHave you tried describing to Claude what it is? The more the detail the better the result. At some point it does become easier to just do it yourself.
- vdfs 7mo agoMost people just forget to tell it "make it quick" and "make no mistake"
- marginalia_nu 7mo agoIt knows what it is, it's a very well known symbol. But translating that knowledge to code is something else. Interesting shortcoming, really shows how weak the reasoning is.
- cat_plus_plus 7mo agoTry writing code from description without looking at the picture or generated graphics. Visual LLM with a suggestion to find coordinates of different features and use lines/curves to match them might do better.
- parvardegr 7mo agoagreed with part that at some point it's better to just do it yourself but for sure they will get better and better
- comex 7mo agoLLMs are really bad at anything visual, as demonstrated by pelicans riding bicycles, or Claude Plays Pokémon. Opus would probably do better though.
- tartoran 7mo agoHow could they be any good at visuals? They are trained on text after all.
- msephton 7mo agoShapes can be described as text or mathematical formulas.
- comex 7mo agoSupposedly the frontier LLMs are multimodal and trained on images as well, though I don't know how much that helps for tasks that don't use the native image input/output support. Whatever the cause, LLMs have gotten significantly better over time at generating SVGs of pelicans riding bicycles: https://simonwillison.net/tags/pelican-riding-a-bicycle/ https://simonwillison.net/tags/pelican-riding-a-bicycle/ But they're still not very good.
- tartoran 7mo agoI have to admit I'm seeing this for the first time and am somewhat impressed by the results and even think they will get better with more training, why not... But are these multimodal LLMs still LLMs though? I mean, they're still LLMs but with a sidecar that does other things and the training of the image takes place outside the LLMs so in a way the LLMs still don't "know" anything about these images, they're just generating them on the fly upon request.
- boxedemp 7mo agoMaybe we should drop one of the L's
- 7mo ago
- jshmrsn 7mo agoConsidering that a fleur-de-lis involves somewhat intricate curves, I think I'd be pretty happy with myself if I could get that task done in an hour. Given a harness that allows the model to validate the result of its program visually, and given the models are capable of using this harness to self correct (which isn't yet consistently true), then you're in a situation where in that hour you are free to do some other work. A dishwasher might take 3 hours to do for what a human could do in 30 minutes, but they're still very useful because the machine's labor is cheaper than human labor.
- marginalia_nu 7mo agoI didn't provide any constraints on how to draw it. TBH I would have just rendered a font glyph, or failing that, grabbed an image. Drawing it with vector graphics programmatically is very hard, but a decent programmer would and should push back on that.
- zeroxfe 7mo ago> TBH I would have just rendered a font glyph, or failing that, grabbed an image. If an LLM did that, people would be all up in arms about it cheating. :-) For all its flaws, we seem to hold LLMs up to an unreasonably high bar.
- marginalia_nu 7mo agoThat's the job description for a good programmer though. Question assumptions and requirements, and then find the simplest solution that does the job. Just about anyone can eventually come up with a hideously convoluted HeraldicImageryEngineImplFactory<FleurDeLis>.
- ehnto 7mo agoEven with well understood languages, if there isn't much in the public domain for the framework you're using it's not really that helpful. You know you're at the edges of its knowledge when you can see the exact forum posts you are looking at showing up verbatim in it's responses. I think some industries with mostly proprietary code will be a bit disappointing to use AI within.
- internet2000 7mo agoI got Opus 4.6 to one shot it, took 5-ish mins. "Write me a python program that outputs an svg of a fleur-de-lis. Use freely available images to double check your work." It basically just re-created the wikipedia article fleur-de-lis, which I'm not sure proves anything beyond "you have to know how to use LLMs"
- robertcope 7mo agoSame, I used Sonnet 4.6 with the prompt, "Write a simple program that displays a fleur-de-lis. Python is a good language for this." Took five or six minutes, but it wrong a nice Python TK app that did exactly what it was supposed to.
- 64738 7mo agoJust for reference, Codex using GPT-5.4 and that exact prompt was a 4-shot that took ten minutes. The first result was a horrific caricature. After a slight rebuke ("That looks terrible. Read https://en.wikipedia.org/wiki/Fleur-de-lis https://en.wikipedia.org/wiki/Fleur-de-lis for a better understanding of what it should look like."), it produced a very good result but it then took two more prompts about the right side of the image being clipped off before it got it right.
- scuff3d 7mo agoI tried to use Codex to write a simple TCP to QUIC proxy. I intentionally kept the request fairly simple, take one TCP connection and map it to a QUIC connection. Gave a detailed spec, went through plan mode, clarified all the misunderstandings, let it write it in Python, had it research the API, had it write a detailed step by step roadmap... The result was a fucking mess. Beyond the fact that it was "correct" in the same way the author of the article talked about, there was absolutely bizarre shit in there. As an example, multiple times it tried to import modules that didn't exist. It noticed this when tests failed, and instead of figuring out the import problem it add a fucking try/except around the import and did some goofy Python shenanigans to make it "work".
- hrmtst93837 7mo ago[flagged]
- flerchin 7mo agoYes plausible text prediction is exactly what it is. However, I wonder if the author included benchmarking in their prompt. It's not exactly fair to keep hidden requirements.
- g947o 7mo agoAttributing these to "hidden requirements" is a slippery slope. My own experience using Claude Code and similar tools tells me that "hidden requirements" could include: * Make sure DESIGN.md is up to date * Write/update tests after changing source, and make sure they pass * Add integration test, not only unit tests that mock everything * Don't refactor code that is unrelated to the current task ... These are not even project/language specific instructions. They are usually considered common sense/good practice in software engineering, yet I sometimes had to almost beg coding agents to follow them. (You want to know how many times I have to emphasize don't use "any" in a TypeScript codebase?) People should just admit it's a limitation of these coding tools, and we can still have a meaningful discussion.
- flerchin 7mo agoYeah I agree generally that the most banal things must be specified, but I do think that a single sentence in the prompt "Performance should be equivalent" would likely have yielded better results.
- grey-area 7mo agoThe training data is full of ‘any’ so you will keep getting ‘any’ because that is the code the models have seen. An interesting example of the training data overriding the context.
- theshrike79 7mo agoThen you add a biome rule to say "no any ever" and the LLM will fix it before claiming the job is done.
- lukeify 7mo agoMost humans also write plausible code.
- deleted 7mo ago[deleted]
- tartoran 7mo agoLLMs piggyback on human knowledge encoded in all the texts they were trained on without understanding what they're doing. Humans would execute that code and validate it. From plausible it'd becomes hey, it does this and this is what I want. LLMs skip that part, they really have no understanding other than the statistical patterns they infer from their training and they really don't need any for what they are.
- owlninja 7mo agoThey probably at least look at the docs?
- deleted 7mo ago[deleted]
- stevenhuang 7mo agoLLMs can execute code and validate it too so the assertions you've made in your argument are incorrect. What a shame your human reasoning and "true understanding" led you astray here.
- red75prime 7mo agoCould we stop using vague terms like “understanding” when talking about LLMs and machine learning? You don't know what understanding is. You only know how it feels to understand something. It's better to describe what you can do that LLMs currently can't.
- stevenhuang 7mo agoAt least it's an easy way for those who don't know that they're talking about to out themselves. If they'd bother to see how modern neuroscience tries to explain human cognition they'd see it explained in terms that parallel modern ML. https://en.wikipedia.org/wiki/Predictive_coding https://en.wikipedia.org/wiki/Predictive_coding We only have theories for what intelligence even means, I wouldn't be surprised there are more similarities than differences between human minds and LLMs, fundamentally (prediction and error minimization)
- FrankWilhoit 7mo agoEnterprise customers don't buy correct code, they buy plausible code.
- kibwen 7mo agoEnterprise customers don't buy plausible code, they buy the promise of plausible code as sold by the hucksters in the sales department.
- marginalia_nu 7mo agoI think SolarWinds would have preferred correct code back in 2020.
- qup 7mo agoOkay, but what did they buy?
- marginalia_nu 7mo agoCode, from their employees.
- qup 7mo agoPlausible code from their employees.
- 2god3 7mo agoThey're not buying code. They are buying a service. As long as the service 'works' they do not care about the other stuff. But they will hold you liable when things go wrong. The only caveat is highly regulated stuff, where they actually care very much.
- cat_plus_plus 7mo agoThat's very impressive. Your LLM actually wrote a correct code for a full relational database on the first try, like it takes 2.5 seconds to insert 100 rows but it stores them correctly and select is pretty fast. How many humans can do this without a week of debugging? I would suggest you install some profiling tools and ask it to find and address hotspots. SQL Lite had how long and how many people to get to where it is?
- bluefirebrand 7mo agoI could "write" this code the same way, it's easy Just copy and paste from an open source relational db repo Easy. And more accurate!
- snoob2021 7mo agoIt is a Rust reimplementation of SQLite. Not exactly just "copy and paste"
- cat_plus_plus 7mo agoThe actual task is usually to mix something that looks like a dozen of different open source repos combined but to take just the necessary parts for task at hand and add glue / custom code for the exact thing being built. While I could do it, LLM is much faster at it, and most importantly I would not enjoy the task.
- thisguySPED 7mo ago[flagged]
- thisguySPED 7mo ago[flagged]
- comex 7mo agoBased on a search, the SQLite reimplementation in question is Frankensqlite, featured on Hacker News a few days ago (but flagged): https://news.ycombinator.com/item?id=47176209 https://news.ycombinator.com/item?id=47176209
- scotty79 7mo agoFlagging on HN is getting insane.
- mmaunder 7mo agoBut my AI didn't do what your AI did. Cherry picked AI fail for upvotes. Which you’ll get plenty of here an on Reddit from those too lazy to go and take a look for themselves. Using Codex or Claude to write and optimize high performance code is a game changer. Try optimizing cuda using nsys, for example. It’ll blow your lazy little brain.
- oofbey 7mo agoIt’s easy to get AI to write bad code. Turns out you still need coding skills to get AI to write good code. But those who have figured it out can crank out working systems at a shocking pace.
- serious_angel 7mo agoI am sorry for asking, but... is there guide even on how to "figure it out"? Otherwise, how are you so sure about it?
- mmaunder 7mo agoThat's actually a great question. Truth be told the best way right now is to grab Codex CLI or Claude CLI (I strongly prefer Codex, but Claude has its fans), and just start. Immediately. Then go hard for a few months and you'll develop the skills you need. A few tips for a quickstart: Give yourself permission to play. Understand basic concepts like context window, compaction, tokens, chain of thought and reasoning, and so on. Use AI to teach you this stuff, and read every blog post OpenAI and Anthropic put out and research what you don't understand. Pick a hard coding problem in Python or Typescript and take a leap of faith and ask the agent to code it for you. My favorite phrase when planning is: "Don't change anything. Just tell me.". Save this as a tmux shortcut and use it at the end of every prompt when planning something out. Use markdown .md docs to create a planning doc and keep chatting to the agent about it and have it update the plan until you're super happy, always using the magic phrase "Don't change anything. Just tell me." (I should get myself a patent on that little number. Best trick I know) Every time you see an anti-AI post, just move on. It's lazy people making lazy assumptions. Approach agentic coding with a sense of love, excitement, optimism, and take massive leaps of faith and you'll be very very surprised at what you find. Best of luck Serious Angel.
- deleted 7mo ago[deleted]
- gzread 7mo agoEarly LLMs would do better at a task if you prefixed the task with "You are an expert [task doer]"
- serious_angel 7mo agoHoly gracious sakes... Of course... Thank you... thank you... dear katanaquant, from the depths... of my heart... There's still belief in accountability... in fun... in value... in effort... in purpose... in human... in art... Related: - <http://archive.today/2026.03.07-020941/https://lr0.org/blog/p/gpt/ http://archive.today/2026.03.07-020941/https://lr0.org/blog/...> (I'm not consulting an LLM...) - <https://web.archive.org/web/20241021113145/https://slopwatch.com/posts/bad-programmer/ https://web.archive.org/web/20241021113145/https://slopwatch...>
- KatanaLarp 7mo ago[dead]
- skybrian 7mo agoYou can ask an LLM to write benchmarks and to make the code faster. It will find and fix simple performance issues - the low-hanging fruit. If you want it to do better, you can give it better tools and more guidance. It's probably a good idea to improve your test suite first, to preserve correctness.
- pornel 7mo agoTheir default solution is to keep digging. It has a compounding effect of generating more and more code. If they implement something with a not-so-great approach, they'll keep adding workarounds or redundant code every time they run into limitations later. If you tell them the code is slow, they'll try to add optimized fast paths (more code), specialized routines (more code), custom data structures (even more code). And then add fractally more code to patch up all the problems that code has created. If you complain it's buggy, you can have 10 bespoke tests for every bug. Plus a new mocking framework created every time the last one turns out to be unfit for purpose. If you ask to unify the duplication, it'll say "No problem, here's a brand new metamock abstract adapter framework that has a superset of all feature sets, plus two new metamock drivers for the older and the newer code! Let me know if you want me to write tests for the new adapters."
- stingraycharles 7mo ago> If you ask to unify the duplication, it'll say "No problem, here's a brand new metamock abstract adapter framework that has a superset of all feature sets, plus two new metamock drivers for the older and the newer code! Let me know if you want me to write tests for the new adapters." Nevermind the fact that it only migrated 3 out of 5 duplicated sections, and hasn’t deleted any now-dead code.
- Mavvie 7mo agoSounds like my coworkers.
- Foobar8568 7mo agoThat's the reality nobody really wants to say.
- Jweb_Guru 7mo agoIt's not reality. I'm really not a fan of the way that people excuse the really terrible code LLMs write by claiming that people write code just as bad. Even if that were true, it is not true that when you ask those people to do otherwise they simply pretend to have done it and forget you asked later.
- jqpabc123 7mo agoLLMs have no idea what "correct" means. Anything they happen to get "correct" is the result of probability applied to their large training database. Being wrong will always be not only possible but also likely any time you ask for something that is not well represented in it's training data. The user has no way to know if this is the case so they are basically flying blind and hoping for the best. Relying on an LLM for anything "serious" is a liability issue waiting to happen.
- retired_account 7mo agoIt’s a shame of bulk of that training data is likely 2010s blogspam that was poor quality to begin with.
- 2god3 7mo agoBut isn't that a reflection of reality? If you've made a significant investment in human capital, you're even more likely to protect it now and prevent posting valuable stuff on the web.
- 2god3 7mo agoAye. I wish more conversations would be more of this nature - in that we should start with basic propositions - e.g. the thing does not 'know' or 'understand' what correct is.
- LarsDu88 7mo agoThis is about to change very soon. Unlike many other domains (such as greenfield scientific discovery), most coding problems for which we can write tests and benchmarks are "verifiable domains". This means an LLM can autogenerated millions of code problem prompts, attempt millions of solutions (both working and non-working), and from the working solutions, penalize answers that have poor performance. The resulting synthetic dataset can then be used as a finetuning dataset. There are now reinforcement finetuning techniques that have not been incorporated into the existing slate of LLMs that will enable finetuning them for both plausibility AND performance with a lot of gray area (like readability, conciseness, etc) in between. What we are observing now is just the tip of a very large iceberg.
- ontouchstart 7mo agoI made a comment in another thread about my acceptance criteria https://news.ycombinator.com/item?id=47280645 https://news.ycombinator.com/item?id=47280645 It is more about LLMs helping me understand the problem than giving me over engineered cookie cutter solutions.
- graphememes 7mo agobad input > bad output idk what to say, just because it's rust doesn't mean it's performant, or that you asked for it to be performant. yes, llms can produce bad code, they can also produce good code, just like people
- jqpabc123 7mo agoyes, llms can produce bad code, they can also produce good code, just like people Over time, you develop a feel for which human coders tend to be consistently "good" or "bad". And you can eliminate the "bad". With an LLM, output quality is like a box of chocolates, you never know what you're going to get. It varies based on what you ask and what is in it's training data --- which you have no way to examine in advance. You can't fire an LLM for producing bad code. If you could, you would have to fire them all because they all do it in an unpredictable manner.
- graphememes 7mo agono but you're a human and you're responsible for it, so it's on you you can make horrible images with photoshop that doesn't make photoshop bad
- jqpabc123 7mo agoThe key word here is *you*. Photoshop doesn't make anything --- *you* make the image horrible --- or not. Any results relate directly to *your* skill. A direct comparison to agentic AI is less than equitable. AI is supposedly able to provide skill --- which it often fails to do.
- graphememes 7mo agoyou talk to the llm bro, you are responsible for the outcome
- codethief 7mo ago> Your LLM Doesn't Write Correct Code. It Writes Plausible Code. I don't always write correct code, either. My code sure as hell is plausible but it might still contain subtle bugs every now and then. In other words: 100% correctness was never the bar LLMs need to pass. They just need to come close enough.
- raw_anon_1111 7mo agoThe difference for me recently Write a lambda that takes an S3 PUT event and inserts the rows of a comma separated file into a Postgres database. Naive implementation: download the file from s3 and do a bulk insert - it would have taken 20 minutes and what Claude did at first. I had to tell it to use the AWS sql extension to Postgres that will load a file directly from S3 into a table. It took 20 seconds. I treat coding agents like junior developers.
- svpyk 7mo agoUnlike junior developers, llms can take detailed instructions and produce outstanding results at first shot a good number of times.
- raw_anon_1111 7mo agoWhile I’m pro LLMs over junior developers. The other issue with LLMs is even the most junior developer will learn your business context over time. In my case, in consulting (cloud + app dev), I just start the AGENTS.md file with a summary of the contract (the SOW), my architectural diagram and the transcript of my design review with the customer.
- conception 7mo agoDid you ask it to research best practices for this method, have an adversarial performance based agent review their approach or search for performant examples of the task first? Relying on training data only will always get your subpar results. Using “What is the most performant way to load a CSV from S3 into PostgreSQL on RDS? Compare all viable and research approaches before recommending one.” gave me the extension as the top option.
- raw_anon_1111 7mo agoI knew the best way. I was just surprised that Claude got it wrong. As soon as I told it to use the s3 extension, it knew to add the appropriate permissions, to update my sql unit script to enable the extension and how to write the code
- deleted 7mo ago[deleted]
- D-Machine 7mo agoThis article is great. And the blog-article headline is interesting, but wrong. LLM's don't in general write plausible code (as a rule) either. They just write code that is (semantically) similar to code (clusters) seen in its training data, and which haven't been fenced off by RLHF / RLVR. This isn't that hard to remember, and is a correct enough simplification of what generative LLMs actually do, without resorting to simplistic or incorrect metaphors.
- ozozozd 7mo agoExactly. It’s also easy to find yourself in the out-of-distribution territory. Just ask for some tree-sitter queries and watch Gemini 3, Opus 4.5 and GLM 5 hallucinate new directives.
- ehnto 7mo agoI think this could be the key difference in how people are experiencing the tools. Using Claude in industries full of proprietary code is a totally different experience to writing some React components, or framework code in C#, PHP or Java. It's shockingly good at the later, but as you get into proprietary frameworks or newer problem domains it feels like AI in 2023 again, even with the benefit of the agentic harnesses and context augments like memory etc.
- 2god3 7mo agoYou’ve hit the nail on the head. I characterise llm’s as being black boxes that are filled with a dense pool of digital resources - that with the correct prompt you can draw out a mix of resources to produce an output. But if the mix of resources you need isn’t there - it won’t work. This isn’t limited to just text. This also applies with video models - llms work better for prompts in which you are trying to get material that is widely available on the internet.
- simianwords 7mo agoAny example of how I can get it to hallucinate?
- user3939382 7mo agoI have great techniques to fix this issue but not sure how it behooves me to explain it.
- 88j88 7mo ago100% I found that you think you are smarter than the LLM and knowing what you want, but this is not the case. Give the LLM some leeway to come up with solution based on what you are looking to achieve- give requirements, but don't ask it to produce the solution that you would have because then the response is forced and it is lower quality.
- mirsadm 7mo ago100% dependent on the person driving it
- helsinki 7mo agoThat's why I added an invariant tool to my Go agent framework, fugue-labs/gollem: https://github.com/fugue-labs/gollem/blob/main/ext/codetool/invariants_tool.go#L89 https://github.com/fugue-labs/gollem/blob/main/ext/codetool/...
- deleted 7mo ago[deleted]
- seanmcdirmid 7mo agoI'm using an LLM to write queries ATM. I have it write lots of tests, do some differential testing to get the code and the tests correct, and then have it optimize the query so that it can run on our backend (and optimization isn't really optional since we are processing a lot of rows in big tables). Without the tests this wouldn't work at all, and not just tests, we need pretty good coverage since if some edge case isn't covered, it likely will wash out during optimization (if the code is ever correct about it in the first place). I've had to add edge cases manually in the past, although my workflow has gotten better about this over time. I don't use a planner though, I have my own workflow setup to do this (since it requires context isolated agents to fix tests and fix code during differential testing). If the planner somehow added broad test coverage and a performance feedback loop (or even just very aggressive well known optimizations), it might work.
- STARGA 7mo ago[dead]
- deleted 7mo ago[deleted]
- bamboozled 7mo agoI'm sure this is because they are pattern matching masters, if you program them to find something, they are good at that. But you have to know what you're looking for.
- gormen 7mo agoExcellent article. But to be fair, many of these effects disappear when the model is given strict invariants, constraints, and built-in checks that are applied not only at the beginning but at every stage of generation.
- riffraff 7mo agoTo be fair, people do too.
- sim04ful 7mo agoI've noticed a key quality signal with LLM coding is an LOC growth rate that tapers off or even turns negative.
- nprateem 7mo agoIn the last month I've done 4 months of work. My output is what a team of 4 would have produced pre-AI (5 with scrum master). Just like you can't develop musical taste without writing and listening to a lot of music, you can't teach your gut how to architect good code without putting in the effort. Want to learn how to 10x your coding? Read design patterns, read and write a lot of code by hand, review PRs, hit stumbling blocks and learn. I noticed the other day how I review AI code in literally seconds. You just develop a knack for filtering out the noise and zooming in on the complex parts. There are no shortcuts to developing skill and taste.
- allajfjwbwkwja 7mo ago> I review AI code in literally seconds You've just settled for hackathon standards and told yourself it's okay because you're using AI. Everyone with experience should know that even thorough code reviews only catch stylistic issues, glaring errors, and the most obvious design deficiencies. The only time new code is truly thought about is as it's being written.
- jeff_antseed 7mo ago[dead]
- einrealist 7mo ago> SQLite is not primarily fast because it is written in C. Well.. that too, but it is fast because 26 years of profiling have identified which tradeoffs matter. Someone (with deep pockets to bear the token costs) should let Claude run for 26 months to have it optimize its Rust code base iteratively towards equal benchmarks. Would be an interesting experiment. The article points out the general issue when discussing LLMs: audience and subject matter. We mostly discuss anecdotally about interactions and results. We really need much more data, more projects to succeed with LLMs or to fail with them - or to linger in a state of ignorance, sunk-cost fallacy and supressed resignation. I expect the latter will remain the standard case that we do not hear about - the part of the iceberg that is underwater, mostly existing within the corporate world or in private GitHubs, a case that is true with LLMs and without them. In my experience, 'Senior Software Engineer' has NO general meaning. It's a title to be awarded for each participation in a project/product over and over again. The same goes for the claim: "Me, Senior SWE treat LLMs as Junior SWE, and I am 10x more productive." Imagine me facepalming every time.
- grey-area 7mo agoThis would be a really interesting experiment. I suspect performance is not the only problem with the codebase though.
- KatanaLarp 7mo ago[dead]
- genie3io 7mo ago[dead]
- grey-area 7mo agoThis is a fascinating look into code generated by an LLM that is correct in one sense (passes tests) but doesn't meet requirements (painfully slow). Doesn't use is_ipk to identify primary keys, uses fsync on every statement. The problem with larger projects like this even if you are competent is that there are just too many lines of code to read it properly and understand it all. Bravo to the author for taking the time to read this project, most people never will (clearly including the author of it). I find LLMs at present work best as autocomplete - The chunks of code are small and can be carefully reviewed at the point of writing Claude normally gets it right (though sometimes horribly wrong) - this is easier to catch in autocomplete That way they mostly work as designed and the burden on humans is completely manageable, plus you end up with a good understanding of the code generated. They make mistakes I'd say 30% of the time or so when autocompleting, which is significant (mistakes not necessarily being bugs but ugly code, slow code, duplicate code or incorrect code. Having the AI produce the majority of the code (in chats or with agents) takes lots of time to plan and babysit, and is harder to review, maintain and diagnose; it doesn't seem like much of a performance boost, unless you're producing code that is already in the training data and just want to ignore the licensing of the original code.
- theshrike79 7mo ago> This is a fascinating look into code generated by an LLM that is correct in one sense (passes tests) but doesn't meet requirements (painfully slow). Why isn't requirements testing automated? Benchmarking the speed isn't rocket science. At worst a nightly build should run a benchmark and log it so you can find any anomalies.
- KatanaLarp 7mo ago[dead]
- mentalgear 7mo ago> I write this as a practitioner, not as a critic. After more than 10 years of professional dev work, I’ve spent the past 6 months integrating LLMs into my daily workflow across multiple projects. LLMs have made it possible for anyone with curiosity and ingenuity to bring their ideas to life quickly, and I really like that! But the number of screenshots of silently wrong output, confidently broken logic, and correct-looking code that fails under scrutiny I have amassed on my disk shows that things are not always as they seem. Same experience, but the hype bros do only need a shiny screengrab to proclaim the age of "gatekeeping" SWE is over to get their click fix from the unknowingly masses.
- spullara 7mo agohuman developers work best when the user defines their acceptance criteria first.
- KatanaLarp 7mo ago[dead]
- consumer451 7mo agoNitpick/question: the "LLM" is what you get via raw API call, correct? If you are using an LLM via a harness like claude.ai, chatgpt.com, Claude Code, Windsurf, Cursor, Excel Claude plug-in, etc... then you are not using an LLM, you are using something more, correct? An example I keep hearing is "LLMs have no memory/understanding of time so ___" - but, agents have various levels of memory. I keep trying to explain this in meetings, and in rando comments. If I am not way off-base here, then what should be the term, or terms, be? LLM-based agents?
- xlth 7mo agoYou're not off-base at all. The way I think about it: - LLM = the model itself (stateless, no tools, just text in/text out) - LLM + system prompt + conversation history = chatbot (what most people interact with via ChatGPT, Claude, etc.) - LLM + tools + memory + orchestration = agent (can take actions, persist state, use APIs) When someone says "LLMs have no memory" they're correct about the raw model, but Claude Code or Cursor are agents - they have context, tool access, and can maintain state across interactions. The industry seems to be settling on "agentic system" or just "agent" for that last category, and "chatbot" or "assistant" for the middle one. The confusion comes from product names (ChatGPT, Claude) blurring these boundaries - people say "LLM" when they mean the whole stack.
- dragonwriter 7mo ago> Nit pick/question: The LLM is what you get via raw API call, correct? You always need a harness of some kind to interact with an LLM. Normal web APIs (especially for hosted commercial systems) wrapped around LLMs are non-minimal harnesses, that have built in tools, interpretation of tool calls, application of what is exposed in local toolchains as “prompt templates” to transform the context structure in the API call into a prompt (in some cases even supporting managing some of the conversation state that is used to construct the prompt on the backend.) > If you are using an LLM via a harness like claude.ai, chatgpt.com, Claude Code, Windsurf, Cursor, Excel Claude plug-in, etc... then you are not using an LLM, you are using something more, correct? You are essentially always using something more than an LLM (unless “you” are the person writing the whole software stack, and the only thing you are consuming is the model weights, or arguably a truly minimal harness that just takes setting and a prompt that is not transformed in any way before tokenization, and returns the result after no transformations or filtering other than mapping back from tokens to text.) But, yes, if you are using an elaborate frontend of the type you enumerate (whether web or CLI or something else), you are probably using substantially more stuff on top of the LLM than if you are using the providers web API.
- alexhans 7mo ago> The vibes are not enough. Define what correct means. Then measure. Pretty much. I've been advocating this for a while. For automation you need intent, and for comparison you need measurement. Blast radius/risk profile is also important to understand how much you need to cover upfront. The Author mentions evaluations, which in this context are often called AI evals [1] and one thing I'd love to see is those evals become a common language of actually provable user stories instead of there being a disconnect between different types of roles, e.g. a scientist, a business guy and a software developer. The more we can speak a common language and easily write and maintain these no matter which background we have, the easier it'll be to collaborate and empower people and to move fast without losing control. - [1] https://ai-evals.io/ https://ai-evals.io/ (or the practical repo: https://github.com/Alexhans/eval-ception https://github.com/Alexhans/eval-ception )
- dillonsmartdev 7mo agoHumans work best like this too
- JasonHEIN 7mo agoBro you are like saying "OH LLM can't do X within 10 days which few people spend over decades" Live a life bro applause and change the title to "it can do xyz" instead of adding the "critical and critical" ...
- swiftcoder 7mo agoWhat's up with the (somewhat odd) title HN has gone with for this article? it's implying a very different article than the one I just read
- akoboldfrying 7mo agoThe following paragraph appears twice: > Now 2 case studies are not proof. I hear you! When two projects from the same methodology show the same gap, the next step is to test whether similar effects appear in the broader population. The studies below use mixed methods to reduce our single-sample bias.
- ollybrinkman 7mo agoThis maps directly to the shift happening in API design for agent-to-agent communication. Traditional API contracts assume a human reads docs and writes code once. But when agents are calling agents, the "contract" needs to be machine-verifiable in real-time. The pattern I've seen work: explicit acceptance criteria in API responses themselves. Not just status codes, but structured metadata: "This response meets JSON Schema v2.1, latency was 180ms, data freshness is 3 seconds." Lets the calling agent programmatically verify "did I get what I paid for?" without human intervention. The measurement problem becomes the automation problem. Similar to how distributed systems moved from "hope it works" to explicit SLOs and circuit breakers. Agents need that, but at the individual request level.
- jt2190 7mo agoInteresting, but couldn’t the agent be given access to tools that allow it to make those evaluations without having to modify the API responses? (Maybe I’m not visualizing “API” the same way you are.)
- pmarreck 7mo agoYes, which is why TDD is finally necessary
- teucris 7mo agoThis article hits on an important point not easily discerned from the title: Sometimes good software is good due to a long history of hard-earned wins. AI can help you get to an implementation faster. But it cannot magically summon up a battle-hardened solution. That requires going through some battles. Great software takes time.
- newzino 7mo ago[flagged]
- vicchenai 7mo ago[dead]
- shablulman 7mo ago[flagged]
- treetalker 7mo agoThis is my experience with how LLMs "draft" legal arguments: at first glance, it's plausible — but may be, and often is, invalid, unsound, and/or ill-advised. The catch is that many judges lack the time, energy, or willingness to not only read the documents in detail, but also roll up their sleeves and dig into the arguments and cited authorities. (Some lack the skills, but those are extreme cases.) So the plausible argument (improperly and unfortunately) carries the day. LLM use in litigation drafting is thus akin to insurgent/guerilla warfare: it take little time, energy, or thinking to create, yet orders of magnitude more to analyze and refute. (It's a species of Brandolini's Law / The Bullshit Asymmetry Principle.) Thus justice suffers. I imagine that this is analogous to the cognitive, technical, and "sub-optimal code" debt that LLM-produced code is generating and foisting upon future developers who will have to unravel it.
- FpUser 7mo ago>" justice suffers" Possible. It also suffers when majority simply can not afford proper representation
- deaux 7mo ago> This is my experience with how LLMs "draft" legal arguments: at first glance, it's plausible — but may be, and often is, invalid, unsound, and/or ill-advised. Correct, and this of course extends past just laws, into the whole scope of rules and regulations described in human languages. It will by its nature imply things that aren't explicitly stated nor can be derived with certainty, just because they're very plausible. And those implications can be wrong. Now I've had decent success with having LLMs then review these LLM-generated texts to flag such occurences where things aren't directly supported by the source material. But human review is still necessary. The cases I've been dealing with are also based on relatively small sets of regulations compared the scope of the law involved with many legal cases. So I imagine that in the domain you're working on, much more needs flagging.
- deleted 7mo ago[deleted]
- roarcher 7mo ago> LLM use in litigation drafting is thus akin to insurgent/guerilla warfare: it take little time, energy, or thinking to create, yet orders of magnitude more to analyze and refute. The same goes for coding. I have coworkers who use it to generate entire PRs. They can crank out two thousand lines of code that includes tests "proving" that it works, but may or may not actually be nonsense, in minutes. And then some poor bastard like me has to spend half a day reviewing it. When code is written by a human that I know and trust, I can assume that they at least made reasonable, if not always correct, decisions. I can't assume that with AI, so I have to scrutinize every single line. And when it inevitably turns out that the AI has come up with some ass-backwards architecture, the burden is on me to understand it and explain why it's wrong and how to fix it to the "developer" who hasn't bothered to even read his own PR. I'm seriously considering proposing that if you use AI to generate a PR at my company, the story points get credited to the reviewer.
- seanmcdirmid 7mo agoOk, I’ll bite: how is that different from humans?
- strken 7mo agoHuman behaviour is goal-directed because humans have executive function. When you turn off executive function by going to sleep, your brain will spit out dreams. Dream logic is famous for being plausible but unhinged. I have the feeling that LLMs are effectively running on dream logic, and everything we've done to make them reason properly is insufficient to bring them up to human level.
- whoamii 7mo agoSome of my best code comes from my dreams though.
- satvikpendem 7mo agoA prompt for an LLM is also a goal direction and it'll produce code towards that goal. In the end, it's the human directing it, and the AI is a tool whose code needs review, same as it always has been.
- basch 7mo agoId argue humans have some sort of parallelness going on that machines dont yet. Thoughts happening at multiple abstraction levels simultaneously. As I am doing something, I am also running the continuous improvement cycle in my head, at all four steps concurrently. Is this working, is this the right direction, does this validate? You could build layers and layers of LLMs watching the output of each others thoughts and offering different commentary as they go, folding all the thoughts back together at the end. Currently, a group of agents acts more like a discussion than something somewhat omnipotent or omnitemporal.
- spiderfarmer 7mo agoAnd yet LLM’s are incredibly useful as they are right now.
- bitwize 7mo agoYou: Claude, do you know how to program? Claude: No, but if you hum a few bars I can fake it! Except "faking it" turns out to be good enough, especially if you can fake it at speed and get feedback as to whether it works. You can then just hillclimb your way to an acceptable solution.
- andai 7mo agoIterative Faking™ — now with plausible-looking test suite!
- satvikpendem 7mo agoOftentimes, plausible code is good enough, hence why people keep using AI to generate code. This is a distinction without a difference.
- andai 7mo agoThere appears to be a similar approach in UX... plausible user experience is close enough.
- satvikpendem 7mo agoYes, especially because in UX there is no "correct" approach to it, it's all relative.
- bluetomcat 7mo agoNo. Plausible code is syntactically-correct BS disguised as a solution, hiding a countless amount of weird semantic behaviours, invariants and edge cases. It doesn't reflect a natural and common-sense thought process that a human may follow. It's a jumble of badly-joined patterns with no integral sense of how they fit together in the larger conceptual picture.
- satvikpendem 7mo agoWhy do people keep insisting that LLMs don't follow a chain of reasoning process? Using the latest LLMs you can see exactly what they "think" and see the resultant output. Plausible code does not mean random code as you seem to imply, it means...code that could work for this particular situation.
- tovej 7mo agoBecause they don't. The chain-of-reasoning feature is really just a way to get the LLM to prompt more. The fact that it generates these "thinking" steps does not mean it is using them for reasoning. It's most useful effect is making it seem to a human that there is a reasoning process.
- andai 7mo agoIt writes statistically represented code, which is why (unless instructed otherwise) everything defaults to enterprisey, OOP, "I installed 10 trendy dependencies, please hire me" type code.
- ZeroGravitas 7mo agoDoes it work if you get the agent to throw away all of its actual implementation and start again from scratch, keeping all the learning and tests and feedback? Gemini seems to try to get a lot of information upfront with questions and plans but people are famously bad at knowing what they want. Maybe it should build a series of prototypes and spikes to check? If making code is cheap then why not?
- freedomben 7mo agoThis does work but it requires prompts to instruct on it. It's also not perfect, though it is pretty good. What I've found when doing exactly this, is that the cost of the initial code makes me hesitant to throw it away. A better workflow I've been using is instead to iterate on very detailed planning documents written in markdown and repeatedly iterating on that instead (like, sometimes 50+ times for a complex app). It's really quite amazing how much that helps. It can lead to a design doc that is good enough that I can turn the agent loose on implementation and get decent results. Best results are still with guidance throughout, but I have never once regretted hammering out a very detailed planning document. I have many times regretted keeping code (or throwing code away).
- siliconc0w 7mo agoJust a recent anecdote, I asked the newest Codex to create a UI element that would persist its value on change. I'm using Datastar and have the manual saved on-disk and linked from the AGENTS.md. It's a simple html element with an annotation, a new backend route, and updating a data model. And there are even examples of this elsewhere in the page/app. I've asked it to do why harder things so I thought it'd easily one-shot this but for some reason it absolutely ate it on this task. I tried to re-prompt it several times but it kept digging a hole for itself, adding more and more in-line javascript and backend code (and not even cleaning up the old code). It's hard to appreciate how unintuitive the failure modes are. It can do things probably only a handful of specialists can do but it can also critical fail on what is a straightforward junior programming task.
- deleted 7mo ago[deleted]
- maremmano 7mo agothis won't age well.
- seba_dos1 7mo agos/code/stuff/
- jswelker 7mo agoI also write plausible code. Not much of a moat.
- giancarlostoro 7mo agoThis is why I used to use Beads and now GuardRails (shameless plug[0]). You brain dump to the model what you want, it breaks it down into discrete tasks, you have it refine them with you. By the time you have the model work on everything it can spawn workers in parallel that know what to do. In hindsight I should have called it BrainDump. [0]: https://giancarlostoro.com/introducing-guardrails-a-new-coding-agent-task-companion https://giancarlostoro.com/introducing-guardrails-a-new-codi...
- thrill 7mo agoIncreasing plausibility tends towards correctness.
- msvana 7mo agoI think there is one problem with defining acceptance criteria first: sometimes you don't know ahead of time what those criteria are. You need to poke around first to figure out what's possible and what matters. And sometimes the criteria are subjective, abstract, and cannot be formally specified. Of course, this problem is more general than just improving the output of LLM coding tools
- plandis 7mo agoYeah it’s extremely helpful to clarify your thoughts before starting work with LLM agents. I find Claude Code style plan mode to be a bit restrictive for me personally, but I’ve found that creating a plan doc and then collaboratively iterating on it with an LLM to be helpful here. I don’t really find it much different than the scoping I’d need to do before handing off some work to a more junior engineer.
- ramoz 7mo ago> Claude Code style plan mode to be a bit restrictive Hey thats why i built plannotator: https://github.com/backnotprop/plannotator https://github.com/backnotprop/plannotator I like staying within Claude Code for orchestrating its plan mode, but I needed a better way to actually review the plan, address certain parts, see plan diffs, etc all in a better visual way. The hooks system through permissionrequest:exitplanmode keep this fairly ergonomic. see it in action: https://www.youtube.com/watch?v=a_AT7cEN_9I https://www.youtube.com/watch?v=a_AT7cEN_9I
- deleted 7mo ago[deleted]
- arikrahman 7mo agoUncle Bob made this concept clear to me when he introduced to me that code itself IS requirements specification. LLMs are the new intermediary, but the necessity of the word and the machine persists.
- plandis 7mo agoI’ve found this to be critical for having any chance of getting agents to generate code that is actually usable. The more frequently you can verify correctness in some automated way the more likely the overall solution will be correct. I’ve found that with good enough acceptance criteria (both positive and negative) it’s usually sufficient for agents to complete one off tasks without a human making a lot of changes. Essentially, if you’re willing to give up maintainability and other related properties, this works fairly well. I’ve yet to find agents good enough to generate code that needs to be maintained long term without a ton of human feedback or manual code changes.
- malkia 7mo agoAre we now at the bottom of the the Uncanny Valley of AI?
- worik 7mo agoThis is becoming clear, now? I have had similar experiences, and I read over and over others experiences like this. A powerful tool...
- jbergqvist 7mo agoProducing the most plausible code is literally encoded into the cross entropy loss function and is fundamental to the pre-training. I suppose post training methods like RLVR are supposed to correct for this by optimizing correctness instead of plausibility, but there are probably many artifacts like these still lurking in the model's reasoning and outputs. To me it seems at least possible that the AI labs will find ways to improve the reward engineering to encourage better solutions in the coming years though.
- geysersam 7mo agoThere's also such a thing as being too ambitious. 99% of developers can not rewrite SQLite in rust even if they spent the rest or their lifetime doing it. Expecting an AI do to a good job vibe-coding a Sqllite clone over a few weekends just isn't realistic. Despite that, it's useful technology.
- namuol 7mo agoThese LLM prompting tip articles write themselves if you just take the last decade of project management articles and replace “IC” with “agent”.
- cadamsdotcom 7mo agoThis is a bit unfair - to generate a bunch of code but not give the model data/tools and direct it to optimize it; then compare it to the optimized work of thousands over decades. Feels like an extremely high effort hit piece, even though I know it’s not.
- KatanaLarp 7mo ago[dead]
- jamesblonde 7mo agoThe reference in the text to Anthropic’s “Towards Understanding Sycophancy in Language Models” is related to RLHF (reinforcement learning with human feedback). Claude code uses primarily different "pathways" in Anthropic LLMs that were not post-trained with RLHF, but rather with RLVF (reinforcement learning with verifiable rewards). So, his point about code being produced to please the user isn't valid from where I am sitting.
- KatanaLarp 7mo ago[dead]
- devonkelley 7mo ago[dead]
- nickcoffee 7mo agoThe acceptance criteria point translates directly outside of coding too. Using Claude Code for sales and operational workflows, having acceptable criteria upfront (along with some manual checks along the way depending on the task) definitely helps the output.
- Shyaamal11 7mo agoOne thing I’ve noticed while working with data/AI workflows is that the “acceptance criteria first” idea applies even more strongly once you move beyond code generation into data pipelines and analytics. LLMs can generate queries, transformations, or even Spark jobs that look reasonable but if the underlying data contracts, schema expectations, or evaluation criteria aren’t defined, you end up with something that looks correct but is semantically wrong. In practice, the teams that get the most value from AI-assisted development tend to have: clearly defined datasets reproducible data pipelines well-defined outputs / metrics Once those pieces are in place, AI becomes much more useful because it’s operating inside a structured system instead of guessing context. That’s also why there’s been a lot of interest lately in lakehouse-style platforms that combine data engineering, analytics, and AI workflows in one place (e.g. platforms like IOMETE). When the data layer is structured and reproducible, AI tooling becomes far more reliable. Curious if others here have seen the same pattern when using LLMs for data engineering or analytics work.