13 ms·
Software Engineering fundamentals matter more
- pavelevst 2mo agoAwesome!
- hirvi74 2mo ago> In the past year, agent harnesses crossed the “can it be done” rubicon. Brother, I'm still in "Can you get it right?"-mode. What am I doing wrong? (Rhetorical, but advice welcomed).
- MattGaiser 2mo agoWhat is “it” specifically and what languages are you using?
- hirvi74 2mo ago"It" is a lot of things. Languages: - C#/.NET: Sufficient sometimes, but not how I'd write things. Most results at least compile, but I have noticed plenty of defiance towards particular instructions, e.g, "Do not use <x>, use <y>" -> code contains <x> and not <y>. - C#/Godot: I have noticed the greatest amounts of defiance here. Not to mention most results are an 80/20 implementation of what I asked for. And no, I am not trying to one-shot a full game or anything. - AArch64 and x86: great results surprisingly, though only small amounts were produced. Mainly, assistance with RE-ing and cracking some binaries from https://crackmes.one https://crackmes.one or where ever. - The Lord's Language (Swift): Maybe the LLMs are better at SwiftUI/Swift, but I have had some rough results going down the opposite direction of the software stack. I have on/off been working on a personal, FOSS "productivity" tool for macOS, e.g., mouseless navigation, window management, GUI automation, etc.. This type of development requires a significant amount work with C APIs like CoreGraphics, Accessibility, CoreFoundation, etc.. The code isn't the problem for me, it's the lack of useful debugging. LLMs, last I have tried (around Opus 4.6 times), seemed to really struggle with things like CoreGraphics Y-axis coordinates being inverted compared AppKit's and other stuff like that. - Applescript (GUI automation): Do not even waste your time trying (I fault no LLMs for this either). - elisp: the code is usually sufficient, though package config can be a little dicey. - Shell scripts (Zsh, Bash, Powershell): great results. - Python: I try to avoid this language unless necessary, but the results have been great. These days, I use the plain Web chat interfaces for about 95% of my usage compared to the CLI harnesses. Sometime ago, I realized I get better results that way. With the web chat, I would say my results have been outstanding.
- jaggederest 2mo agoDirect feedback: You have to give up on style. "not how I'd write things" is not a blocker. Defiance of instructions is normal, you just have to steer it and correct. There's no substitute for diligence yet. 80/20 - this means your scope was too large, split the scope or tell the agent to revert, split the scope, and try again. RE and assembler: it's really good at this stuff. It can patch almost any binary with the right tools Swift: you have to give it tool usage in whatever result you're wanting. If it's a macos app, you have to let the LLM pilot it to get feedback, or build an extensive end to end test suite that it can drive autonomously. If you get into the loop on changes it'll feel awful and like no time savings. Review at the level of using the app and looking at the code, not in process or reviewing every tool call or diff. applescript: works great, I have a bunch of automation set up this way, what problems are you seeing? elisp: tough language, llms kinda hate parentheses unless you're really tight on the linting, and elisp is enough of its own animal that the training for e.g. common lisp isn't great. python: will suck unless you enable all the typechecking, make it use bdd, and have a linter/formatter run precommit and yell at the robot for you. Web vs cli: you should use the cli 100% because it lets you change the environment, if you're getting better results on web, you haven't set your local environment up very well. My personal preference is to run my own dev server on aws but that's spendy.
- al_borland 2mo agoI’ve found some success is small projects, with limited scope, in a greenfield. I’m terrified to attempt agentic anything in the repo my job actually cares about. I triggered it once by accident, when the agent was first rolled out and enabled by default… it broke everything. Now I just use ask mode, and even that is wrong half the time, and once it goes wrong it just keeps getting worse. I saw a post from Dave Plumber who vibe coded up a new cross platform task manager. He said his spec document for the AI was 107 pages long. So maybe what I’m doing wrong is not giving the AI a literal novel of spec.
- applfanboysbgon 2mo ago> He said his spec document for the AI was 107 pages long. This sounds like programming but with extra steps that make it take longer with less reliability.
- 0x696C6961 2mo agoIkr, at that point the code itself is a better way of encoding the information.
- hirvi74 2mo agoI concur with your first sentence. I have found success creating some sort of MVP, but I have had virtually no success with taking something from initiation to completion. My employer won't even provide LLMs for us, let alone allow us to use agentic coding on our repos. All our code is still USDA certified, organic, free-range code. > He said his spec document for the AI was 107 pages long. Absolutely not. My ADHD forbids such temptations of the dark arts. I'll feed any LLM a 107 page spec list, but I won't be writing nor reading that spec list.
- simonw 2mo agoTell it to use red/green TDD and start things off with an already configured test suite, maybe with a single test that asserts 1+1==2. Make sure it know how to run the tests before it starts writing any additional code. Then set it a clear goal.
- slopinthebag 2mo agoBasically all the examples of LLM's building impressive things have been because they have human written tests to base the implementation on. If you have an LLM write the tests the results are far less impressive or valuable.
- bharatsuthar 2mo agoYes and LLMs are known to cheat on tests written by them.
- slopinthebag 2mo agoIt's not always cheating either. They aren't intelligent, so they don't actually understand the purpose of the tests or can build them to define the actual semantics of the problem space. It's literally just next-token prediction based on the codebase and prompt. Cheating implies that they have agency, and ironically agents don't.
- andai 2mo agoLast year when they added computer use to Claude web I was excited to try it out. I just asked it for a code snippet and it ended up setting up a whole repo in a docker container or something. Even volunteered a test suite. This genuinely amazed me. ...until I checked the tests. It was just console.log("Tests passed!") AGI 2027
- simonw 2mo agoHave you seen that recently? I used to see that happen, but I've not caught it with the more recent (Opus 4.5+, Fable 5, GOT 5.5/5.6) models.
- mw888 2mo agoYou're appealing to ambiguity. All you've said is you have failed—how is anyone supposed to know what went wrong?
- hirvi74 2mo agoI suppose they aren't, but I am perfectly fine with reading what has worked for others should they feel inclined to share.
- mw888 2mo agoGenerally constraining scope and providing enough existing material until the LLM is productive. The common counter-argument is that specifying to the sufficient level is more work than just not using an LLM at all. I find that is not a universal rule.
- jaggederest 2mo agoI'd be happy to screenshare with you if you like, we can work on something trivial or open source. Half an hour should be more than enough to see whether you're doing anything obviously self-sabotaging.
- dosisking 2mo agoThere are two 'camps' with respect to AI. One camp already knows that Neural Nets don't work and are a dead end. The other camp hasn't yet figured out that Neural Nets don't work, but are convinced that they do (or eventually will), because they think everything always improves over time in a linear fashion.
- mortalapeman 2mo agoWith generated code, the directory structure, interface design and general state management is usually a haphazard mess. Even with the best frontier models. But what really gets me is the model often tries to make assumptions for me that I didn't specify in the prompt. Subtle things like which error states are "oh shit we need to bail" vs "this isn't a deal breaker." Sometimes it will ask, but more often than not it will just make a decision and it's often the wrong one. If I don't have a fully kitted out test suit and a good type checker to verify the final product against, the the whole looping thing is just useless to me and I'm back to reviewing every line of code it puts out and having to draw on my years of architecture experience to make sure we don't build a giant pile of trash.
- Gigachad 2mo agoBecause they are designed to be used by managers who don't know how to answer these questions and don't want to be asked them. Just have the magic answers box pick something.
- slopinthebag 2mo agoThey're RLHF'ed to an inch of their lives to be able to one-shot complete tasks, since requiring human input defeats the purpose of being able to replace the labor force. But once the insanity ends LLMs will be packaged as tools for developers to use to boost their productivity, and we'll consider them as we do IDE's and debuggers and stuff. But we have to get through this hype cycle first.
- siva7 2mo agowake up, slopinthebag. wake up..
- ed_elliott_asc 2mo agoYou think LLM’s will be more than a tool?
- 2mo ago
- theteapot 2mo ago> It helps to know that LLMs don’t “reason”. They predict .. Semantics. Prediction is the training objective. The ability to reason can be, and very arguably is, an emergent property of that.
- jayd16 2mo agoEven if that was true, you'd have to still prove it has emerged.
- krackers 2mo agoWhat would be your test to determine that?
- slopinthebag 2mo agoWhy would "reasoning" be an emergent property of prediction?
- hsn915 2mo agoHow do you predict without reasoning?
- slopinthebag 2mo agoWhere is the reasoning in linear regression?
- bluegatty 2mo ago"They’re foundationally incapable of always and consistently preventing prompt injection attacks. “Alignment work”, safety harnesses, and sandboxes all help to add barriers against the worst, but there are fundamental gap" ... They seem to be very good at a lot of rudimentary best practices, more so than humans, but more accurately - if you run and audit pass with specific instructions ... they're very good at that. I mean - it's what they're the best at which is applying 'fuzzy heuristics' in a mechanical way. If can describe issues concisely, the patterns, the styles, the rules then LLMs can very mechanistically and methodologically grind through them. I don't even see how this is controversial - without getting into 'what their reasoning means' - we can all agree that their synthetic reasoning is pretty good at narrow scales, and they've been 'trained by compilers' and are extremely good at spotting common patterns. If you back that up with a lot of tokens ... they excel. Designing architecture, that's difficult, but hammering away at all the 'known-knows across a system' especially to identify things ... they're pretty good at that.
- dmitrijbelikov 2mo agoLLM is the new Excel
- Krei-se 2mo agohey that's my line.
- rstuart4133 2mo agoyou must have stolen it off me.
- vismit2000 2mo agohttps://knowyourmeme.com/memes/spider-man-pointing-at-spider-man https://knowyourmeme.com/memes/spider-man-pointing-at-spider...
- user43928 2mo agoThe article says what many here like to hear, but in my opinion the core arguments are false. > Making software debuggable, maintainable, layered, and composable – that’s still quite a trick Not really. I have been working on a mobile app for months, and I stopped even glancing at the code about two months ago. 150k LOC, around half of that in tests, and the AI still has no problem maintaining the code on my behalf. Debuggable? It can add extensive instrumentation in seconds. None of this requires expertise, prompting, or mention of TDD. It's the default. Frankly I do not believe the author tried developing a large codebase fully agentic and without reviewing the code. I believe many here look at the code produced, deem it substandard, and go hands on. > They’re foundationally incapable of always and consistently preventing prompt injection attacks From Anthropic's article about the Auto mode: > We commissioned an evaluation from a third party, Trajectory Labs, who tested different models within the latest publicly available versions of Claude Code and Codex as of July 17th 2026.1 They tested 72 indirect prompt injection scenarios held out from Anthropic > In this evaluation, none of the 720 attack attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode. On the other hand, 5.83% of the attacks succeeded against GPT-5.6 Sol running Codex's Auto-review mode. Notably, this is greater than the 0.09% average attack success rate against our latest models running in bypassPermissions mode without additional safeguards. The tests showed a 19.03% attack success rate against GPT-5.6 Sol when running in Full Access mode I'm sure someone is going to reply with how they do not trust Antrophic's research, but lacking other data, prompt injection appears to be largely solved already.
- hbcdbff 2mo agoHow do you expect us to take your views on LLM code quality and durability seriously when a) you don’t even look at the code and b) you’ve only been doing this for two months?
- user43928 2mo agoI've been working on the app for four months, and I am clearly not talking about code quality. I am talking about product quality and maintainability. Both are more than adequate. I know this because I have worked on it for an estimated 300 hours. Has the author practiced a similar approach for even a week? I doubt it.
- Alien1Being 2mo agoAI generated code is like IKEA furniture. IKEA furniture embodies many elements of good cabinet making but skips many nonessential elements. And does this more consistently than cabinet makers who can be bored, incompetent, depressed, burnt out, resentful, tired, having a bad day. In the future AI code inevitably will embody most good software engineering practices. And will do this more consistently than software engineers who can be bored, incompetent, depressed, burnt out, resentful, tired, having a bad day. Just look at the messages on HN or around you at your colleagues to see how mediocre the average software engineer is.. Today's IKEA is good enough for most people. Tomorrow's AI coding will be good enough for most corporations. Good enough to vastly reduce the need for fine craftsmen and women / software engineers. Good enough to deskill those who call themselves cabinet makers / senior software engineers. These days the cabinet makers I personally know just do contract kitchens for project builders. But IKEA is and AI will be, bad enough that at the high end with special requirements / taste / money / an inflated sense of self worth, some furniture makers still exist and thrive. Perhaps 1% percent of current software engineers of today will be needed in the future when AI code inevitably has the ability to follow good software engineering practice....... And as usual it will mainly be the mediocrities that remain ( so there is hope for you too ), with occasional islands of excellence.
- hugodan 2mo ago...ah the classic middle manager analogy of code is X, where X is nothing like code at all but is being used to drive a point that is just standing on poor grounds. keep it, we have been through "is like building a house", "like following a recipe", "like a living organism", "like $SOMETHING_WITH_COMPONENTS", etc... we can handle your IKEA furniture, thanks you for your contribution
- amelius 2mo agoHaving not been formally verified, almost all software today feels cheap. Maybe an AI can change that at some point.
- ChrisGreenHeur 2mo agoYay let’s lock ourselves into the formal verification toolsets, so that we can never use new language features again.
- deleted 2mo ago[deleted]
- amelius 2mo agoBut do the same SWE fundamentals apply if the one doing the programming is many times smarter than us?
- baliex 2mo agoAn interesting question. We (humans) care about maintainability because the codebase will be adapted by teams of us for many years, based on new feature and bug fixes. Rewriting from scratch is basically never an option after a certain amount of time. Maybe the machines could just start from scratch each time and come at maintainability from a totally different angle
- mrkeen 2mo agoIt wouldn't make sense to confuse attributes of the software with attributes of the programmer. Otherwise, one could simply declare one's IQ in a const somewhere, and have all unit tests follow the form: if the programmer's IQ is high enough, then the method under test is likely correct.
- justincormack 2mo agoLLMs are not smart in any way, but they are a weird kind of effective thats very different.
- hirvi74 2mo agoI have worked with a few people like this.
- mikgp 2mo agoI think SWE fundamentals matter - because it will be a long time before software is a closed system. And the problem is that as long as humans are in the loop building software that dynamic will have to be maintained. We use Loki for logging at work. There’s certain types of queries it just doesn’t support. And so the question becomes - Do you change logging providers Adapt to Loki’s capabilities Create a third layer / tiered storage. And each of those decisions have multiple downstream consequences. It’s not that LLMs can’t make those decisions per say, it’s that What does an LLM do when five different people ask for a system optimized to do five different things. It could figure it out itself, but like I don’t think that’s how the human software contract works.
- brabel 2mo ago> Making software debuggable, maintainable, layered, and composable – that’s still quite a trick. Quite a lot of that work requires extensive, thoughtful reasoning. And that’s where the LLM’s today, even the leading edge of the “capability” from frontier models, fall short. It’s been my quest during my career to figure out what is maintainable software, what is composable or not, and how the two things, and many other things, are in direct conflict. There is no single answer. If your goal is to take over the market quick, as many here would like, maintainability is a very low priority aspect of your code base. Composability may matter for integrators but can be entirely ignored in your CRUD backend. Beyond that, I don’t know of a good way to measure most of these intangible properties. Highly competent software developers disagree in even basic things, like whether OOP is a good idea, should we all be using pure functional programming etc. Hence, how would you expect an LLM to get good at figuring this out for you? If you describe exactly what trade offs you are willing to make, and give it ways to measure how well it’s doing, then I do think LLMs will be able to not fall short. Given the current state of things, it’s just a matter of opinion whether LLMs fall short, or humans fall short for that matter.
- louiskottmann 2mo agoThoughtfully and coherently factorized, abiding a set of architectural rules (i.e: we compose "this" way here, re-evaluated as we go) and following up-to-date framework conventions. LLMs are terrible at this. Chosing OOP or FP is irrelevant, fundamentals matter more
- dolni 2mo ago> If your goal is to take over the market quick, as many here would like, maintainability is a very low priority aspect of your code base. It's low priority in the sense that people in charge tend not to value it. That's different from saying it is not impactful, or wouldn't lead to a good business outcome, if the software were built better. Software that needs to be babysat is an ongoing opportunity cost.
- mekoka 2mo ago> If your goal is to take over the market quick, as many here would like, maintainability is a very low priority aspect of your code base. Is it? How long do you expect to keep your momentum after "taking over the market"? Anecdata time. I once joined a 3 year old project that had ground itself to a near halt with this philosophy. The project's lead seemed almost allergic to the word "refactoring". It had accrued so much tech debt that I was the third "new guy" to join in less than two years, after the previous attempts to hire had successively faltered within 6 months, because my predecessors couldn't deal with the unmaintainable mess. I made it to 9 months. Maintainability is not tied to OOP or functional, but rather to how much a team cares to manage the cognitive load that comes attached to having to deal with the code base. When that becomes a genuine priority, the code tends to be written with concern for the next human mind's ability to interact with it. And when it makes sense in that one pursuit, functional, OOP, DRY, WET all become valid -- even seductive but toxic affordances like inheritance can sometimes be useful in the right context.
- mstank 2mo agoI have an open question for software engineers out there: As someone that has never studied CS but has written basic code most of my life (accelerated now with AI), where is the best place to learn software engineering fundamentals?
- arcanemachiner 2mo agoI don't think there is such a thing as "software engineering fundamentals", as the fundamentals differ based on the type of software you want to make. Do you want to make websites, or work on embedded systems? Do you need to squeeze every OK ounce of performance out of the machine running your code, or is developer velocity more important to you? Will your code run on a single machine, or does it need to be networked/distributed? The thing you want to make determines what the fundamentals will look like for that area of study. There isn't enough time to learn the fundamentals of everything needed to make good software across all domains. And even if there was, you would be wasting time learning all the principles that don't apply to 99% of things you would be working on at a given moment.
- adam_arthur 2mo agoThere are tons of core principles that can be learned that largely apply across fields. One prime example, single source of truth for data/concepts. To be violated only when performance is meaningfully improved (denormalized databases). But when you do so, you should definitely recognize you're opening up out of sync issues for that performance gain. Though for the majority of code, there is no performance benefit to adding multiple sources of truth. Yet it's the most common error I see re: quality. The sad thing is that software engineering fundamentals and best practices never became widespread or widely taught in school prior to LLMs
- lilbigdoot 2mo agoMaking lots of things and learning what works. It's good to read and find ideas to grow, but a volume of work is the most important thing. You'll discover a lot of the ideas on your own out of need. Keeping an eye out for tools and ideas related to what you enjoy can help broaden your horizons, but time spent making things is best
- AgenticVnus 2mo agoIts good
- tomwuu 2mo agoTime will tell.
- jamesforestwest 2mo ago[dead]
- wabstractions 2mo ago[flagged]
- jartan2002 2mo ago[flagged]
- CodeWithLeo 2mo ago[flagged]
- erichocean 2mo agoScenario: an AI lab develops ASI for software development internally. The ASI produces bug-free software and human-readable specs. It's so reliable the company can guarantee the code matches the spec. Rather than provide tokens to developers, they instead sell finished software to the companies: specs in, software out. Companies no longer need to employ software developers. Instead, they buy bespoke, bug free, guaranteed quality software from the AI lab. Seems plausible to me.
- johnnyanmac 2mo ago>Seems plausible to me. As plausible as any cyberpunk Utopia model you read up on, sure. Nevermind that AGI/ASI is still pie in the sky thinking as of now.
- monitorion 2mo ago[flagged]
- styrum 2mo ago[flagged]