9 ms·
Why Software Factories Fail (or: harness engineering is not enough)
- _doctor_love 3mo ago> I haven't been able to dig up any definitive data/findings from StrongDM on how that whole dark factory went. The weather-report has a few sparse updates between February and June of this year. This was easy to find out I thought. And just with an old-fashioned google search too, no deep research agent needed. See here: https://diffusion.io/ https://diffusion.io/ Seems like it went pretty well if a consulting company is now being started. I agree with a lot of what Dex Horthy is saying here but on some fronts I feel like he's missing something. Coding well with LLMs, it's not a skill issue, it's an effort/laziness/rigor issue. In order for coding with LLMs to go well, there has to be more rigor, more discipline, more good engineering hard-assedness. To reiterate, the teams seeing the best results with AI were already high-discipline and high-hygiene. AI works on data. The better the data, the better the likelihood of a desirable outcome. Code is data. If you have bad code, no matter how awesome the model you let loose on it, you can't get as good a result as if you had good code to start with. This principle has been well known in AI/ML circles since the 20th century. e.g., if you are doing spec driven development and not seriously investigating formal verification, IMHO you will come up short. Prompts are simply not enough to steer a coding agent to the level of precision needed. Without deep programmatic verification - at all levels, formal verification is just one slice - the solutions the agent produces will always be just slightly (or very) out of true.
- edot 3mo ago“Seems like it went pretty well if a consulting company is now being started.” You interpreted this backwards. Software companies offer consulting when their product cannot stand on its own. See Palantir, Salesforce, etc. They are successful companies, yes, but not successful products. The product needs to be instantiated and maintained by sales engineers and consultants and customized into something so bespoke that it’s hardly the company’s product anymore.
- stellar_jay 3mo ago> Prompts are simply not enough to steer a coding agent to the level of precision needed. Without deep programmatic verification - at all levels, formal verification is just one slice - the solutions the agent produces will always be just slightly (or very) out of true. I found this to be exactly right, and in my work I’ve come up with a taxonomy of constraint mechanisms which I keep in mind when guiding agents: generative to constrain the output of the model, interpretive to constrain how the model ‘understands’ code, and elicitative to help it ask the right questions of users. Full write up is here: https://www.research.autodesk.com/blog/constrain-agent-not-user/ https://www.research.autodesk.com/blog/constrain-agent-not-u...
- _doctor_love 3mo agoThat's an excellent writeup. Haven't gotten all the way through it yet but so far I'm with you.
- stellar_jay 3mo ago[dead]
- sythe2o0 3mo agoSome more context on the consulting company: StrongDM was sold earlier this year, about a year after the dark factory was first announced, and the former CTO moved on to this (presumably) in order to continue the idea. Disclaimer: I'm a former StrongDM employee
- navanchauhan 3mo agoapg?
- dhorthy 3mo ago> In order for coding with LLMs to go well, there has to be more rigor, more discipline, more good engineering hard-assedness. To reiterate, the teams seeing the best results with AI were already high-discipline and high-hygiene. hard agree. But i don't think this is sufficient. Even formal verification has its limitations. > AI works on data. The better the data, the better the likelihood of a desirable outcome. Code is data. If you have bad code, no matter how awesome the model you let loose on it, you can't get as good a result as if you had good code to start with. This principle has been well known in AI/ML circles since the 20th century. hard agree. but also RL data is shaped differently than SFT data that has driven the majority of AI/ML innovations since ~2000, and its where there's so much room for innovation still. e.g. ImageNet was all just hand-labeled answer pairs. > it's not a skill issue, it's an effort/laziness/rigor issue I'm sorry but this feels like a semantic argument - the point of "skill issue" is "you didn't put in the effort or learn the techniques"
- _doctor_love 3mo ago[dead]
- _doctor_love 2mo agoNot sure what happened but I wrote you back a whole reply that I can see posted when I am logged in but not when I am logged out. In any case, tldr of that comment was that I don't think we have any fundamental disagreement.
- deleted 3mo ago[deleted]
- jaytaylor 3mo agoHi, I'm one of the trio from the StrongDM AI Lab. Just a minor thing I want to clarify about the Weather Report [1] - it's framed in kind of a negative light in the article ("sparse updates"), but we've been updating it as frequently as we find a meaningful improvement in a relevant dimension. Since we launched it in February it has averaged about one update per month, as frontier labs keep racing forward! [1] https://factory.strongdm.ai/weather-report https://factory.strongdm.ai/weather-report
- dhorthy 3mo agoappreciate that context! I definitely did not mean to come out and say "its definitely not working" or anything, but would love to hear from y'all a retrospective on the ~5-6 month anniversary - what was right, what did we get wrong, etc
- jaytaylor 3mo agoWe are working on new articles to share our latest findings, so stay tuned! Overall our outlook continues to be bullish. Almost all software problems yield to a combination of the Factory Techniques covered on the strongdm.ai website. More powerful models work even better...
- dhorthy 3mo agoawesome - i have updated the post with a link to this thread!
- ghostinit 2mo ago[flagged]
- vanuatu 3mo agoThis is one of the best writeups I've seen of this a lot of the model's constraints come down to how they are RLed. Discussions online would be a lot better if everyone understood how the labs train the models in a high level (or did a lil data labeling)
- dhorthy 3mo agoyeah i like this and others in the thread mentioned that understanding RL and RLHF and the shape of the data is really important (at least the fundamentals, I'm sure there's quite complex industrialization of RL inside labs as Nathan Lambert says)
- vkaku 3mo agoNecessarily, better data is what we need, more importantly, better collaboration and better specialization at all. While the title is a bit misleading and clickbait-y, the message is decent. I disagree with the way that big models are trained on noisy relationships and RL is applied to tone it back down, it represents a stupid amount of compute thrown at this problem at a scale that is often unnecessary. The rest of it is on point.
- syndacks 3mo agoDex you aren't part of the slop cannon, you _are_ the slop cannon
- dhorthy 3mo agoi can't tell if this is a compliment or not
- molsongolden 3mo agoJust in case this isn't a compliment, I'll note that I think Dex has been pretty good about admitting when they were wrong, explaining why they were wrong, and what they're doing now instead.
- dhorthy 3mo agowe out here trying
- rglynn 3mo agoTo me, the thing that stands out about the whole state we're in here is PR review. Yes, in an ideal world, PRs read well, are a joy to review, reflect what you discussed etc etc. We have to be real; there is only so much we can do to that end. I'm not sure how the best teams do PR review, from my perspective it sucks. I'm talking specifically about the UX. I've always hated Github's PR page, so I typically reviewed by pulling down the branch and opening the diff with $EDITOR. These days I think there's really no excuse for the awful UX. Linear (a company that isn't even in the domain of code review) put out a basic PR review feature[0] that is already better than what GH offers. It's simple: point a small model at the PR, group file changes together based on theme, add some commentary and sort by importance (schema changes > openapi spec). Immediately, so much mental load has been reduced without the reviewer or the requester doing anything. This feature is pretty damn basic, and I think there are obvious next steps like generating visualisations which a dedicated product could find the time to implement. Keen to hear others thoughts on why this is the wrong approach, or if there are tools in wide use that solve for this, or why this isnt the right problem to focus on. 0 - https://linear.app/docs/diffs#guides https://linear.app/docs/diffs#guides
- xorcist 3mo ago> group file changes together based on theme, add some commentary Isn't that what commits are? Or ... should be?
- lozenge 3mo agoNo, commits always display the files in a fixed order, and then display changes from line 1 to line N. An AI could select the order to display in, add per hunk commentary and automatically adjust how much context lines are displayed.
- xorcist 3mo agoBut if you need to split the commit in hunks, and add a commentary per hunk, isn't that just a sign that you really should just split your commits? That's the most killer feature of git, that it's so easy to slice your commits any way you desire, and then redo again. The use case of taking a chunk and commit separately is so common it even got a special mode in the add command. That, and the super fast jumping between branches is what set it apart from contemporary version control systems. The extra context provided by the review tool is gone when the review is done anyway. Review systems come and go, but the commit log is for eternity.
- AIorNot 3mo agoWait these arent “software factories” they are strung together ai rube goldburg machines Its crazy to me people write these articles and create standards like this is some kind of engineering standard with years of research and experience This is like calling these folks the experts on aviation: https://youtu.be/M9Yww9LG3gw?is=xgtA-xMpNy-09Asu https://youtu.be/M9Yww9LG3gw?is=xgtA-xMpNy-09Asu Its still so early in the game for de facto standards - engineering teams need to experiment and see what works for their own quality metrics not just parrot “standards and methodologies” This is still the very early days of AI and AI engineering
- dhorthy 3mo agointeresting - i'd say my main goal is to put the current "agentic software factory" hype in the historical context of "we've actually been rube-goldberging software deploys for a while now"
- M4R5H4LL 3mo ago[flagged]
- dang 3mo ago"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html (and please particularly avoid personal attacks on this site)
- fishtoaster 3mo agoThere's some good ideas and points in here, but this bit threw me: > # We tried this > In July 2025 we went full lights-off Isn't it pretty well-accepted at this point that the models underwent a step-change in usefulness around fall 2025 / spring 2026? I know that I was able to start handing agents whole features after that, but not before. I feel like any perspective/experience on "what agents can/can't do" from before that period is... maybe less than relevant to the modern era. TFA calls it out a few sections later with "But surely the models have gotten better since then", but then just writes off any improvement. That does not match my experience.
- 2001zhaozhao 3mo agoI had a bit of this impression when reading the post as well as the authors' product website. A lot of it does seem to be stuck in 2025. For instance I think their post "long-context isn't the answer" on their website post straight-up isn't accurate, and gives me the impression they are just extrapolating previous performance to new models. In my experience, Opus 4.6 and newer have worked very reliably for long context to me (i don't perceive any intelligence drop at 700-900K tokens). Yeah it's extremely cost-inefficient, but it works.
- deleted 3mo ago[deleted]
- hansvm 3mo agoFWIW, I notice the intelligence drop drastically with long contexts with Opus 4.6. It's barely usable for anything intricate and long. That long window is good for _something_, but it's not as good as a short window.
- Terretta 3mo agofor synthesis. read the entire whatever in one gulp and boil it down or plan what to do about it. not for steps.
- 2mo ago
- 2001zhaozhao 3mo ago> When I say maintainability, I mean the specific thing where it becomes really, really hard to change one part of the codebase without breaking another part. The corollary of agents being bad at maintainability but good at coding is that you can vibecode all the parts where maintainability doesn't matter. So if you build a (domain-specific) modular architecture for your software first you can then just let your software factories loose on building the modules.
- dhorthy 3mo agoyeah this is along the lines of what some friends of mine call "core vs. pragmatic modules" or even s/modules/codebase zones/ the idea that if you have a solid core and decoupled modules, you can have "zones" in your codebase where you allow the model to run wild and do a little slop, because you know the blast radius is contained
- throwatdem12311 3mo agoUntil you inevitably need a cross cutting concern. “Oh just this one time” Then an agent sees the pattern and assumes it’s a best practice. Then your beautiful architecture is ruined. “Just this once” indeed.
- dhorthy 3mo agothis is 100% right. you have to guard the codebase patterns with your life. because the codebase is part of the prompt.
- layer8 3mo agoExcept that there is typically a feedback loop between the implementation and the module interfaces. While implementing, you discover aspect that makes you adjust, and sometimes completely alter, the interfaces and module boundaries. And that’s not just for an initial implementation, it continues to happen as new requirements come in over the lifetime of the software. Ostensible implementation details continue to inform the architecture.
- Makeph 3mo ago[dead]
- mrbnprck 3mo agoI've recently started experimenting with grounding LLM driven implementation/verification on RFC based normative specifications, to avoid having to manually steer the LLM during implementation and dealing with reviewing sloppy pull requests. It works quite well, as it puts your entire focus on writing (hopefully) unambiguous specifications vs. having to discuss unwanted changes with an LLM during code-review. One flaw is that this only works great if you know exactly what you want, which is not always the case.
- dhorthy 3mo agonormative specifications can help, but the thesis here is that specs that define behavior of the product or even architecture are helpful but there's MORE that can be done and even though "program design" feels too in the weeds it's still essential if you care about maintainability
- mrbnprck 3mo agoIsn't maintainability mostly about applying basic engineering principles e.g. separation of concerns, single responsibility, dependency inversion, open-closed etc..? If these topics are addressed in the normative representation, and correctly translated by an LLM into code (especially by slicing up the specifications into measurable outcomes to set the intended foundation), then future change e.g. maintainability becomes essentially easy as well, no? Historically I know that the majority maintenance problems occur from slow continuous evolution of a system that it initially was never designed for. And the only way to address this was continuous system design.
- dhorthy 3mo ago> Historically I know that the majority maintenance problems occur from slow continuous evolution of a system that it initially was never designed for. And the only way to address this was continuous system design. yes exactly - this is what I'm advocating for - that you can't skip the system design, and that actually good system design goes down to the typedefs and object graph at the code level, not just mermaid charts and db schemas and service contracts. I will highlight what a few others have said along the lines of "a sufficiently detailed spec IS code" - that is, to make the spec guaranteed to produce the code you want, the spec will look a lot like code (and will be roughly the same effort to review as the code itself anyway, saving you no time) what I'm proposing is "how can you maximized the odds that the code WILL be good or close enough to good that its easy to get there, with the LEAST amount of human effort/attention" - how can you move fast without skipping what matters https://haskellforall.com/2026/03/a-sufficiently-detailed-spec-is-code https://haskellforall.com/2026/03/a-sufficiently-detailed-sp...
- jfkfisksnsb 3mo ago[dead]
- jadar 3mo ago> So, why can't models do software maintainability? I feel like the explanation does nothing to actually elucidate why models can't do it. Is it an inherent weakness of LLMs? Training processes? The typical "this is crap" that we constantly hear? It goes on to write about RL and how there's no penalty for bad design. But that sort of side-steps the question and makes you ask: "why not do RL and make a penalty for bad design?" Of course the models aren't good at it ... they're not good at anything until you've tuned them and put them in a harness that rewards good edits and throws away (improves) bad edits. That doesn't explain why "models can't do software maintainability." The real question is why harnesses can't do software maintainability, and how to build a system that can do it. (I suppose that's the purpose of the ad at the bottom of the page.)
- dhorthy 3mo agofair point, this is the thing I struggled most to extract out while writing it - if you can propose an RL environment that penalizes a model for bad design, then I'm all ears - right now there's no fast oracle/verifier for this (as stated in the post) My current evolving take on "how would you build such a thing" is you need to tee up a roadmap of 20 features and feed them to a model one at a time, so it can't design up front for what's coming. That way if it builds the first 10 features and the codebase goes to slop, it get's penalized when it can't build features 11-20, or when those features take wayyy more tokens/time/cycles than a model that maintains a clean codebase can do. This is how most real software is built by most teams - incrementally, getting feedback from users along the way, and steering goals in response.
- jadar 3mo agoIsn't that how AI written software gets better too? By steering the model towards the goals of the user? I don't know if it's a question of how many features to feed to the model, either. Of course, overwhelm the context with too many features and it will get confused. But that's where the memory management idea that G. Huntley talks about is helpful. You're trying to steer the model within its memory limits towards a certain goal. The problem is getting it to produce "good" code. Formalizing what that means is the task of the programmer. How do you steer a model towards always, or more often, producing good code so that you don't have to do rework? That's the same problem as with a junior engineer, but the way you do it is different. Right now we're trying to do it with mountains of prompts — which sort of works but has diminishing returns — and with onerous code reviews. We've seen this get better over time, but I think some more mechanical methods will help as we figure out how best to steer the models.
- firasd 3mo agoI think there is a fundamental issue here of what building software even means If you think you can just assign Github tickets to AI agents and go drink daiquiris on the beach I think you'll find that you end up with more and more towers of abstraction and indirection. There are 'points of view' that emerge during coding I think. And at some point you as a human have to be like "wait... what if we use Redis here". "Wait.. the API is already returning the data we need". "Wait... let's not add customers to the report who have not been active in the past year". Stuff like that
- dhorthy 3mo agoyeah I 100% agree - and I think the most popular coding agent workflows / skill kits are designed to pull those insights and intuition out of humans in a way that optimizes for the developer's experience building the plans or building the code, e.g. - claude code plan mode - mattpocock/skills - obra/superpowers - research/plan/implement etc etc
- AmericanOP 3mo agoThe machine gives you what you ask for even when that thing doesn’t exist yet. Rather than “lights off,” utilizing information theory, decision-making theory and creativity theory makes me better at asking for the right things. Memory is not transcribed to weights like when humans sleep. Memory is notes handed to someone on groundhog’s day who can’t remember yesterday. We hope they believe us. Don’t be too surprised when a highly entropic system introduces entropy to a project over time.
- 3mo ago
- d_silin 3mo agoMy radical opinion is that LLMs are harmful for software development - they are the ultimate "goto" operator. All actual code should be written by a human developer. Instead, use them in adversarial mode - run QA scenarios using LLM agent as a substitute for end user to do bug discovery.
- arm32 3mo agoI love how we need to preface such an opinion as being "radical" nowadays.
- 0xblacklight 3mo agowhy?
- antonvs 3mo ago> All actual code should be written by a human developer. This seems arbitrary. Why don’t you say the same thing about machine code? Developers use tools so they can avoid writing machine code. What is causing you to draw a line in the sand about use of tools? The obvious answer is it’s just a function of the time period you grew up in, and a lack of willingness to adapt to change.
- Archer6621 3mo agoThis is indeed an important angle to consider as well. I think there is some nuance though, but it's hard to articulate. I think that is because there are some long term effects that are currently invisible, but that can be anticipated. A lot of it is the human factor, and how changes to how people think and act may ripple through the organizations as well. Things such as skill atrophy which may reduce not only immediate skills, but also adjacent skills that were maybe important for critical thinking and making good decisions. Or the fact that learning to use an LLM is a skill that is not truly grounded in reality and can therefore not translate well to other areas (you're essentially learning skills within the "reality" of the LLM, based on its weights). That is very different from machine code --> tools, where the tools make explicit assumptions that are based on the workings of the machine code (which is analytical in nature), and where you can, if you wish to, jump in to override those assumptions where needed without having to adapt to a mental model that operates in its own reality (i.e. the LLM's "thought process").
- ozhero 3mo agoThis is a very well written article and he makes his arguments backed up by data. We may choose to disagree but thats the point of healthy debate based on clearly expressed opinions. Key point is I don't think this is AI slop which is way too common in long form articles these days and in keeping with the whole point of his article.
- rapatel0 3mo agoThe dude is selling an IDE. Also it's missing the point of a software factory concept A software factory will not work infinitely forever for everything. A software factory isn't a solve anything button (aka god). In a conventional factory, things break and fail. Process machines get poluted. Extruders get jammed. You still need to establish intent, define what you care about, define guardrails, and of course manage the factory.
- dhorthy 3mo agoand build it incrementally! You don't have to build the entire software factory at once. You don't have to mastermind the whole future system, instead you're actually stacking and layering these small, isolated problems. I think that's a really good approach to start getting value tomorrow or this week without saying, "I'm going to revolutionize how we ship." It's just: 1. Find places where you can use agents. 2. Figure out where the right leverage points are for humans and where the right leverage points are for agents. 3. Just start building those things and plugging them into each other. One day you'll wake up, and 80% of all of your stuff is automated.
- rapatel0 3mo agoTotally. Also to add to that. I think (like any factory) you need to build in systems to allow you to have observe and service the machines make process changes as needed. You're the foreman/factory manager. You can optimize the factory over time upgrade machines, place machines closer or add conveyor belts for more productivity. The only think that slightly worries me is that the Labs are clearly incorporating the best in class logic from their datasets.
- zingar 3mo agoEnjoyed most of this but unconvinced by the program design part. If I see an agent writing function signatures or listing which functions to edit in a plan that tells me that I’ve given it too big a vertical slice. I always delete the code guesses. The thing that writes the code must always spend some time discovering where/what to write or it won’t have the right context. Or put another way: “decide first, act later” always feels worse than act-learn-act.
- dhorthy 3mo agoone thing I probably didn't mention is we do the program design having already done an in-depth codebase research, with current patterns and architecture surfaced - that actually seeds every step of the flow including even the product part - but yes if you're working in very small slices then I think it's very feasible to skip program design and just review the code as you go, and resteer live. I do this all the time for tasks that are too big for a oneshot but on the smaller side overall.
- zingar 3mo agoDo you think that the program design adds a lot on top of pointing out the current patterns? I might be anchored by working with humans, but if I saw a plan for humans that included function signatures and what calls what I would say that is way too detailed. As a result I don't put it in my plans for AI either.
- sergiotapia 3mo agoPersonal anecdote: I was able to set things up in such a way that multiple non technical people at my job were able to build features into our project. One person even created our own CRM. And I'm not a turbo-genius savant. The one downside was that it required their computers to install dev dependencies, postgres, infisical for secret management, etc. That's the next frontier. What I'm working on now independently. I plan to open source this solution, but again I'm not a savant genius. I guarantee many people are working on the same thing it's converging. Our job as engineers is becoming more of a higher level facilitator and AI "plumbing" maintenance work. How can we get the AI to empower everybody at the company while keeping the wheels turning. That's my aim.
- Robdel12 3mo agoI 100% whole heartily agree with this. Anyone saying the models made a huge leap in fall 2025 / spring 2026 still aren't looking at whats going on. Just this weekend I thought I had a solid plan and evals/tests to let 5.6 sol work unsupervised over night while I slept and it made an absolute mess. Somewhere in the loop it had to make a decision and it made a wrong one. Making everything beyond that trash. The code 'worked' but it had the wrong system design and wrote the most brittle tests around its assumption, validating its own decision. It turned into a nasty feedback loop for the model. This is 5.6 sol high. The models write good/great code. I'm very happy to never write code again but models are no where near good enough to run off on their own without a human in the loop. Models LOVE to cheat. Just peek at the tests they write. And before you come at me, I have built 7 products in the past 1.5/2 years with agentic engineering, all with users. One of those projects is _dead_ because I let the vibe go too hard at the same time Anthropic decided to nerf both their harness and their models silently. If you care to look at the source: https://github.com/Robdel12/OrbitDock https://github.com/Robdel12/OrbitDock I spent a week or so and like a billion+ tokens trying to refactor and save it. It just wasn't worth it. I wish people would be pragmatic about this. I get the dream is to let it do everything and not to be in the loop, because being in the loop is exhausting. But if you want to make whatever you're building be robust and survive more than 6 months, you have to. I don't care how good your tests, plans, skills, etc are. At some point the model will have to decide something and it'll be the wrong one. Compounding the slop from there forward.
- dhorthy 3mo ago> I spent a week or so and like a billion+ tokens trying to refactor and save it. It just wasn't worth it. this is exactly what happened to github.com/humanlayer/humanlayer - it was overslopped and we reset from scratch to build it right - spent 2 weeks in VS CODE - not even cursor, plumbing the core patterns from scratch. codebase is part of the prompt, yada yada > I wish people would be pragmatic about this that might be the tl;dr for the whole post haha
- maerF0x0 3mo agoI'm currently learning about Cost functions and Regression in the Machine learning course on coursera[1], and I cannot help but be struck by the similarities of how gradient descent seeks to minimize the error between the model and the training data, how agents do the same (far less mathematically) to seek an acceptable solution to software problems, and how light evolutionary pressures seem to guide species towards a better fit solution to the problem of existence and reproduction. (I'm no expert on the evolutionary example so go easy on me!) [1] - this course: https://www.coursera.org/learn/machine-learning/home/welcome https://www.coursera.org/learn/machine-learning/home/welcome
- forlorn_mammoth 3mo ago> how agents do the same (far less mathematically) to seek an acceptable solution to software problems, except that there is no gradient towards 'better' software, at least not in a mathematical sense.
- Altern4tiveAcc 3mo ago>It's easy to be a little bummed by the core thesis here: "for now we're stuck reading the code". >I was pretty excited for a world where we could just ask for things and let the models cook and not read the code and get beautiful production software I feel so disconnected reading those things. Reading and writing code is what brings me joy. I'd never feel "bummed" or "stuck" with it.
- stillpointlab 3mo agoIt is comforting to find other people experiencing the exact same reality as me, since I see so much in this post that matches my own experience. It reminds me of all the recent talk about "taste". Architecture "quality" may not be objective in a right/wrong sense, in the same way that fashion isn't right/wrong. It's like we are all going to have to relinquish reason/rationality to the machine and start to study up on aesthetics. Even historically, my big struggles have usually been deciding between two nearly-equivalent options. I get this a lot now with LLMs because there is no break in-between these decisions that implementation used to force. I feel I'm constantly making "taste" calls between tradeoffs that have no clear objective criteria, and it is as exhausting as the code review this post (and my experience) suggests are still necessary, even with Fable/GPT-5.6 level models. In many cases, I do what I've done with junior engineers whose code I reviewed pre-agent: make on-the-spot judgement calls. When I see a broken window, I call it out. But when I see minor issues, I sometimes just let it pass, note it in memory and tackle it wholesale once an accumulation of similar minor issues get to a certain size. As a tangential aside, I consider two dev shops from pre-agent days. One decides to hire 7 extremely talented engineers and gets them to work closely together. The other decides to outsource to 100 decent engineers and tries to silo them into modules. I think we are facing a similar choice with agents. You can either work extremely closely with a handful of agents, collaborating on design, review, etc. Or you can spin up a fleet of sub-agents and YOLO, then try to separate the wheat from the chaff in some automated way. My taste is the former, small highly coordinated shop. But time will tell if I am right or wrong.
- dhorthy 3mo agoyeah my best articulation of taste is something i got from Jake Nations[1] while he was still at netflix - "you know a bad pattern when you see it because at some point you were up at 2am debugging it" taste is the hard-earned intuition about every anti-pattern and landmine that has blown up in your face since you started doing software 1 - https://www.youtube.com/watch?v=eIoohUmYpGI https://www.youtube.com/watch?v=eIoohUmYpGI
- stillpointlab 3mo agoThat is a good point and I don't mind getting a bit philosophical when I point out that experience is distinct from rationalism. Underneath this there is an argument about empiricism vs. rationalism (or realism vs. idealism). The hand-wringing on the 50-50 cases is almost always theoretical, in the sense that I am trying to reason instead of rely on memory/experience. But that is a double edge sword because sometimes memory/experience are closer to trauma and can lead one away from a solution that is correct now in some new environment. It also reminds me of the phrase "use in anger". You don't really know about an approach/architecture until you've had to deal with something in an urgent or high-stakes moment (your 2am debugging). But to get that experience, you have to have shipped the thing first. It is a chicken-and-egg problem, you can't debug something at 2am unless it is live, and you can't know if it will cause you a problem until you've dealt with it. AI (LLMs, agents, etc) is this giant alteration to the environment that shakes up everything. In some sense, I feel I have to throw out my taste and "use in anger" all over again. It is painful but may be the only way. And I have to accept the risk that going slowly (like this post suggests) may be the wrong way, and I may watch the young untraumatized new-comers blast by me riding their 100 agent orchestrations to massive success. What I mean here is, in the final scenario from my original comment (7 top guys vs 100 decent guys) - I know the result from experience. But am I just traumatized and it will be different with agents?
- basketbla 3mo ago> Claude Code won because of Reinforcement Learning inside the harness Doesn’t this gloss over the fact that token subsidization is a thing? I assume the reason a lot of people switched from other harnesses to CC was all the weirdness around setup-token and whether third party harnesses would continue to work without api pricing.
- dhorthy 3mo agothat's probably part of it but every enterprise in the world (even teams as small as 10-20 engineers) are paying per token, not with subscriptions. Claude code did make it into those orgs because people played with it at home with subscriptions first, but the bulk of that revenue is almost certainly coming from people paying per token, not the ones getting the ~90% subsidy on subscriptions
- cadamsdotcom 3mo agoLove the idea of RL for codebase health. And a benchmark to measure against! Imagine a "MaintainabilityBench" that rewards models which detect code duplication while working on a task and perform some refactor instead of glibly duplicating; or that detect the need for a new architectural layer, or that hoist a type constraint so there's no need for dumb casts. You can keep on imagining scenarios. There are probably a few hundred distinct elements to RL for. The books "Working With Legacy Code" and "Architecture of Open Source Applications" would be great fodder. Sadly don't have time to build it, there's this mountain of reviews in front of me...
- dominotw 3mo ago> perform some refactor instead of glibly duplicating; wouldnt that be part of original RL though. why would it be a seperate thing.
- cadamsdotcom 3mo agoBecause writing code that passes a test is different to passing the test while also noticing and performing a refactor.
- spion 3mo agosometimes you forsee the code developing vastly differently between the two copies so you don't want the refactor. it really all comes down to lack of online learning and contextual awareness; memento mori notes are about as effective as developers with no expertise reading the design patterns book (well, ok, they are effective, but not sufficiently and not always directionally correct as its hard to accurately encode the nuance with language)
- johnxianren 3mo agoSo true. So when you see that split coming, do you drop a quick warning, or just copy-paste and leave the mess for tomorrow?
- 3mo ago
- orsenthil 3mo agoI find it amusing that people who are talking about Dark Software Factories, are talking about productivity in terms of number of pull requests or commits as a unit. If we are going in the Dark Software Factory route, why aren't we calling the code units as bos (bunch of shit) yet.
- dhorthy 3mo agoits the optimizing utilization instead of overall throughput all over again. eli goldratt talked about this in the 1970s[1]. we still haven't learned 1 - https://en.wikipedia.org/wiki/The_Goal_(novel) https://en.wikipedia.org/wiki/The_Goal_(novel)
- reinitctxoffset 3mo agoNo one serious about the idea can afford either vanity metrics nor ignorance of the code. The bar is higher, not lower, to operate with this much automation in the water supply. It's mostly a lot more math and a lot more work. It's an extreme form of any startup: you trade off capital for years of your life.
- throw10920 2mo agoInteresting, I'd like to discuss further - can you please add an email to your HN profile?
- dboreham 3mo agoSurely "pile of...?"
- _pdp_ 3mo agoI have mixed feelings about software factories! On one hand, our core product is just simply not fit for them at its scale. We've tried but the project is large enough to require human input for every change. But we have AI automations for light code refactoring, writing tests, UI changes etc. and they work. On the other hand, I started a number of small experiments to see how far software factories can be pushed and while the code produced so far is nothing spectacular I could easily imagine how this can be extended in the near future. Perhaps if you start from the ground up with the idea that the code will be written that way then you can come up with strategies and architectures that accommodate it. At least this is my thinking right now. Anyway, it is all open source and documented here https://relentless.works/ https://relentless.works/ I am not sure for long I will keep this running. I provide zero direction to where this is going. I have no idea what it will happen next. It is a fun experiment. I have another such experiment with a trading agent. I thought it will loose all of the money in short time. For a while it was stuck with no open positions after it lost a bit. I decided not to intervene and just observe the behaviour. Recently it opened new positions which was an interesting development. It is still loosing money (~ -3%) but it has not lost all of them and given the current market circumstances I would say this ain't bad at all. It just shows that perhaps we might be a bit impatient when it comes to AI. So I think it is probably possible to build software factories but we need new concepts and a bit of change of mindset. I hope this helps.
- bavell 3mo agoGood read and matches my experience pretty well, but I'd humbly suggest the author look up the definition of 'vertical' and 'horizontal' and swap their terminology + fix their videos :)
- dhorthy 3mo agoi humbly disagree - horizontal means touching one plane of the stack across, vertical means cutting down through it and touching multiple layers https://en.wikipedia.org/wiki/Vertical_slice https://en.wikipedia.org/wiki/Vertical_slice
- janalsncm 3mo agoEither you need to understand how your codebase works or you don’t. Claude can write the code for you but it can’t understand it for you. That part has to happen at human speeds. There are cases where you don’t have to understand everything, but I think that’s a more nuanced question. All of the above is true even if Claude writes perfect code.
- dboreham 3mo agoMy experience conflicts with this assertion. I've used Claude to achieve an understanding of two large codebases (that I mostly wrote, and certainly came up with most of the concepts therein) to the point that it's far superior to my understanding. I now get it to explain things to me that I have long forgotten.
- weatherlite 3mo ago> Either you need to understand how your codebase works or you don’t. It's an interesting point. We can also think about it perhaps as a non binary thing - you need X amount of understanding in a specific codebase to be effective. Even before LLMs in large codebases no one understood it all; but we at least mostly understood our own PRs and our own areas of expertise in the codebase.
- nevertoolate 3mo agoI have a very good mental model in my head of each and every codebase I have ever worked on. Nothing to do with my PRs or even my team’s direct responsibilities. I could always direct a thought provoking idea to challenge the existing status quo and get a reasonable nuanced answer from code owners on slack or during a water-cooler talk. I build that mental model also by reading the code. This is my job, so I know when to tear down what or how to respond to a random product idea immediately on a meeting to asses complexity.
- pydry 3mo agoBefore LLMs you either built good abstractions to make it possible to not understand large chunks of the code base or you flailed. A lot of the time people flailed. The one thing Ive never seen an LLM do well is shape clean, coherent abstractions. To be fair it's a rare human skill as well but in LLMs if they don't have a direct analog in their training data they flail.
- zhonglin 3mo agoFable, Sol can write any code, it is not the software factories is dead, AI built too many software factories....
- sathish316 3mo agoI call it the Intent-Implement-Quality problem. Software factories can implement anything given a one-liner requirement. That one-liner requirement can be a complete app/product, epic, feature, bug, design change or refactoring. But these one liner requirements are requirements coming from a human who has an intent or requirement or direction for the product to evolve in mind. Can Software factories manufacture intent that reflects the exact requirements of the person using it or their vision of how the product or software should evolve? Implementation in any language is as easy as generating a summary for an LLM. If coding is only math and there is only one way to translate a requirement into an implementation, the problem boils down to just providing or verifying generated intent. But, there is definitely more than one way to implement a thing and one of the ways leads to a design and architecture that is coherent with rest of the system, design and architecture that is extensible and evolves, code that you can keep in your head when you want to change things next time or tomorrow, software that can serve millions of users and securely guarantee millions of dollars of revenue. There is a combinatorial explosion of ways and the one right way is also subjective of the person and the problem. Software factories can of course improve Quality by generating more unit tests, more integration tests, fix code violations, generate Proof of work as videos, screenshots etc, but there is no feedback loop or test suite to verify and correct the other subjective notions of Quality.
- sathish316 3mo agoIf you’re working on an app or software that has few users, no revenue or minimal revenue, tolerance to bugs is higher and just another Claude prompt away, Software factories are a perfect fit. Most personal software or hobby software or 0-1 yet-to-find-PMF startups belong in this category. You can even take a stand that you’ll never look at the code and just ship. This is a perfect equilibrium for a Software factory, where the only input or feedback is a one-liner/specced requirement and the only output is outcome of whether the one-liner requirement worked or not. For almost any other software that does not fit this criteria, software factories are yet to solve the Intent and Subjective Quality problem.
- aleph_minus_one 3mo ago
- reinitctxoffset 3mo agoWho said we failed. If we had succeeded, would we tell you?
- swyx 3mo agothe talk version of this writeup was just released today: https://youtu.be/Ib5GBkD555M https://youtu.be/Ib5GBkD555M
- chicagopluto 3mo agoThis is a great write up. I buy the argument against the dark factory approach, but I would be curious to hear if their process has changed at all with Fable/GPT-5.6. I believe the keynote this is based on (www.youtube.com/watch?v=Ib5GBkD555M) was given while Fable was still banned and GPT-5.6 had not yet been released. Has HumanLayer found they still need to follow as deliberate of a pre-work planning process as before? Or do they find that Fable/GPT-5.6 are able to do more with slightly less handholding?
- dhorthy 3mo agoi did a write up on fable while it was out - it can do big refactors, but it does not know what to change without human steering. For that, you need humans to know what to ask for. lights off is still out for me https://x.com/dexhorthy/status/2064747631885398231 https://x.com/dexhorthy/status/2064747631885398231
- softwaredoug 3mo agoIf you built a real factory, you’d basically never want it to be dark. You’d want a culture of getting wrenches out to inspect the cars being built. You’d want to care about small details. Not because we need to build cars by hand. But because looking at the real product (cars, code, etc) is the best way to make the factory better. It’s the best way to know what problems aren’t being measured, what processes need to improve, and how automation can produce a better end product. https://softwaredoug.com/blog/2026/07/09/write-code.html https://softwaredoug.com/blog/2026/07/09/write-code.html
- dhorthy 3mo agodon't forget the andon cord
- ffsm8 3mo ago> If you built a real factory, you’d basically never want it to be dark This statement seems to be factually untrue as the fully automated factories in China are indeed not illuminated for most of the time In USA/Europe we just don't have that level of automation
- softwaredoug 3mo agoThe “fully automated” factories have humans involved in QA, maintenance and engineering. They often have to stop the factory and investigate and troubleshoot.
- ffsm8 2mo agoi mean youre most likely not wrong (i dont have insider knowledge of those places), however your previous comment said "never want it to be dark" - when the reality is that its normal operation _is_ dark. only when the automated QA detects issues or theyre doing random inspections are the lights turned up/on this is an outdated video for waht was state of the art a few years ago https://www.youtube.com/watch?v=MCBdcNA_FsI https://www.youtube.com/watch?v=MCBdcNA_FsI again, i dont have insider knowledge so i cannot speak on whats SOTA in chinese automated factories beyond whats covered publicly
- bisonbear 3mo agoTo me, this comes down to verifiability. How do we measure the quality of what an agent is doing on our codebase rather than simply measuring task accomplishment? > Verifying quality is orders of magnitude harder than "did the tests pass" Agree that agentic grading is the future here. Cognition's Frontier Code is probably the best large public benchmark at this. You attribute agent quality issues to RLVR's binary pass/fail, however I wouldn't be surprised if labs are already supplementing that with rubrics as rewards to train more 'tasteful' models like Fable. What can a practitioner do? I think there's promise in turning the optimization machine to the harness itself - building out a representative dataset of tasks on your repo, grading agent quality on them across various configurations, and optimizing [AGENTS.md / SKILLS.md / workflow / model / harness / tools] on that signal. High quality grading is still very hard, but it's more tractable at smaller, repo-level scale, and you can afford slower, more expensive verification for each task. You only need it to be right about your codebase's standards. > In fact, it's not hard to imagine that if a model could reliably tell good code from bad, it might have written the good version to begin with Pushing back slightly - detecting slop and discriminating quality is easier than generating it (why code review is so effective), and why grading is viable at repo eval scale even if it's much harder at RL scale. Everyone is flying blind. For example, I am genuinely interested in trying HumanLayer, but would likely want some harder evidence (beyond anecdotes) that it's actually making my agents more effective before rolling out to an enterprise team. I'm building this harness optimization loop @ https://stet.sh https://stet.sh if curious
- dsifry 3mo agoThis is why I built Metaswarm and Metareview and why they work so well, and the are building and supporting production infrastructures and sites for months on end. Give them a try: https://github.com/dsifry/metaswarm https://github.com/dsifry/metaswarm and https://github.com/dsifry/metareview https://github.com/dsifry/metareview
- ModernMech 3mo agoIn autonomous navigation there's a concept taken from sailing called "dead reckoning", where you just use an internal model of the robots dynamics to make controller commands. It works for short distances, but without feedback from sensors, the path the robot takes quickly diverges from the intended one as errors accumulate quadratically over the distance travelled. If the robot travels far enough without any external feedback, the localizer can "diverge", meaning its belief about where the robot is becomes wildly off from where the robot actually exists, making safe control virtually impossible. Now, if you've got a really good model you can get much further with dead reckoning compared to a worse model. But a model is not reality so no matter how good it is, so without feedback eventually you still run into this problem no matter what. I imagine that's a lot like what goes on in these agentic loops.
- ChicagoDave 3mo agoI’ve been postulating that knowledge of old school principles like Method One along with test driven development and domain-driven design are the external factors allowing a subset of teams to succeed. Absent those skills, the agile movement is flailing with GenAI.
- luciana1u 3mo ago[flagged]
- rahulladumor 3mo ago[flagged]
- becomevocal 3mo agoLike the "building blocks" mentality vs. "factory from scratch" mentality. Because... that's just how the real world works and if we are modeling intelligence off ourselves then why wouldn't that be the right approach?
- abusada 3mo agohttps://youtu.be/Ib5GBkD555M?is=3IFJRnshyZ0tFG_v https://youtu.be/Ib5GBkD555M?is=3IFJRnshyZ0tFG_v
- trenchgun 3mo agoCMD+F "reward hacking" -> 0 hits. Fail.
- Msurrow 3mo agoHonest question: For let’s say a senior developer having to go through all of these steps for the model to implement a feature: > Product Design > System Architecture > Program Design > Vertical Slices By the time it takes to go through these steps and agree with the agent, wouldn’t the senior dev not just be able to implement the feature by themselves? I mean, in general; I know there’s always features with complex logic etc, so the question is just about the “general case”
- gblargg 3mo agoThe compiler is the software factory. It builds the executable given the specification (code).
- yunbiao 3mo ago[flagged]
- sgt101 3mo agoThis is interesting and all, but why should we believe any of it? 1) This guy has a track record (confessed) of making shit up, yapping on about it, and pushing it on the innocent. He's done a bunch of damage with his bullshit and now wants us to pay attention again. I mean - something something off fella. 2) There is NO EVIDENCE AT ALL that his ideas are good. He's just making stuff up. Give me a reason. Also shouldn't he be ostracised and stripped of his wealth for his previous rubbish?
- romanovcode 3mo agoIt's just a long ad for his product - humanlayer.
- gyulai 3mo ago> NO EVIDENCE AT ALL It's funny how, when the hype is strong enough, the burden of proof around the need for EVIDENCE suddenly shifts. Normally, the burden of proof is on $NEWFANGLED_THING to prove it's better than $TRIED_AND_TESTED. Software dark factories where no one looks at code are that unproven newfangled thing and all he's really saying is that, in his experience/assessment, those don't work, so he's trying to find other modes of human-ai-collaboration that might actually work that capitalize better on the things (humans) that weren't broke and didn't need fixing when gen-ai coding came along; and then getting the word out about that. Under normal circumstances there would be absolutely no ground for any controversy around such a stance. Instead, he's having to contend with reactions like: "But $HYPED_UP_THING, and how great that is is all that anyone is talking about! Don't listen to him, he's just trying to sell you something (other than what everyone else is trying to sell). ...naysayer probably thinks he's smarter than everyone."
- sgt101 3mo agoIt's just he's qualified himself as an unreliable witness. This is not someone who deeply understands and thinks about what he does or what he advises. The things that I know about, that he writes about, where I know he's got it a bit wrong.... I find myself questioning if he's right or not. Because he's a persuasive writer. But winning arguments doesn't make you right. This isn't a supported, balanced discussion it's a self aggrandising mismash of commonplace ideas and issues. He's learned how to grab the mic, he believes he has the right to the mic, people are listening to him.. all three of these are wrong and until we call this out and put a stop to people like this shooting their mouths off damage is going to carry on being done.
- claud_ia 3mo ago[flagged]
- iamwil 3mo agoI've been building and running my software factory for 8 months now. Granted, there's no automated pulling down tasks and pushing PR right now (soon!). But after specifying what I want, it mostly goes to shipping on its own. On occasion, it does raise issues that I have to make a decision on. After doing systems evals on review, I've stopped looking at code during review for 4 months now. I do spend a lot of time up front specifying what I want. My prompts aren't one-liners, but rather an interview process where we work through all the open questions and ambiguity. I haven't hit the wall that the OP talked about (when agents just can't seem to make the right changes, and it's impossible for me to go in and change things manually). I used a lot of guardrails such as plan reviews, browser-based QA, adversarial reviews, unit tests, linters, typecheckers, post-commit hooks, and formal method traces. I also specified engineering principles that steers the code base to minimize state and side-effects: functional core; imperative shell, make impossible states impossible, use pure functional style, etc. There are times, when I can feel a part of the code base is messy without looking at it, because the agent will make recurring mistakes in the same part of the code base over time. What I found the agent was doing over time is that it's been layering state variables as requirements were discovered. So what helps is to ask it to refactor all these state variables into a single sum type. And if the state machine for it is complicated, I'll ask it to write a formal model of the state in Quint. Then I'll generate traces that get run as unit tests, and ask it to write the code against that. So while the code base isn't exactly Brownfield, it's over a year old now. As for the code base, there's a backend and a frontend. I think it helps that I established a clear pattern I wanted. You code are like memes: agents will just copy patterns they see in the code base. When it does have to create a new part of the system, I found Sonnet-level models tend to draw system boundaries in all the wrong places. Opus is better. I don't yet know about Fable. Happy to answer any questions about my workflow.
- chrisweekly 3mo agoI'd be interested in seeing more of your setup, eg if you published a long blog post and/or repo.
- iamwil 2mo agoOther people have expressed interest too. Will let you know when I'm done.
- jkwang 3mo agoInteresting framing. As coding agents move from demos to production, the bottleneck usually isn't the model but the harness around it: observability, rollback, and intent validation.
- felixlu2026 3mo ago[dead]
- antonvs 3mo ago> We haven't even hit AI yet, and there are already several loops in this picture. What a surprise! And here I was thinking that loops had only just been invented, to use in agent harnesses, as described in the famous paper “Loops Are All You Need.”
- feiz45607 3mo ago[flagged]
- spacecadet 3mo agoIt's a skill issue... governance and culture skills. Doesn't matter at this point what model you use... a thoughtful harness with appropriate levels of HITL and strong security, governance, discipline. "We" have a very successful dark factory running... no I cant tell you about it. It's trade secrets at this point. It was incredibly hard to build, roll out, and maintain. People have no idea that 2 years ago we hit peak performance with omni models. We won competitions with GPT-4o-mini, mini... it does not matter. Context matters. Discipline matters. Observability matters. Most of all, PEOPLE matter. All of the fools who tried this and laid people off. Fools. It's about empowering your team, making their lives easier, less frustrating, and higher impact.
- nassirkhan 2mo ago[flagged]