6 ms·
The Future of AI Software Development
- fuzzfactor 8mo agoLooks to me like the people that are filthy rich [0] can afford to move so fast that even the people who are very rich in the regular way can't keep up. [0] Which is not even enough, these are the ones with truly excess money to burn.
- bilekas 8mo agoI'm not sure you read the article, it's not referring to financials, but tech debt.
- fuzzfactor 8mo agoI like Fowler and reviewed it well. Are you assuming tech debt has no financial cost?
- bilekas 8mo agoOh sure but it just usually doesn't show up on a financial statement so just seemed a bit strange to be commenting on the financials is all, maybe I misunderstood your context.
- fuzzfactor 8mo ago>it just usually doesn't show up on a financial statement Exactly. That's one of the reasons it's gotten so out-of-hand. Edit: Thought I'd add that once wealth rises up to a stratospheric degree, the higher it is the less there is a need to get their money's worth. If you could put a dollar figure on it, the biggest can afford more technical debt than anybody theoretically, and they still won't actually have to deal with it. It'll be somebody elses' problem since it's not on the balance sheet. I guess I was commenting on the movers & shakers, not the financials specifically. Good observations you make though.
- adregan 8mo agoIn the section on security: > One large enterprise employee commented that they were deliberately slow with AI tech, keeping about a quarter behind the leading edge. “We’re not in the business of avoiding all risks, but we do need to manage them”. I’m unclear how this pattern helps with security vis-à-vis LLMs. It makes sense when talking about software versions, in hoping that any critical bugs are patched, but prompt injection springs eternal.
- bilekas 8mo ago> but prompt injection springs eternal. Yes, but some are mitigated when discoverd, and some more critical areas need to be isolated from the LLM so taking their time to provision LLM into their lifecycle is important, and they're happy to spend the time doing it right, rather than just throwing the latest edge tech into their system.
- ethin 8mo agoHow exactly can you "mitigate" prompt injections? Given that the language space is for all intents and purposes infinite, and given that you can even circumvent these by putting your injections in hex or base64 or whatever? Like I just don't see how one can truly mitigate these when there are infinite ways of writing something in natural language, and that's before we consider the non-natural languages one can use too.
- bilekas 8mo agoFull mitigation seems impossible to me at least but the obvious and public sandox escape prompts that have been discovered and "patched" out just making it more difficult I guess. But afau it's not possible to fully mitigate.
- lambda 8mo agoThe only ways that I can think of to deal with prompt injection, are to severely limit what an agent can access. * Never give an agent any input that is not trusted * Never give an agent access to anything that would cause a security problem (read only access to any sensitive data/credentials, or write access to anything dangerous to write to) * Never give an agent access to the internet (which is full of untrusted input, as well as places that sensitive data could be exfiltrated) An LLM is effectively an unfixable confused deputy, so the only way to deal with it is effectively to lock it down so it can't read untrusted input and then do anything dangerous. But it is really hard to do any of the things that folks find agents useful for, without relaxing those restrictions. For instance, most people let agents install packages or look at docs online, but any of those could be places for prompt injection. Many people allow it to run git and push and interact with their Git host, which allow for dangerous operations. My current experimentation is running my coding agent in a container that only has access to the one source directory I'm working on, as well as the public internet. Still not great as the public internet access means that there's a huge surface area for prompt injection, though for the most part it's not doing anything other than installing packages from known registries where a malicious package would be just as harmful as a prompt injection. Anyhow, there have been various people talking about how we need more sandboxes for agents, I'm sure there will be products around that, though it's a really hard problem to balance usability with security here.
- riffraff 8mo agoI think the title on HN doesn't reflect all that is in TFA, but rather the linked article[0]. Fowler's article is interesting tho. I do like the idea that "all code is tech debt", and we shouldn't want to produce more of it than we need. But it's also worth remembering that debt is not bad per se, buying a house with a mortgage is also debt and can be a good choice for many reasons. [0]: https://thenewstack.io/ai-velocity-debt-accelerator/ https://thenewstack.io/ai-velocity-debt-accelerator/
- senko 8mo agoI like the "cognitive debt" idea outlined here: https://margaretstorey.com/blog/2026/02/09/cognitive-debt/ https://margaretstorey.com/blog/2026/02/09/cognitive-debt/ (from a participant of the retreat) and especially the pithy "velocity without understanding is not sustainable" phrase.
- simonw 8mo agoYeah that editorialized title is entirely wrong for this post. Problem is the real title is "Fragments: February 18" which is no good here either. I suggest something like "Tidbits from the Thoughtworks Future of Software Development Retreat" (from the first sentence, captures the content reasonably well.)
- eru 8mo agoTech debt is totally misnamed. 'Tech debt' behaves more like equity than debt: if you project goes nowhere, the 'tech debt' becomes a non-issues.
- nthypes 8mo agoIMHO, it doesn't, but I have changed the title to avoid any confusion.
- senko 8mo agoWhat's with the editorialized title? The text is actually about the Thoughtworks Future of Software Development retreat.
- nthypes 8mo agoIMHO, it doesn't, but I have changed the title to avoid any confusion.
- simonw 8mo ago> LLMs are eating specialty skills. There will be less use of specialist front-end and back-end developers as the LLM-driving skills become more important than the details of platform usage. Will this lead to a greater recognition of the role of Expert Generalists? Or will the ability of LLMs to write lots of code mean they code around the silos rather than eliminating them? This is one of the most interesting questions right now I think. I've been taking on much more significant challenges in areas like frontend development and ops and automation and even UI design now that LLMs mean I can be much more of a generalist. Assuming this works out for more people, what does this mean for the shape of our profession?
- akkanzn 8mo ago[dead]
- AutumnsGarden 8mo agoI’ve become the same way. Instead of specializing in the unique implementations, I’ve leaned more into planning everything out even more completely and writing skills backed by industry standards and other developer’s best practices (also including LOTS of anti-patterns). My work flow has improved dramatically since then, but I do worry that I am not developing the skills to properly _debug_ these implementations, as the skills did most of the work.
- mjr00 8mo agoIMO debugging is a separate skill from development anyway. I've known plenty of developers in my career who were fully capable of writing and shipping code, especially the kind of boilerplate widgets/RPCs that LLMs excel at generating, yet if a bug happened their approach was largely just changing somewhat random stuff to see if it worked rather than anything methodical. If you want to get/stay good at debugging--again IMO--it's more important to be involved in operations, where shit goes wrong in the real world because you're dealing with real invalid data that causes problems like poison pill messages stuck in a message queue, real hardware failures causing services to crash, real network problems like latency and timeouts that cause services which work in the happy path to crumble under pressure. Not only does this instil a more methodical mentality in you, it also makes you a better developer because you think about more classes of potential problems and how to handle them.
- chadash 8mo ago> Will LLMs be cheaper than humans once the subsidies for tokens go away? At this point we have little visibility to what the true cost of tokens is now, let alone what it will be in a few years time. It could be so cheap that we don’t care how many tokens we send to LLMs, or it could be high enough that we have to be very careful. We do have some idea. Kimi K2 is a relatively high performing open source model. People have it running at 24 tokens/second on a pair of Mac Studios, which costs 20k. This setup requires less than a KW of power, so the $0.8-0.15 being spent there is negligible compared to a developer. This might be the cheapest setup to run locally, but it's almost certain that the cost per token is far cheaper with specialized hardware at scale. In other words, a near-frontier model is running at a cost that a (somewhat wealthy) hobbyist can afford. And it's hard to imagine that the hardware costs don't come down quite a bit. I don't doubt that tokens are heavily subsidized but I think this might be overblown [1]. [1] training models is still extraordinarily expensive and that is certainly being subsidized, but you can amortize that cost over a lot of inference, especially once we reach a plateau for ideas and stop running training runs as frequently.
- embedding-shape 8mo ago> a near-frontier model Is Kimi K2 near-frontier though? At least when run in an agent harness, and for general coding questions, it seems pretty far from it. I know what the benchmarks say, they always say it's great and close to frontier models, but is this other's impression in practice? Maybe my prompting style works best with GPT-type models, but I'm just not seeing that for the type of engineering work I do, which is fairly typical stuff.
- fullstackchris 8mo agoregardless its been 3 years since the release of chatgpt. literally 3. imagine in just 5 more years how much low hanging (or even big breakthroughs) will get into the pricing, things like quantization, etc. no doubt in my mind the question of "price per token" will head towards 0
- crystal_revenge 8mo agoI’ve been running K2.5 (through the API) as my daily driver for coding through Kimi Code CLI and it’s been pretty much flawless. It’s also notably cheaper and I like the option that if my vibe coded side projects became more than side projects I could run everything in house. I’ve been pretty active in the open model space and 2 years ago you would have had to pay 20k to run models that were nowhere near as powerful. It wouldn’t surprise me if in two more years we continue to see more powerful open models on even cheaper hardware.
- siliconc0w 8mo agoEven with the latest SOTA models - I still consistently find issues. Performance, security, memory leaks, bad assumptions/instruction following, and even levels of laziness/gaslighting/dishonesty. I spend less time authoring changes but a lot more time reviewing and validating changes. And that is using the best models (Opus 4.6/Codex 5.3), the OSS/flash models are still quite unreliable at solving problems. Token costs are also non-trivial. Claude can exhaust a $20/month session limit with one difficult problem (didn't even write code, just planned). Each engineer needs at least the $200/mo plan - I have multiple plans from multiple providers.
- raphaelmolly8 8mo ago[dead]
- christkv 8mo agoMy bet is that the amount of work needed per token generated will decrease over time and the models will become smaller for the same performance as we learn to optimize so cost and needed hardware will go down
- anthonypasq 8mo agoWhat is up with all this nonsense about token subsidies? Dario in his recent interview with Dwarkesh made it abundantly clear that they have substantial inference margins, and they use that to justify the financing for the next training run. Chinese open source models are dirt cheap, you can buy $20 worth of kimi-k2.5 on opencode and spam it all week and barely make a dent. Assuming we never got bigger models, but hardware keeps improving, we'll either be serviing current models for pennies, or at insane speeds, or both. The only actual situation where tokens are being subsidized is free tiers on chat apps, which are largely irrelevant for any sort of useful economic activity.
- deleted 8mo ago[deleted]
- simonw 8mo agoThere exist a large number of people who are absolutely convinced that LLM providers are all running inference at a loss in order to capture the market and will drive the prices up sky high as soon as everyone is hooked. I think this is often a mental excuse for continuing to avoid engaging with this tech, in the hope that it will all go away.
- louiereederson 8mo agoReferring to my earlier comment, you need to have a model for how to account for training costs. If Anthropic stops training models now, what happens to their revenues and margins in 12 months? There's a difference between running inference and running a frontier model company.
- simonw 8mo agoTraining costs are fixed. You spend $X-bn training a model and that single model then benefits all of your customers. Inference costs grow with your users. Provided you are making a profit on that inference you can eventually cover your training costs if you sign up enough paying customers. If you LOSE money on inference every new customer makes your financial position worse.
- deadbabe 8mo agoThere have been some back of the napkin estimates on what AI could cost from the major platforms once no longer subsidized. It does not look good, as there is a minimum of a 12x increase in costs. Local or self hosted LLMs will ultimately be the future. Start learning how to build up your own AI stack and use it day to day. Hopefully hardware catches up so eventually running LLMs on device is the norm.
- taeric 8mo agoI really hate that we allowed "debt" to become a synonym for "liability." This isn't a case where you have specific code/capital you have borrowed and need to pay for its use or give it back. This is flat out putting liabilities into your assets that will have to be discovered and dealt, someday.
- greymalik 8mo agoThe headline misrepresents the source. It’s not the title of the page, not the point of the content, and biases the quote’s context: “ if traditional software delivery best practices aren’t already in place, this velocity multiplier becomes a debt accelerator”
- nthypes 8mo agoIMHO, it doesn't, but I have changed the title to avoid any confusion.
- acomjean 8mo agoSo do we need new abstractions / languages? It seems clear that a lot of things can be pulled together by AI because it’s tedious for humans. But it seems to indicate that better tooling is needed.
- mamma_mia 8mo agomamma mia! out with the old in with the new, soon github will be like a warehouse full of old punchcards
- clockworkhavoc 8mo ago[flagged]
- g8oz 8mo ago1. Singham is not a fugitive from American justice just yet - although refusing to cooperate with Congress may lead him to be. 2. Is it a problem if a rich guy funds activities in America that suspiciously align with a foreign power? That has interesting implications for many pro Israel billionaires and organizations. 3. Only a paranoid MAGA troll would characterize the left wing groups he funds as domestic terrorists. Code Pink? Pro Palestinian protest groups? Come on.
- PaulHoule 8mo agoGet over your FOMO: I walked into that room expecting to learn from people who were further ahead. People who’d cracked the code on how to adopt AI at scale, how to restructure teams around it, how to make it work. Some of the sharpest minds in the software industry were sitting around those tables. And nobody has it all figured out. People who say they have are trying to mess with your head.
- tcgv 8mo agoThat’s fair at the “adopt AI at scale / restructure orgs” level. Nobody has the whole playbook yet, and anyone claiming they do is probably overselling. But I’d separate that from the programmer-level reality: a lot is already figured out in the small. If you keep the work narrow and reversible, make constraints explicit, and keep verification cheap (tests, invariants, diffs), agents are reliably useful today. The uncertainty is less “does this work?” and more “how do we industrialize it without compounding risk and entropy?” I wrote up that “calm adoption without FOMO, via delegation + constraints + verification” framing here, in case it helps the thread: https://thomasvilhena.com/2026/02/craftsmanship-coding-five-stages-of-grief https://thomasvilhena.com/2026/02/craftsmanship-coding-five-...
- PaulHoule 8mo agoI use agents all the time but I keep my feet on the ground. The thing is doing that you do not get the radical explosion in productivity that influencers want you think they are getting.
- empath75 8mo agoSo here are a few things i have been thinking of: --- It's not 2 pizza teams, it's 2 people teams. You no longer need 4 people on a team just working on features off of a queue, you just need 2 people making technical decisions and managing agents. --- Code used to be expensive to create. It was only economical to write code if it was doing high value work or work that would be repeated many times over a long period of time. Now producing code is _cheap_. You can write and run code in an automated way _on demand_. But if you do that, you have essentially traded upfront cost for run time cost. It's really only worth it if the work is A) high value and B) intermittent. There is probably a formula you can write to figure out where this trade off makes sense and when it doesn't. I'm working on a system where we can just chuck out autonomous agents onto our platform with a plain text description, and one thing I have been thinking about is tracking those token costs and figuring out how to turn agentic workflows into just normal code. I've been thinking about running an agent that watches the other agents for cost and reads their logs ono a schedule to see if any of what the agents are doing can be codified and turned into a normal workflow, and possibly even _writing that workflow itself_. It would be analogous to the JVM optimizing hot-path functions... --- What I do know is that what we are doing for a living will be near unrecognizable in a year or two.
- ShazCoder 8mo agoThis sounds very fascinating. One of the more interesting ideas I have come across.
- tabs_or_spaces 8mo ago> Will this lead to a greater recognition of the role of Expert Generalists? I've always felt that LLMs can make you average in a new area/topic/domain really quickly. But you still need expertise to make the most out of the LLM. Personally, I'm more interested in whether software development has become more or less pay to win with LLMs?
- mehagar 8mo agoIt's refreshing to hear people say "We're not really sure" in public, especially from experts. I agree that AI tools are likely to amplify the importance of quick cycles and continuous delivery.
- markoa 8mo agoOne thing that I'm sure of is that the agentic future is test-driven. Tests are basically executable specs the agent can follow and verify against. When we have solid tests, the agent output is useful and we can trust it. When tests are thin or missing, the agents still ship a lot of code, but we spend way more time debugging and fixing subtle bugs.
- mattmanser 8mo agoGreat, so now I have to design the API for the AI, think of all the edge cases without actually going through the logic, and then I'm invariably going to end up with tests tightly coupled to the implementation.
- deleted 8mo ago[deleted]
- mchaver 8mo agoThis is why I think they work so well with strongly typed languages like Haskell and OCaml. You say do this until it compiles and passes a set unit tests for business logic. I find I am using even more verification tools like JSON schema validators. The more guardrails and hard checks you give an agent, the better it can perform.
- the_harpia_io 8mo ago[flagged]
- wseqyrku 8mo ago> When I began in software in the 1980s I was dismissed as an “object guy” by database folks and as a “data modeler” by object folks. I've since been dismissed as a “patterns guy”, “agile guy”, “architecture guy”, “java guy”, “ruby guy”, and “anti-architect guy”. I'm now a past-it gray-beard surviving on drinking the intellectual blood of my younger colleagues. It's tasty. I don't think you can find that level of ego anywhere in the software industry or any other industry for that matter. Respetc.
- trashymctrash 8mo agoi was initially confused because i couldn’t find it in the article, but then i found what you quoted here: https://martinfowler.com/articles/expert-generalist.html https://martinfowler.com/articles/expert-generalist.html
- wseqyrku 8mo agoYeah, sorry my head got dizzy I didn't realize I clicked away from the linked article.
- Copenjin 8mo agoHe forgot the "micro-shit" guy. At least he is not uncle Bob but not really much far from it. Slop well before AI.
- tcgv 8mo agoMartin’s framing (org and system-level guardrails like risk tiering, TDD as discipline, and platforms as “bullet trains”) matches what I’ve been seeing too. A useful complement is the programmer-level shift: agents are great at narrow, reversible work when verification is cheap. Concretely, think small refactors behind golden tests, API adapters behind contract tests, and mechanical migrations with clear invariants. They fail fast in codebases with implicit coupling, fuzzy boundaries, or weak feedback loops, and they tend to amplify whatever hygiene you already have. So the job moves from typing to making constraints explicit and building fast verification, while humans stay accountable for semantics and risk. If useful, I expanded this “delegation + constraints + verification” angle here: https://thomasvilhena.com/2026/02/craftsmanship-coding-five-stages-of-grief https://thomasvilhena.com/2026/02/craftsmanship-coding-five-...
- zi2zi-jit 8mo agoWe are saying goodbye to Scrum or Agile for sure. That is designed by human and too slow and performative.
- _alaya 8mo ago> The more radical possibility is that source code as we know it could become a transient artifact, generated on demand and never stored. The retreat was divided on this. Some saw source code disappearing within a decade. Others argued that deterministic validation requires a stable artifact to test against, and that artifact is effectively source code regardless of what we call it. Here’s a free idea I’ve had that I have no idea how to implement. I hope somebody much smarter than me will come along, think it’s a great idea, and steal it. I highly encourage you to do so, and I wish you well. The idea is to have some kind of substrate—like a superpowered AST—that is the true code: the thing that actually gets compiled and run. Humans never look at this directly. Instead, we look at a representation of this code, and we can toggle between different representations of it. I’m borrowing ideas from topology in mathematics here: if I look at a shape one way, I should be able to transform it into a different shape, but isomorphically, everything is still the same. That would let me look at the same thing in different ways, understand it from different angles, critique it more easily, and maintain it more easily. Gemini tell me that this idea has already been tried in the past? Projectional Editing? Intentional Programming?
- opsmeter 7mo ago[dead]