13 ms·
Wut? I pilot LLMs all day but there's no way in hell I'd agree to be at the helm of a finance product. That first pillar is still there. Maybe the author isn't
by iandanforth 4mo ago
Wut? I pilot LLMs all day but there's no way in hell I'd agree to be at the helm of a finance product. That first pillar is still there. Maybe the author isn't aware of the impact they have, but I know, with the evidence of reverted PRs, that when I step outside my area of deep knowledge I can no longer call BS on the agents. Our most capable agent, with access to the same kind of distributed systems the author talks about, is regularly wrong, frequently myopic, and just outright dumb constantly. It's the expertise of engineers on the team that push it back on track.
- keyle 4mo ago[flagged]
- iandanforth 4mo ago"I ended up working in software development roles in the domains of finance, bookkeeping and payment processing, where I had great autonomy and a close and candid relationship with Product Managers and stakeholders. I learnt a lot about the domain and how to effectively write programs for it: PCI compliance, double-entry ledgers, escrows, reconciliation, payment lifecycles, bank transfer idempotency, etc. It was, then, obvious that I should focus my career on becoming an expert on that domain to stand out as a professional and differentiate myself in a field that showed signs of an increasing need for domain specialists."
- stuaxo 4mo agoThe backend is the bit that "does stuff" so it's the part that needs to be correct. He said "Last year, I got hired by a company in the finance workspace.".
- jalev 4mo agoUnfortunately every software related industry is embracing LLM/Codegen. Your banks, fintechs, insurance. Everyone. Your concerns are the same I'm having, yet it's regularly dismissed or hand-waved away as "don't worry about it the delivery velocity/ROI is worth it"
- Hamuko 4mo agoAre banks that concerned about velocity? Because moving fast and breaking things in the banking sector can get extremely expensive. It's also not a who-gives-a-shit industry like operating a taxi service or hosting images, but a very tightly regulated sector.
- jalev 4mo agoI might have been a bit broad with the brush. I can't speak for banks, but I can speak for the the fintech/money-movement space (e.g. Remitly, Wise, Revolut). It's a race to get first-to-market for backend integrations/features. It's given rise to a culture of "move fast break things" where safety is only for some core features, but absolutely not for the constellation of other services we provide. Failure rates have increased almost a percentage point since Codegen/LLM adoption was mandated from up top. You would think regulators would be on top of this, but our industry runs on all actors "self reporting" their outages. Most don't unless they can't hide it (>1h)
- mrkeen 4mo ago'Keeping up with regulations' may as well be a separate field from the core stuff. It has the same pressures as any other development effort. Managers will want the integration to the KYC service LLM'd as quickly as possible.
- bigthymer 4mo ago> Are banks that concerned about velocity? Yes
- hn_throwaway_99 4mo agoNot in the universe where I live. Having worked in a variety of web tech, and then working at a fintech with a partner bank, traditional banks move incredibly slow compared to nearly every other tech company out there, and for good reason.
- simon84 4mo agoIt's not so much about velocity or quality, both of which LLM do (or will) provide. The real question is about accountability and liability. When a major data leak is going to happen, who will they sue or fire ? That is the value engineers provide. They understand, confirm, and take ownership.
- lelanthran 4mo ago> Wut? I pilot LLMs all day but there's no way in hell I'd agree to be at the helm of a finance product. Dunno how much longer that is going to remain true for your specific employer - all the fintech companies I deal with personally have had some sort of AI account for their devs since last year. Even places like jane street have employees posting blogs (one of which was on HN frontpage about 60m ago) saying they mostly direct agents. How long do you think your specific employer is going to hold out?
- iandanforth 4mo agoSorry if I was unclear. I don't work in finance. I do work with agents. I think expert engineers in finance who are guiding agents are adding a lot of value because of their knowledge of finance. Because I lack that knowledge of finance, even given access to agents, I would not accept a role guiding agents in a finance company because I wouldn't be able to guide the agents well and my/our output would be bad.
- dakiol 4mo ago> It's the expertise of engineers on the team that push it back on track. But how are you so sure your colleagues are not more "expert" than you? Prior LLMs there was room for very good engineers and mediocre engineers to work together in 99% of the companies out there. With LLMs, only the "best" engineers will survive, because nobody needs mediocre engineers anymore. This being HN, I imagine every engineer reading this thinks they are in top the 10-5% of their company/city/country, and therefore they think they are not "mediocre" engineers that can get affected by the introduction of LLMs. Statistically, they are probably wrong. So, it's all about ego. Chances are you are not a rockstar and LLMs will eventually take over your job. As usual, the only winners here are corporations and executives. Most of us are the last monkeys in the chain, and so we'll get screwed.
- aleqs 4mo agoThe corporations and executives are already winning if you swallowed the concept of 'rockstar' engineer. Sure there are more and less experienced engineers, but even interns can and often do provide good input and spot mistakes made by seniors. The 'rockstar' engineer at most tech companies simply equates to the somewhat autistic guy with a brown nose who's working 15 hour days for a pat on the head from management (and making many mistakes in the process).
- throwaway7783 4mo agoEven if we forget "rockstar", there are certainly different levels of engineers. More experience doesn't automatically mean better either. That is not to say experience doesn't matter. It matters quite a bit. Sure , good interns can sometimes have good feedback or spot mistakes. But not consistently enough. All of this to say that it's not just experience that makes one a better engineer.
- aleqs 4mo agoExperience is one of the only objective signals we have, but you're right it's not the only one. I've seen plenty great junior engineers and interns, and plenty of incompetent staff/principal engineers.
- deleted 4mo ago[deleted]
- t34t34r43 4mo agoPosting this under a burner so I don't dox myself: I work in FinTech on a regulated product. We have access to Mythos. Mythos identified part of our codebase that it confidently asserted was not complaint with a particular regulation and we were at grave risk by allowing it to operate the way it was. Except this was not the case, it had of course hallucinated what the regulation actually required (I know this because the code in question had already been reviewed by human counsel). This is (supposedly) the most bleeding-edge model available. We use a lot of genAI to help us write code, but there is no way in the mid-term we could ever rely on these tools to actually build compliant financial products. We'd have to be totally mad. Yes, lots of Fintech companies are using these agents to accelerate, but anyone who's using them to actually ship product without a human actually digging into it is opening themselves up to a world of risk.
- galactushonor 4mo ago> it had of course hallucinated what the regulation actually required Did it do the correct job once you put the regulations doc(s) in the context?
- loloquwowndueo 4mo agoWhat I usually do when in doubt is challenge the AI. “Please quote the section of regulation the product is non compliant with”. It usually admits it hallucinated the whole thing.
- mattmanser 4mo agoIt sometimes says that even if it hasn't though, so like everything with LLMs, you can't actually rely on that.
- Chu4eeno 4mo agoYou should just hit retry, usually they either latch onto the correct part of the latent space (surprisingly often) or they admit they don't know (or call a tool, depending on the model).
- znpy 4mo agoYou pilot LLMs all day but that might not last. A lot of companies are investing money on “ai factories” that are join to automate a lot of software development (that is, steer LLMs) on the basis of jira tickets (or linear/trello cards or whatever).
- iandanforth 4mo agoFunny you mention that, since that's what I pilot LLMs to build :)
- SlinkyOnStairs 4mo ago> That first pillar is still there. Maybe the author isn't aware of the impact they have, but I know, with the evidence of reverted PRs, that when I step outside my area of deep knowledge I can no longer call BS on the agents. Our most capable agent, with access to the same kind of distributed systems the author talks about, is regularly wrong, frequently myopic, and just outright dumb constantly. It's the expertise of engineers on the team that push it back on track. I'd posit there's another layer. You have domain knowledge, certainly. But more valuable still is the wisdom to find more. Anthropic and OpenAI can stick financial regulations in the training data all they want, but the AI systems will never learn to anticipate the future, or reach out to clients, partners, or regulators in complicated situations.
- baq 4mo ago> AI systems will never learn to anticipate the future Citation needed. I don’t see any reason these systems shouldn’t be able to speculate; indeed some would say that’s all they do, even about the past.
- micromacrofoot 4mo agoa year ago I would have agreed, but the gap is getting smaller all the time... these things can do 90% of the work, and how many people does a company really need for the remaining 10%? certainly not as many as they needed before
- realusername 4mo agoThe things can do 90% of the work ... but only if used by the right people. I've seen first hand what less experienced developers produce using the same models, your 90% accuracy suddenly drops to 50%...
- Quothling 4mo agoWith opus 4.8 we're frankly aproaching the 100% of the work, but only if tasked by the right people. A decade ago I worked as an enterprise architect and left it because I preffered coding. Now I'm an enterprise architect again, and we're at the point where I've setup a Microsoft Fabric and integrated a ADLS Gen2 with a Lakehouse building Dimension and Fact tables for our Business Intelligence people with Cowork. A month ago I didn't know what Dimension and Fact tables were in a datawarehouse and now I've not only setup a flow for it I've made it more accurate than what they had before because I understood how BC365 worked and the previous consultants didn't. We had a PoC in place to get fabric, it had like 500 hours allocated for what I did in a week with cowork, and my product is actually on secure vnet network with Azure identity security with both a test and a production environment delivering actual data. Cowork even made the damn powerpoint slideshows for decision makers. The single saving grace right now is that it apparently isn't easy for everyone to do this yet. But I didn't use a whole lot of my knowledge on software engineering to make any of it happen, not even the pandas and arrow code that moves the data behind the scenes. I mainly used my knowledge of NIS2 compliance and general data architecture in a step-by-step process. To me anyone with common sense should be able of doing this, and I really don't think I'm special... but then I teach other people AI at our company and they can barely get it to create a running program. Which is fine for now, but I have to work another 20ish years before I retire, and by then a lot of young people will have grown up with AI, and like I said, I'm not special. I think the only thing that differentes me is that I mash the buttons until it works but also have decades of security and compliance hammered into me.
- abhgh 4mo agoReg PRs - for the ones with complex requirements what I am seeing is that time to initial PR is very short, and a ping-pong between the reviewer and developer begins, because in my cases (not all) the developer vibe-coded parts, and they didn't really understand the requirements deeply or their code, and it takes multiple iterations for them to fix it. You can argue this is a human problem but this is the net effect I'm seeing. I am not sure but for complex cases it seems to me that the earlier sum of moderately long PR time + moderately long review time has been replaced by very short PR time + even longer review time. I am not sure if there's a net gain in these cases. Sometimes even if the code is functionally correct, it's verbose enough (e.g., too many intermediate functions) that I think they will impact future reviews.
- jkwang 4mo ago[flagged]
- bwfan123 4mo ago> I pilot LLMs all day Love the metaphor. Planes are sophisticated machines capable of auto-piloting, but humans are still needed to ultimately pilot the beast.
- esafak 4mo agoThere is a product called Microsoft Copilot...
- PantaloonFlames 4mo agoa slightly different metaphor. Copilot suggests it is next to you, helping you pilot... something else. The computer? The system? But "piloting the LLM" changes the relationship. The LLM is the thing that is being piloted.
- djeastm 4mo agoI think their metaphor is apt. Microsoft Copilot is, in modern terms, the "harness" that "co-pilots" the LLM it manages sessions for, with you as the "pilot". So you are still piloting the LLM. At least that's how it reads to me.
- chipsrafferty 4mo agoI thought it was, you're the pilot, copilot is your helper.
- altmanaltman 4mo agoWhat about autonomous drones then? They are not planes? The real reason we have pilots in commerical jets is that they are a failsafe because a mistake in your automated system can kill 100+ people and guess what? It still does from time to time even with a human piloting the "beast." It is not impossible to have fully automated planes. We just cannot afford to make any mistakes otherwise people will die so we keep pilots as failsafes. Nobody will die if your claude code hallucinates.
- abustamam 4mo agoYeah I'm constantly shocked at how simultaneously smart and dumb Opus can be. It can tell me a LOT about my codebase but it will miss very critical clarifications that I begin with. And when I call it out it obviously remembered it, it just ignored it.
- jrockway 4mo agoI agree with this experience. LLMs are great and save me a lot of time, but they need frequent nudges to avoid going down a completely wrong path. I just don't feel like the management dream of "every engineer has 3 agents working for them full time" is quite a reality yet. I'm not saying it won't get there, or that I feel secure being a software engineer until I'm of retirement age, but I also think it's important to understand the limitation of the tools. You do need to know your codebase. You do need to iterate on small chunks of it at a time. You do need to carefully understand every line of code you're putting into production. LLMs are amazing at generating a lot of proposals, but you need to carefully consider each one. Most surprising to me about the article was the desire for OP's company to use AI for design docs. I feel like AI-generated design docs are some of the worst -- basically treating English as a programming language. They aren't enjoyable to read, and they often miss the forest for the trees. A human written sketch explaining why we're here and what we're working towards is still meaningful and important. If you want code-level details of every decision and algorithm, we have code for that. I have mixed feelings on whether these documents are useful LLM inputs. I did a project where I carefully paired with Claude Code on producing a specification that another model would actually implement. I'm not sure it saved me any time, and it was very un-fun. (I kind of blame Opus 4.7 xhigh for this. It ain't speedy.) I feel like I can nitpick code to get exactly what I want, but defining exactly what I want an auto-mode LLM to go and do, in English, is much more difficult. I don't think the PLAN.md I generated would have been useful for a human trying to understand the system (too verbose), and Claude Code still made its usual mistakes that I have reminded it a billion times not to make (t.Context() in tests, not context.Background()!), so I'm just not sure it was worth it. I would say I probably wouldn't do it again in the near future. A rough sketch to get humans on board and to get the high level details worked out, written by hand, and then pairing with the LLM on actually typing in the code seems the most productive to me. But I do try to go outside my comfort zone once in a while to test the edges of these tools. They are very impressive and are worth a lot of the hype. (I know I will never write a YAML file again. I hate it more than anything, and Claude is amazing at it. But I worry I wouldn't feel the same way if I hadn't already had 8 years of k8s experience.)
- deleted 4mo ago[deleted]
- root-parent 4mo ago[dead]
- quijoteuniv 4mo agoNorwegians have a saying: “Den som er ferdig utlært, er ikke utlært – men ferdig.” Meaning if you are finished with learning the one that is finished is you. Typical scandinavic hard cold truth… I understand the frustration of spending years nurturing a skill and then seeing its value decline.But this isn’t really an LLM problem. The same thing happened to factory workers, typists, draftsmen, and many others before. The technology changes, but the underlying issue is the economic system we live in, where the market can suddenly decide that something you’ve spent years mastering is worth much less than before. LLMs are not creating that dynamic. They’re just accelerating it.
- throwaway201606 4mo agoIn software dev, for a big finance corp I like your comment, want to try to expand on it Comment long but there is a TL;DR at the bottom My theory is that there are 4 areas to domain knowledge worth taking about here - there may be more but I like 2*2 matrices 1) explicit internal requirements - core of how the how the app should work towards achieving your business objectives - code expresses what should be done and to a pretty large extent, why it should be done - from business unit requirements - we are building a tool to do “X” 2) implicit internal requirements - core of how the how the app should work towards achieving your business parameters and constraints eg profit = selling price - ( total of costs ) - code expresses what should be done but really can’t express why. At best it is in the comments eg if market is EU then tax = 30% (or some value for a table), AI can see what is being done but rationale is not explicit 3) implicit internal requirements - core of how the how the app should work towards achieving your business constraints - code expresses what should be done but really can’t express why. At best it is in the comments eg if item is “rocket” , shipping = $1m ( we only make rockets in Antarctica and shipping from there is $1m) 4) implicit external requirements - core of how the how the app should work towards achieving your business constraints - code expresses what should be done but really can’t express why. At best it is in the comments eg if item is “rocket” , add a 3 month gating stage to get approval from government to sell the item and do not collect payment till gating approved - AI can see the code but has no idea why it has to be done These come from partners, regulation, compliance, auditability etc So, my theory AI can be good at the explicit stuff trivially (1, 2) but cannot be good at the implicit stuff (3,4) It might be able to figure out implicit stuff needs to be done but will probably not be able to figure out why it needs to be done and it will definitely not be able to definitively figure out edge cases for when to do it / not do it As long as you focus on implicit stuff, you will be fine for a little bit TL;DR - become good and keep being good at being the person who understands the implicit external drivers of software dev
- raducu 4mo ago> is regularly wrong, frequently myopic, and just outright dumb constantly. I give LLMs snippets of text messages exchanges with my wife and I can't believe how dumb the LLMs are of getting basic facts right let alone nuance. I'm 100% not one those "LLMs are just stochastic parrots" people, but coding and coding-like activities are extraordinarily well fit for LLMs, but for things that there's less training data, LLMs probably do a lot worse
- cookiengineer 4mo agoAn LLM agent will only be as good as the environment it operates in. If you build your environment to be specification based, you have to make sure you have good specifications. If your "memory solution" uses freeform markdown notes, you already lost from the start. Also choose languages with good unit testing built-in, and languages with unified code styles, and unified toolchains. If you use C++, assume that there's a million ways to build your algorithm. If you use JS, assume 10 different build pipelines. If you use java, assume bloat by dependency hell. LLMs mimic the ecosystem's variety and variadicity(?). Languages like Go shine so well because it's a very opinionated language, where there's only one proposed way on how to implement things. And that's a good harness to begin with. LLMs are like children on the playground. You have to build better rulesets and fences to make them behave how you expect them to. Also, check out qwen3.6 coder and heretic models. 30b is the sweet spot for coding and unit testing. For planning and designing, gemma4 is pretty good.
- insane_dreamer 4mo agoit doesn't much matter what you and I think about LLM's quality and output it only matters what upper management think, and its clear that more and more companies value "good enough" and reducing costs over "good"
- deleted 4mo ago[deleted]
- yes1would 4mo agoLmao please do not say you "pilot" what is a glorified ELIZA. You're a joke that people will have forgotten about in 2 years.