9 ms·
Posting this under a burner so I don't dox myself: I work in FinTech on a regulated product. We have access to Mythos. Mythos identified part of our codebase th
by t34t34r43 4mo ago
Posting this under a burner so I don't dox myself: I work in FinTech on a regulated product. We have access to Mythos. Mythos identified part of our codebase that it confidently asserted was not complaint with a particular regulation and we were at grave risk by allowing it to operate the way it was.
Except this was not the case, it had of course hallucinated what the regulation actually required (I know this because the code in question had already been reviewed by human counsel). This is (supposedly) the most bleeding-edge model available.
We use a lot of genAI to help us write code, but there is no way in the mid-term we could ever rely on these tools to actually build compliant financial products. We'd have to be totally mad. Yes, lots of Fintech companies are using these agents to accelerate, but anyone who's using them to actually ship product without a human actually digging into it is opening themselves up to a world of risk.
- galactushonor 4mo ago> it had of course hallucinated what the regulation actually required Did it do the correct job once you put the regulations doc(s) in the context?
- loloquwowndueo 4mo agoWhat I usually do when in doubt is challenge the AI. “Please quote the section of regulation the product is non compliant with”. It usually admits it hallucinated the whole thing.
- mattmanser 4mo agoIt sometimes says that even if it hasn't though, so like everything with LLMs, you can't actually rely on that.
- Chu4eeno 4mo agoYou should just hit retry, usually they either latch onto the correct part of the latent space (surprisingly often) or they admit they don't know (or call a tool, depending on the model).
- deanc 4mo agoI've worked on projects in the airline and health industry which are highly regulated too. The regulations can be incredibly difficult to process and implement, and make sure you adhere to everything correctly. I've been involved in multiple scenarios where people have made false assertions about compliance or lack of. I'd still place a bet that the SOA models make _far_ less mistakes than humans.
- deleted 4mo ago[deleted]
- genxy 4mo agoThey might make fewer mistakes, but they aren't evenly distributed. They don't use logic when making mistakes, it is gaps in the training data and now large of a span they have to bridge in the latent space. Just as they aren't smart like humans, they aren't stupid like humans. Don't mistake rate for quality.
- Terr_ 4mo agoYeah, this starts to overlap with some autonomous vehicle stuff, where I like to say that the rate of errors is not the shape or distribution of errors. We have long historical experience and innate tools for detecting and mitigating errors made by humans. If we can't apply those to automation, then even fewer total mistakes may end up being a worse outcome.
- realusername 4mo ago> I'd still place a bet that the SOA models make _far_ less mistakes than humans. Well too bad, the problem is that they also produce things much faster than humans so errors will compound quicker.
- sillyfluke 4mo ago>I'd still place a bet that the SOA models make _far_ less mistakes than humans. Genuine question: your top coder seems to be producing the most error-free code from your perspective, has the deepest knowledge of the architecture and codebase, and is faster on the trigger than the others. But your top coder has proven and verifiable dementia, where they will confidently assume the existence of apis and code that do not exist, mix up the purpose of others and forget other things, and you can't predict when and how they will introduce errors into the system or the severity of such errors. Are you really comfortable letting this person with dementia generate most of your codebase in the airline and health industry? I also hope you have an iron-clad agreement that prevents the model provider from doing silent updates because all your evidence of correctness you collected thus far goes out the window in that case. Another genuine question: You have witnessed a human coder and the AI you're using make the same important mistake. Assuming you do not have the time and resources to retrain, fine tume, and test your frontier model: Who would you trust not to make the same mistake multiple times in the future after you have warned them that their job depends on it, the AI or the human?
- SuperV1234 4mo agoIs that all that Mythos did? Did it find any real potential issue, optimization/simplification opportunities, or sparked any thought-provoking discussion within your organization? Or was it purely a net negative experience?
- troupo 4mo agoIn regulated industries none of those matter if the tool invents compliance issues or breaks compliance. The only thought-ptovoking discussion should be "why the hell do we have this stochastic parrot anywhere near out codebase"
- brookst 4mo agoOdd take. So if it identified 17 real gaps and helped fix them, the fact it was wrong about one gap, and the appropriate humans caught it and no harm was done, the whole thing is useless? Not saying that is the situation, I don’t know. But if “one error is too many” is your point of view… do you think the humans in these orgs are 100% perfect 100% of the time?
- troupo 4mo ago> So if it identified 17 real gaps and helped fix them, the fact it was wrong about one gap, and the appropriate humans caught it and no harm was done How many gaps have humans not caught? > But if “one error is too many” is your point of view Yes, in regulated industries "one error is too many" is the only right approach. Yes, humans also make errors, and there you have a range of options: from tracing and finding the causes of the error (and tightening processes) to literally jailing those responsible. Your hallucination machine will happily "identify" 17 gaps, and create 34 more. And no, there are no processes to make it better. The "make no mistakes" incantation will happily be ignored for obvious reasons, regardless of how many forms of it you throw at it.
- ToValueFunfetti 4mo agoIt doesn't seem like you're engaging with the material circumstances described above. What does it mean for a human to not catch that a part of a codebase is actually compliant with regulations? What does it mean for the hallucination machine to create 34 more gaps when it doesn't appear to have more than read access? How would it not be useful to have a machine that identifies 17 real crimes that your highly regulated business is unintentonally committing even with a 90% false positive rate?
- Lionga 4mo agoHave you added "Make no mistakes" to the proompt? Mythos can't go wrong then, must be a skill issue.
- cheschire 4mo agoits shocking people don't realize you're being ironic
- SpicyLemonZest 4mo agoI realize they’re being ironic, it’s just a poor contribution to an otherwise productive conversation.
- steveBK123 4mo agoAI cannot fail, it can only be failed
- iugtmkbdfil834 4mo agoMy current favorite in that area ( because I saw it in the wild ) is: "Make it better" with no additional or reasonable previous explanation of what better might mean. "AI will figure it out" not for pattern extraction, but for a full blown analysis with equally generic prompt all confidently stated by an executive telling people working it how it works
- steveBK123 4mo agoIf you talk to it like a programmer talks to a computer, it works a lot better. So the question remains if non-programmers will adapt, the LLMs will accept wider range of input styles, or .. its just another abstraction layer for devs to use. I've observed this in the wild where someone is iterating with an LLM and giving it only negative feedback. For example responding to edits with "don't make it blue" rather than "keep the existing button shape, and change the color back to green". The LLM doesn't really come back the way a human would and say "so what color do you want?".. it just, guesses. Now abstract that to more complex tasks.
- gaiagraphia 4mo agoIsn't that a net positive though? (not sure about the cost human and tech cost). I'm guessing that without using Mythos, those conversations would never have been had, and confidence in the compliance of the product would've been lower. I love using AI tools as casinos. It's epic in helping to forge ideas and kickstart thought processes. You basically have the entirety of world knowledge at your fingertips to have a pint with.
- cucumber3732842 4mo ago> I'm guessing that without using Mythos, those conversations would never have been had, and confidence in the compliance of the product would've been lower. The conversations had already been had and the product made compliant. Mythos just pulled new rules out of its ass and of course the product wasn't compliant with those. So they do a fire drill and find that to be the case at great expense. Yeah you can frame it as "more checking is always better" if you wanted but that's just the same old "other people's resources are valueless" slight of hand we see on everything. It probably was mostly wasteful work.
- hedora 4mo agoThere's a chapter in Simple Sabotage about how to undermine a white collar organization from the inside. One of the key tactics is to hold meetings that revisit decided upon points, and to invent unnecessary process / checking. So, in this case, the LLM's behavior was equivalent to the behavior of the resistance during WWII. I think that book should be required reading for all engineering students.
- vulcan01 4mo agoyour parent: > the code in question had already been reviewed by human counsel
- johnbarron 4mo agoThey cant read all comments they comment on...
- rvz 4mo ago100%. Unfortunately those not in the depths of mission critical systems or regulated products will continue to believe that producing tons of code quickly using LLMs without humans in these systems is acceptable. Here's an example of what we will continue to see with folks fully immersed in gen AI psychosis: "The creator of claude code said that he no longer writes code for about 6 months and now has Claude doing all his work now. He also said recently that he no longer prompts Claude and now has it running in loops and it is self-improving itself and performing better than a human!" If the code produced by the LLM is perfect, the LLM takes the credit. But when a disaster happens, you cannot blame the LLM and it then falls on the human who did it. I don't think SWEs heavily vibe-coding with LLMs realize the risk in not understanding what the code the LLM being produced is doing even after generating tests (lol). We will see more of this too. [0] [0] https://sketch.dev/blog/our-first-outage-from-llm-written-code https://sketch.dev/blog/our-first-outage-from-llm-written-co...
- oceanplexian 4mo agoWhy is it such a dramatic statement for Boris to claim that he no longer writes code? Are people on HN still typing out functions by hand one character at a time? It would be like a developer in 2020 claiming that he only writes assembly because compilers can’t be trusted. No one is taking that person seriously. If you chose a career in tech you made a decision to work in one of the fastest moving fields in human history. Now it’s time to get over it, learn the new tools and adapt.
- bigstrat2003 4mo ago> Now it’s time to get over it, learn the new tools and adapt. No, thank you. I have used the new tools, determined that they aren't helpful to me, and set them aside as I would with any other bad tool. I don't feel the need to let hype take the steering wheel.
- chipsrafferty 4mo agoLet me guess, you used them 12 months ago?
- franze 4mo agowhat am i missing? you take a spec and create tests, every little thing you use another ai to verify these tests against the spec you review the tests vs the spec (at one point human review) you put the tests off limits to change / wall them. you let the ai write the software that fulfills the tests. there will be some gaps where you repeat the cycle above if the tests fulfill the spec, the code will fulfill the spec
- steveBK123 4mo agoIf each step requires micro-steps iterating with an LLM with human review to prevent hallucinations creeping in.. at some point you might just be better off letting the human do the work. Particularly as tokenmaxxing has ended and people are being charged more economic prices. If the pricing 5-10x the way Uber,etc did on the path to profitability.. even more so.
- Daishiman 4mo ago> If each step requires micro-steps iterating with an LLM with human review to prevent hallucinations creeping in.. at some point you might just be better off letting the human do the work. I mean for a lot of spec code people define the API signatures and let the llm run with it which is an excellent tradeoff.
- officialchicken 4mo agoIME, regulatory compliance is something you are rarely able to test for in a nice little box or with well-known suite. So there's no easy "this complies" in many situations, no matter how many lawyers, compliance officers, and llm's you run it past.
- franze 4mo agoso, whats the difference to human engineering? other than there are "internal micro feedback loops" during development?
- deleted 4mo ago[deleted]
- trumpdong 4mo agoIt was my impression that a whole lot of products are only pretending to be compliant, and that it's much more profitable to operate like that.
- rpicard 4mo agoIn my experience this is not representative of most fintechs. Of course there are both cases of real intentional noncompliance, and accidental, but by and large it seems like everyone’s trying to innovate within the law.
- scott_w 4mo agoThis makes sense because these companies want to become large companies and contract with large companies. Large companies, by and large, try to follow the law (while trying to bend it to the limit) because they're aware they have a big target on their back and no CEO wants to be on the front page of the papers for tanking a company in such a stupid fashion.
- sandworm101 4mo agoCompanies that are growing tend towards faking compliance. Many financial rules like pci only kick in at certain scales. So a company growing very quickly will often be behind the curve but will do everything to seem like they are compliant. Then they would hire people like me to come in and make them actually compliant. More often than not, making an effort at improvement was enough to keep the ball rolling.
- mattmanser 4mo agoI think it's the same throughout startup software to be honest. It's just easier to point out when there's clear rules. Security, GDPR, backups, build pipelines, disaster recovery, most of it will be faked, half-heartedly done once or ignored entirely. Then there's the more abstract things like scalability, idempotency when integrating with external APIs, error recovery, accessibility, UX, etc. Almost always that sort of stuff will have been entirely ignored, or there will be a fig leaf over a real mess of misunderstood standards or manual intervention steps. Startup developers usually have to be generalists as they often wear many hats, so things that need deeper domain knowledge get done to a bare minimum.
- mbbutler 4mo agoFalse-positive rate is so high with Mythos according to friends and other reporting I have seen. The original Mythos release used ASan to filter false-positives so it was able to maintain a good FPR, but when Mythos moves into domains that don't have a readily available oracle to help filter hits, the result is a deluge of false bullshit.
- ericmcer 4mo agoThe dynamic of agent codes human reviews does seem like the only sane one for the foreseeable future. Even Anthropic themselves still fall back to this. The problem is that sucks, even if all software engineers keep their jobs and salaries, the floor is still pulled out from under us. Imagine if a surgeons job was to supervise robot surgeons from a remote computer, or a woodworker just signs off on work before the machines do all the cutting and assembly. Sure they still have important jobs in their field but the soul & humanity of their skill is gone.
- adrianN 4mo agoAfter a couple of years of this their expertise will be gone too and then nobody is qualified to supervise the clankers.
- odeono 4mo ago"Soul and humanity" is doing a lot of work here. Does the woodworker who shape using a handsaw use less "soul" than the one who uses a machine? Does the musician who use a DAW and VSTs instead of analogue tape recorders create music with less "soul"? Does the painter who buys acryllic paint instead of synthesizing their own dye from plants use less "soul"? As technological innovation progresses, the barrier to creation falls. The process of creating something is not to be conflated with the final piece of art itself.
- jadbox 4mo agoNot _my_ opinion, but I just wanted to share that many people (in the Midwest) do believe that anything synthetic that it not readily made from simple materials has "less soul". It's a sorta test of "if I dropped you off in the jungle, can you still produce works of soul? Or are you just another cog in the machine.".
- ImprobableTruth 4mo agoExcept it's not just a tool. It's when a woodworker, musician or painter completely outsources their work and just marks what's wrong, sending those parts back. Yes, the final art piece might be the same, but the artist definitely uses less of their "soul".
- bobkb 4mo agoIMHO even if we are using auditing tools I believe we must use deterministic tools for critical analysis like this. Such rule and pattern based systems may not scale beyond certain point but they can be accurate.
- solenoid0937 4mo agoI use Opus 4.8 and GPT 5.5 and haven't suffered from hallucinations in months. But we also put a lot of effort into our harness.
- Loic 4mo agoSometimes the harness can only be a human. And this is fine. Developing new software with a really smart intern is the same, you, as an expert, need to bring your experience/expertise on the table to have everything right. Because experience needs time.
- Aeolun 4mo agoOpus 4.8 and gpt constantly hallucinate stuff as well. If you haven’t encountered or caught it that’s something different. Of course these days it’s mostly confidently asserting a wrong thing.
- deleted 4mo ago[deleted]
- rfgplk 4mo agoI primarily use Opus for a Lisp-like DSL codebase (non public, closed source) and it genuinely has _never_ hallucinated. All it pulls from is BNF, language spec + examples. So I have no idea how people are getting it to hallucinate on _popular_ languages.
- Chu4eeno 4mo agoI think you forget that they really are stochastic (I'd wager it has hallucinated things to you that wasn't important so you missed it, or you've just been very lucky), and the people you're arguing with are forgetting that there is significant difference in when and how often even frontier LLMs hallucinate. I catch claude every now and then, but isn't the measured hallucination rates down in low single digit percent for claude now?
- latentsea 4mo ago
- tpoacher 4mo agoIn some sense, you should still act on this, since if an external auditor relies on the same stack, it'll still cause you headaches.
- whatevaa 4mo agoThe models can change at any time and behave differently.
- PeterStuer 4mo agoI have worked on highly regulated areas in finance (risk). Compliance is a highly creative art, often requiring lots of out-of-the-box thinking and non-obvious solutions. The people I found worst at this were IT. They tend to over-interpret regulation, and super-restrict beyond what is needed for actual de-facto compliance. My guess is the model makes the same mistakes as the programmers: taking 'rules' literally, unaware of sectoral joint understanding, validated interpretations and habits. (btw. this is often on the non-tech side also a difference between regulatory and legal. The former are much more result oriented while the latter are primarily risk averse.
- jayd16 4mo agoWho gets in trouble if it turns out you are actually held to the literal rule?
- tsunamifury 4mo agoIf you think rules are literal than you aren’t aware how the world works. There’s a reason it’s called “judgement”
- jayd16 4mo ago...And that judgement could take them literally. So what is your point? My point was simply that it's easy to scoff at someone else being careful if it's their neck and not yours.
- parineum 4mo agoThey could but they don't. That's pretty much the whole job. You can also appeal decisions to a more reasonable party if you draw RobotJudge3000 for your trial
- deleted 4mo ago[deleted]
- rectang 4mo ago
- ilaksh 4mo ago3 years max. Maybe 5 if you are lucky.The models will continue to improve. The exponential gains in compute efficiency that have been ongoing for 70+ years will continue and that will result in even smarter models. There are dramatic hardware changes in the pipeline. But really that particular issue could have been solved by literally just telling it in a markdown file or instructions something like "verify all facts or compliance requirements with web search and include citations in responses".
- suttontom 4mo agoAh yes, the magical equivalent of "you are a senior software engineer who writes bug-free code". IME people would benefit greatly from the process, albeit tedious and time-consuming, of testing out the same prompt sequence/session with the exact same model multiple times. It becomes clear extremely quickly how capable but unreliable and inconsistent a model can be even when given the same context. If you have ever completed a long, complicated task with an agent and then lost the session and tried doing the same thing again from scratch you may have had the experience of seeing the subtle changes that come up in the model's thinking which lead it to accept or reject certain paths and ignore or incorporate prompt instructions like the one you've provided.
- ilaksh 4mo agoChange the temperature to 0 and it will be more consistent.
- deleted 4mo ago[deleted]
- ofjcihen 4mo agoThis is akin to “don’t make mistakes” “Verify all facts and compliance requirements” leaves enormous holes even if you assume the LLM has a concept of facts and requirements (it does not). What facts? What requirements? For what industry? For what subset of that industry? For what country or countries that you will be doing business in? Are these current “facts” and “requirements” or is the LLM referencing a dusty article from 1992 for which the subject matter has been radically overhauled? In my job I regularly see small but incredibly important mistakes like this lead to major issues. Some of those are human driven but increasingly the defense of the person responsible has turned into “Claude said it was fine though!”
- DaedalusII 4mo agoyou should rewrite this comment with chatgpt if you do not want to be dox https://www.tomsguide.com/ai/ai-can-now-identify-anonymous-internet-users-researchers-say-its-surprisingly-accurate https://www.tomsguide.com/ai/ai-can-now-identify-anonymous-i...
- DANmode 4mo ago> the code in question had already been reviewed by human counsel You sure?
- PIY 4mo ago[flagged]
- tenthirtyam 4mo ago> anyone who's using them to actually ship product without a human actually digging into it is opening themselves up to a world of risk. Maybe it's just me, but it seems that companies will happily take existential risks to get a better bottom line short term. Either you're too big to fail or you've already privatised any profits and subsequent losses (due to the risks becoming manifest) are socialised. The motor industry seems to be particularly egregious in this aspect, but also the food industry, construction industry etc. Seems to me even governments make the same choices in many ways - cut back health-care, policing, education, public transport and let the next government deal with the consequences.
- philipallstar 4mo ago> Seems to me even governments make the same choices in many ways - cut back health-care, policing, education, public transport and let the next government deal with the consequences. Definitely - defund the police was an astonishing rallying cry that made the communities it pretended to help much more dangerous for their residents. But the opposite is far more likely, and far more destructive long-term: it's easier to buy votes by spending more on social programmes and rack up debt and/or inflation to cover it, and then spend even more to fix that problem for enough voters that the people paying for it all can't vote it away, and the people who vote for it over the decades just don't understand why their pot feels so uncomfortably warm all of a sudden.
- mkovach 4mo ago[dead]