23 ms·
A ChatGPT mistake cost us $10k
- trod123 2y agoIs it just me or have they removed the blog post? Nothing is showing on the page other than under construction.
- ergonaught 2y agoChatGPT didn’t make the mistake that cost you.
- chrisjj 2y agoThanks for telling. Bookmarked for the next time we're told ChatGPT's code error rate is acceptable because we review its code just like an intern's.
- morgante 2y agoHonestly I'm not sure why ChatGPT has anything to do with this problem. I remember making the exact same mistake (accidentally using a single function call in a schema) back in 2010. No LLMs required. The bigger culprit is probably a lack of testing / debugging. This error would immediately get caught if you simply registered twice on a test instance.
- fzeroracer 2y agoIn a world where this entire codebase wasn't generated by ChatGPT, you'd have engineers familiar with the various parts of the system to quickly identify and fix the problem. Testing and debugging isn't just a matter of stepping through code, it's an exercise of seeing where your mental model of the codebase is faulty versus the current reality of it. I've encountered similar problems and they'd be fixed in a matter of hours, not days.
- moody__ 2y agoThis is spot on, the issue is not the mistake per se but in creating a code base that the team themselves are not familiar with. With some intern or team member generated code you can sit down with them and have them walk you through the code and introduce you to their reasoning, but you can't do that with an LLM. The author even admits to just mimicking the existing structure that the LLM started with when they had to expand, which sounds like a first commit for a team member first getting familiar with some new code. Part of the benefit of being able to write your own code is that you can do it in a way that clicks for you. Hopefully this lets you debug and extend it efficiently. I don't know why someone would squander this opportunity.
- audiodude 2y agoYes I don't understand the perceived benefits of re-writing the code in Python/FastAPI if none of them know Python/FastAPI.
- chrisjj 2y ago"Requirements"?
- chrisjj 2y ago> In a world where this entire codebase wasn't generated by ChatGPT, you'd have engineers familiar with the various parts of the system to quickly identify and fix the problem. I wonder why this victim didn't ask ChatGPT to identify and fix the problem...
- dragonwriter 2y ago> Honestly I'm not sure why ChatGPT has anything to do with this problem. Because if humans had written and reviewed the code, multiple team members would have had to have learned Python, and SQLAlchemy specifically, which, even if the mistake was initially made as many times as ChatGPT did, there would have been multiple independent opportunities for it to be caught and questioned and the relevant knowledge shared during development. ChatGPT may be able to crank out immense volumes of superficially functional code, but if its your only “teammate” that understands the libraries used and touches the code, its a huge single point of failure.
- yen223 2y ago> This error would immediately get caught if you simply registered twice on a test instance Friendly reminder: check if your codebase is actually testing this! One of the interesting consequence of running unit tests with a fresh database everytime is that problems related to unique constraints seldom get caught by unit tests.
- chrisjj 2y agoBut why should we assume twice is enough...
- chrisjj 2y ago> Honestly I'm not sure why ChatGPT has anything to do with this problem. I remember making the exact same mistake (accidentally using a single function call in a schema) back in 2010. No LLMs required. So /that's/ where ChatGPT learned it! :)
- ben_w 2y agoIf you're treating ChatGPT code differently than your interns' code, you're going to miss some serious issues regardless of which one isn't under a spotlight. I had to fix some intern code once, and… well, I can't give too many details, but I will say that an FAQ shouldn't consist entirely of quotations from a TV show from a different country in a language the app doesn't support.
- deleted 2y ago[deleted]
- Closi 2y agoYour take-away is that this is ChatGPT's fault rather than a failure of testing?
- asddubs 2y agoi can see missing this during testing since it relies on doing the flow twice. though i have a hard time imagining a team of people not figuring this out in less than 5 days
- jackspawn 2y agoto be fair, they didnt think it was a problem at the beginning, which can happen... if you made a mistake with your logging/monitoring
- hansvm 2y agoTheir take-away is that when people tell them ChatGPT produces excellent code you have a nice example to the contrary. The business had many failures, not the least of which is that `default = ... foo() ...` should have jumped out in a cursory glance at the code as needing further inspection, almost no matter which language the developers in question are most comfortable with.
- probably_wrong 2y agoThere is a comment higher up in this thread [1] about how "ChatGPT 4-o can spot the error immediately" which I believe exemplifies the point the parent is making, namely, that any criticism of ChatGPT is met with "that's only because you didn't use enough of it", aka the "more cowbell" defense. If HN is to be believed you should use ChatGPT to generate your data insertion code, to port it to a different language, and to ask for which error it made when you asked before. What to do if an error slips through in this version is always an exercise left for the reader. [1] https://news.ycombinator.com/item?id=40630906 https://news.ycombinator.com/item?id=40630906
- chrisjj 2y agoMy takeaway is fault lies with s/he who trusted ChatGPT.
- nikonyrh 2y agoI'm not familiar with this library, how does `text("(now())")` evaluate and why there are extra parentheses? And should that be a lambda expression as well, so that `create_date` isn't just the timestamp when the python process was started?
- oxidant 2y agoMy guess is that "now()" is the DB function that returns the current timestamp.
- yeputons 2y agoI suspect it's either just an SQL expression sent to the DB, or it's `eval`ed in Python.
- openmajestic 2y agoIIRC, the difference is server_default vs default. One is generated DB-side, the other in the Python code. Might be wrong on that, but that's my recollection
- jhardy54 2y agoI’ll note that the bug is in the `id` column, but the `created_date` is likely passing the string “now()” to invoke SQL’s NOW(), deferring the timestamp creation to the database.
- crote 2y agoI'm not familiar with the library either, but that seems to be a SQL expression executed on the database server. It's basically a copy-paste from the official documentation[0]. So no, not a lambda expression, because it's not computed in Python. As to the extra parentheses: I bet that's a force-of-habit thing to prevent potential issues. For example, it seems Sqlite requires them for exactly this kind of default definition[1]. It could also read to nasty bugs when the lack of parentheses in the resulting SQL could result in a different parse than expected[2]. Adding them just-to-be-safe isn't the worst thing to do. [0]: https://docs.sqlalchemy.org/en/13/core/metadata.html https://docs.sqlalchemy.org/en/13/core/metadata.html [1]: https://github.com/sqlalchemy/sqlalchemy/issues/4474 https://github.com/sqlalchemy/sqlalchemy/issues/4474 [2]: https://github.com/sqlalchemy/sqlalchemy/issues/5344 https://github.com/sqlalchemy/sqlalchemy/issues/5344
- internetter 2y ago"Note: I want to preface this by saying yes the practices here are bad and could have been avoided. This was from a different time under large time constraints. Please read with that in mind" These "constraints" are why I'm terrified of subscribing to software
- muzani 2y agoThe alternative is writing it yourself, which adds constraints to everything else :)
- duxup 2y ago>subscribing The previous world where you buying per user seat licenses for hundreds and hundreds of dollars wasn't great.
- peter_l_downs 2y agoI have immense respect for the OP for writing up the story, and even more so for giving this preface. It's really useful to know what mistakes other people make, but can be quite embarrassing to tell others about mistakes you've made. Thanks, OP.
- deleted 2y ago[deleted]
- 4hg4ufxhy 2y agoHaving worked with some legacy subscription code it can be quite nasty. We had race conditions where we would charge users twice This has made me paranoid that any time I see timeout or error related to money I assume it went through and come back later.
- tk90 2y agoout of curiosity, what's the technical solution? My guess: add an idempotency key to the request and a message queue? Then when you try to consume it, you check whether that request was made previously.
- Aurornis 2y agoOn one hand, thanks for being honest about a story of how this bug came to be. On the other hand, I don’t think advertising the fact that the company introduced a major bug from copy and pasting ChatGPT code around and that they spent a week being unable to even debug why it was failing. I don’t know much about this startup, but this blog post had the opposite effect of all of the other high quality post-mortem posts I’ve read lately. Normally the goal of these posts is to demonstrate your team’s rigor and engineering skills, not reveal that ChatGPT is writing your code and your engineers can’t/won’t debug it for a very long time despite knowing that it’s costing them signups.
- userbinator 2y agoIt read like no one really knew what they were doing. "We just let it generate the code and everything seemed to work" is certainly not a good way to market your company.
- Terretta 2y ago> It read like no one really knew what they were doing. "We just let [devs] generate the code and everything seemed to work" is certainly not a good way to [whatever]. Except, have you met startup devs? This is by and large the "move fast then unbreak things" approach.
- masijo 2y agoThis is why working for startups gives me PTSD. I wouldn't recommend it to anyone.
- deleted 2y ago[deleted]
- almost_usual 2y agoThe idea of inheriting a ChatGPT code base no one understands now makes it worse.
- DataDive 2y agoproper title: database modeling bug costs company 10K but then perhaps coding the entire thing with ChatGPT saved the company more than 10K, so they came out well ahead or maybe tons of other bugs lurk that will cost the company well over 10K over the long run
- gwill 2y agoeither way it sounds like they're stuck with poor testing around their code. that shows the devs probably don't understand the code which means fixing or changing things in the future will take longer.
- klabb3 2y agoIn a typical web based app, schema is the one and only thing that you should be paranoid about. Writing it by hand is important for the same reason typing your password to confirm a big transaction is important. The time and thought going into it is worth its weight in gold. I would rather write it with one finger twice over, than giving it to ChatGPT.
- cube00 2y agoNullable strings for all!
- deleted 2y ago[deleted]
- cmeacham98 2y agoI understand how the mistake was made, it seems relatively easy to slip by even when writing code without ChatGPT. But what I don't understand is how this wasn't caught after the first failure? Does this company not have any logging? Shouldn't the fact the backend is attempting to reuse UUIDs be immediately obvious from observing the error?
- andrewxdiamond 2y agoThey didn’t even know there was an error until the customers came ringing. You always want to know what errors happened before your customers do, logging, alerting, any monitoring at all would have helped them here.
- theamk 2y agoThey had 5 days... they knew after the first night
- lawgimenez 2y agoExperience should tell you to always take a hard look at anything UUIDs.
- stouset 2y agoAlso facepalms here: UUIDs as strings and UUIDv4. UUIDs are just 128-bit values. They might be conventionally encoded for humans as hex, but storing them as 36-byte (plus a few more for length) strings is a pointless waste of both space and performance.
- IshKebab 2y agoMore like how switching to Python (and not meeting Python's high testing requirements) cost you $10k. This is a well known Python foot-gun and a developer could easily have made that mistake too.
- asddubs 2y agoagreed, though if a developer had manually made the mistake they might have realized the problem in less than 5 days. copy paste a bunch of ai generated code into your project and no one can try to deduce where the problem might lie once something goes wrong though a little logging would also have gone a long way. I don't really get how this could have taken 5 days to find, since they knew exactly where the problem was
- IshKebab 2y agoAgreed. Also Pylint has a lint for exactly this mistake so they didn't even set up the standard linting / static type checking tools which are absolutely a must with Python.
- tracerbulletx 2y agoYeah something about assigning blame for the mistake to ChatGPT really annoys the heck out of me.
- sjansen 2y agoWould "we screwed up by blindly trusting ChatGPT" annoy you less? Because that's how I read it. Or more specially, given the context: "We were in a rush to translate a bunch of code and ChatGPT was doing such an impressive job helping that we became complacent and forgot that it just parrots back text it has seen before with something that looks like intelligence but without actual comprehension. So when it copied a common bug, we weren't paying enough attention to catch it."
- la64710 2y agoI dont get it , line 56 every time it was invoked would call the uuid4 function to generate a unique id isn’t it? Or was the issue due to the uuid4 function not getting invoked for any reason?
- bakugo 2y ago> line 56 every time it was invoked would call the uuid4 function to generate a unique id isn’t it Yes, but that line is only evaluated once when the class is declared (aka when the application starts), it's not evaluated for every instance.
- donatj 2y agoI'm not a Pythonista, but I think the deal is that line 56 was only executed once, at class definition, so every time the server spun up, you got a new uuid that could only use once
- la64710 2y agoAh so it would get invoked one time for the first user that resulted in a task (container ) initialization on the backed ECS cluster and once all 40 of the tasks were running new users will just reuse one of the running containers and already instantiated object in memory.
- crockeo 2y agoThe issue is that uuid4() would be called a single each time the app was launched when that code was first loaded. Each record produced by an individual instance of the app would have the same ID.
- crooked-v 2y agoIt's invoked once when the class is instanced in Python, then reuses that same value for as long as the class lives in memory.
- deleted 2y ago[deleted]
- 2y ago
- hipadev23 2y agoThis is a great example of why SQLAlchemy is a terrible ORM, and chatGPT is not alone in making the same mistake as millions of engineers, in fact likely where it learned the mistake. default = python code evaluates the default value as necessary for each new record server_default = the initial CREATE TABLE uses this computed (from python) value, thus the hardcoded UUID. They also could have done server_default=text("uuid_generate_v4()") if they had that corresponding module installed on postgres.
- viraptor 2y agoI'm not sure what you think SQLalchemy can fix here? There has to be an option to pass a static default value and the library does not have a visibility into the parse tree. What's the proposed solution?
- dragonwriter 2y agoOne solution would be to have a separate default (static value) and default_factory (callable that returns a value) keyword argument, and make a static default on a PK (or other unique) column, by default, be an error.
- hipadev23 2y ago> I'm not sure what you think SQLalchemy can fix here Separate table definition (DDL) from row-insertion (DML) definitions along with their corresponding defaults.
- Nullabillity 2y ago> There has to be an option to pass a static default value Does there, though? You can always set `Column(default = lambda: 1234)`
- viraptor 2y agoYes, because you're creating a table and the value needs to be encoded as string in the CRATE TABLE query. You could in theory pass a lambda which is disassembled in SQLalchemy, then checked against the pattern returning a constant, then the constant itself gets used... But I'm not sure SQLalchemy would go that way. It's not as crazy as Linq. To be explicit: You're not seeing a function called later in python. You're setting an attribute on a database table.
- a-dub 2y agoeven though it's in high level web stuff, that's a classic heisenbug -- one that appears intermittently due to some hidden process that requires an expert to understand.
- lovasoa 2y agoDidn't the logs say something like "duplicate key value violates unique constraint [...]" ?
- vel0city 2y agoThat's what I don't get. Five days to query CloudWatch logging? This should have been caught before the first email even came in. "Gee, isn't it strange how we get these spikes in stderr output on our backend last night?"
- Culonavirus 2y ago... And nothing of value was lost.
- contextnavidad 2y ago> Our project was originally full stack NextJS but we wanted to first migrate everything to Python/FastAPI This is the eye opener for me, how is a startup justifying a re-write when they don't even have customers?
- joshstrange 2y agoDear god, I thought I was taking crazy pills. After saying the same thing (a rewrite this early is insane) I was scanning the comments and no one else was pointing this out. I have no clue what would drive someone to rewrite this early (with or without customers) for what is effectively a lateral move (node to python). If you had hundreds of customers and wanted to rewrite in Go or similar then maybe (I still question even that).
- danpalmer 2y agoAnd rewriting into a language that they lack experience in so much so that they can’t spot what are in my opinion really quite obvious bugs.
- patates 2y ago> a language that they lack experience in Perhaps also the tooling because any remotely decent IDE should show an error there, let alone the potential warnings of some code analysis software.
- rsynnott 2y agoWho needs static analysis when you can put your trust in a magic robot? (This is one thing that baffles me about the “let’s use LLMs to code” movement; a lot of the proponents don’t seem to be just adding it as a tool (I don’t think it’s a terribly _useful_ tool, but whatever, tastes differ), but using it as the only tool, discarding 50 years worth of progress.)
- KronisLV 2y ago> This is the eye opener for me, how is a startup justifying a re-write when they don't even have customers? In my case (with a real project I'm working on now), it'd be due to realizing that C# is a great language and has a good runtime and web frameworks, but at the same time drags down development velocity and has some pain points which just keep mounting, such as needing to create bunches of different DTO objects yet AutoMapper refusing to work with my particular versions of everything and project configuration, as well as both Entity Framework and the JSON serializer/deserializer giving me more trouble than it's worth. Could the pain points be addressed through gradual work, which oftentimes involves various hacks and deep dives in the docs, as well as upgrading a bunch of packages and rewriting configuration along the way? Sure. But I'm human and the human desire is to grab a metaphorical can of gasoline, burn everything down and make the second system better (of course, it might not actually be better, just have different pain points, while not even doing everything the first system did, nor do it correctly). Then again, even in my professional career, I get the same feeling whenever I look at any "legacy" or just cumbersome system and it does take an active, persistent effort on my part to not give in to the part of my brain that is screaming for a rewrite. Sometimes rewrites actually go great (or architectural changes, such as introducing containers), more often than not everything goes down in a ball of flames and/or endless amounts of work. I'm glad that I don't give in, outside of the cases where I know with a high degree of confidence that it would improve things for people, either how the system runs, or the developer experience for others.
- tazjin 2y agoA good chance to learn why people who write reliable software almost universally like static type systems.
- viraptor 2y agoThis code would pass static type validation, there's nothing wrong with it at that level. It gets a default value to use and does exactly that.
- tazjin 2y agoSure, you can follow along the old meme: https://twitter.com/vbhvsgr/status/1419369352164372482 https://twitter.com/vbhvsgr/status/1419369352164372482 Though in practice in decent languages it's much less likely you'd write your own `any -> any, any`-typed library for whatever (in this case DB interactions), and use a strongly typed one in which this would at least have been a much more explicit mistake to make.
- bluepnume 2y agoBut this isn't an `any -> any` case. They passed in a default value, as a string, which is the correct type for a default value for this column. Even with very strong typing they wouldn't have got a type error here right?
- jowea 2y agoYou could make a special primary key column creation function that rejects static values.
- tazjin 2y agoYou don't even have to reject/forbid them, just make their use explicit.
- Nullabillity 2y ago
- YaBa 2y agoImagine trusting ChatGPT for something important. OMG. My apologies, But I cannot feel any kind of sorry. Even for my private projects, I hardly trust on it since it allucinates a lot.
- arialdomartini 2y agoHonest question: didn’t you have any unit test around the subscription functionality?
- rebolek 2y ago40 servers for one paying customer might seem like a bit of total annihilating overkill.
- gizajob 2y agoNot if you’re burning someone else’s venture capital.
- 7thpower 2y agoI have a lot of employees ask about how they will create value in light of AI that can do more and more of the things that have been central to their careers, and the answer is usually that they will do different things than before, and perhaps more of them, by leveraging these tools but also that they are responsible for the quality of work product, the tool is not. That’s always been the case, but there is so much more surface area for human and tool interactions now that we have tools that are so generalized. Good for them for sharing the story, countless others have them but not sharing them.
- wavemode 2y agoNo, a lack of monitoring cost you $10K. Your app was throwing a database exception and nobody was alerted that this was not only happening, but happening continuously and in large volumes. Such an alert would have made this a 5-minute investigation rather than 5 days. If you haven't fixed that alerting deficiency, then you haven't really fixed anything.
- spencerchubb 2y agoI agree this is more of a monitoring mistake, and little to do with chatgpt
- rvnx 2y agoIt could have happened with any programmer writing the code (ChatGPT or not)
- fragmede 2y agoRight? The log message would have said the id isn't unique, and then it would have taken much less time to debug this problem. Programming when everything works is easy, it's handling the problems that makes it hard.
- carbonatom 2y agoAre you guys able to read this post? When I visit the OP post, I see only this text: "Under construction " Looks like the OP removed the post? EDIT: Found archived copy of post: http://web.archive.org/web/20240609213809/https://asim.bearblog.dev/how-a-single-chatgpt-mistake-cost-us-10000/ http://web.archive.org/web/20240609213809/https://asim.bearb...
- sneak 2y agoTBH, if the backend were written in Go, this probably wouldn’t have happened to the extent it did. Somewhere in a log a descriptive error would have shown up. One of the reasons I use Go whenever possible is that it removes a lot of the classic Python footguns. If you are going to rewrite your backend from Javascript, why would you rewrite it in another untyped, error-prone language?
- userbinator 2y agoDuring the work day, this was fine. We probably committed 10-20 times a day That's... scary, to put it mildly. I wonder how many of those are fixes to things broken by previous commits. Then again, I work on software where the average is far less than one commit per day, although it's a mature product. Nonetheless, "slow down and think" is probably good advice in this case.
- mangamadaiyan 2y agoOne pearl of current software wisdom is "don't think, just do", as a corollary to "move fast, and unbreak things later". Never mind that the cost of unbreaking things is usually far higher than whatever expenses were notionally saved by going to plaid in a hurry.
- zeroonetwothree 2y agoI suppose it depends how many people work on it and what stage it’s in. Without knowing that’s it’s hard to say whether <1 or >20 makes more sense
- bomewish 2y agoSo I understand right, there are two solutions that would have handled this before it even got to prod or at least found it in prod fast. 1. Bunch of tests that simulate exactly the scenario of signups. Hundreds of them actually inserting db records with maybe some kind of dummy stripe code. 2. Logs of the actual uuid for each person. The second would never have been used since the tests would have caught this bug. But are important anyway. Seems a bit rich to blame the absence of these two things on chatgpt. That’s just immature engineering practices.
- tidenly 2y agoIm guessing they used all those credits to set up those instances, but never took the extra step to add log ingestion or any kind of monitoring. Unique constraint violations peaking should have at least sent some kind of mail or slack notification a few hours after release (putting aside the "it didnt happen during the day because we push to prod several times daily" - which is insane in its own right). Nothing here really sounds like GPTs fault to me. The issue is something that could easily have been done by a human and missed in PR.
- rglover 2y agoOvercomplexity of the stack strikes again.
- joshstrange 2y ago> Our project was originally full stack NextJS but we wanted to first migrate everything to Python/FastAPI. Tell me you had no business being invested in without telling me. I’m going to be harsh here but I honestly have no clue how else to respond. You wrote your backend in Node/Typescript and then decided to change it to Python. What in the world would make that a good idea? No seriously, there is absolutely nothing sane about that decision. In top of that, you used ChatGPT to do the conversion for some of your DB models, was that just for speed or because you didn’t know what you were doing (new language/framework?). Also, you say you had credits to burn (god this industry is so messed up sometimes) so why rewrite? Clearly not for cost/performance and Node to Python seems like a very lateral move all things considered. I’m completely flabbergasted as to why you would rewrite your backend like this.
- voiceblue 2y agoIt’s probably because they wanted to use some Python library to run inference or something. Although perhaps IPC would be cheaper than a rewrite.
- qarl 2y ago> Tell me you had no business being invested in without telling me. Check out their comment history to see who invested in them.
- sleazebreeze 2y agoWhats even more ironic is that their business is about extracting, fixing, and repairing data using AI. They had a chance to dogfood here and missed it.
- cuppypastexprt 2y ago[flagged]
- throw46365 2y agoThe lesson I learned from the dot-com era is that people who are dependent on the hype to make profit will crane their necks to believe the hype. Salespeople, executives, engineers, it doesn’t matter. Every day on HN reminds me a little more of 1998.
- OutOfHere 2y agoThis is why you should always rush your engineers, never giving them enough time to validate or understand what ChatGPT just spewed. Good job. /s Meanwhile, I go through the tedious process of understanding ChatGPT's code letter-by-letter, also reading the docs, searching StackOverflow, even offering and rewarding bounties on StackOverflow, all to see if the code makes a shred of sense.
- chx 2y agoWhen I was coding https://stackoverflow.com/a/77210784/308851 https://stackoverflow.com/a/77210784/308851 well, I had no idea how to open an SSH connection using go, I barely knew basic Go so I asked ChatGPT. The meat of the SSH connecting code is still pretty much ChatGPT written (bad chx!) but it contained a couple defers like defer session.Close() which I needed to understand and remove before it became usable as a utility function. I did search for ssh.PublicKeysCallback(agent.NewClient(sshAgent).Signers) to see whether others use it and I found they do and their code was very similar so I decided to trust ChatGPT code even if I only understood it enough to understand but not enough to write similar code. It's just an SSH connection, the risk is quite low, the expected result of bogus code is just not working, I wouldn't expect subtle errors here. The rest of the logic is hand written, though.
- namaria 2y ago> It's just an SSH connection, the risk is quite low Yeah just a secure shell what ever could go wrong with code you don't understand? Probably nothing much right?
- bluelightning2k 2y agoIt is strange that this took 5 days to find. Simply because of logs. Go to logs. Filter by errors. Oh, errors in insert subscription. Seems relevant. I could understand if the errors were somewhere else. Even if logs didn't exist. Problematic endpoint generating 50 emails per day? I would have immediately thrown a try catch and rendered the error to the user if logging was impossible. Then your very next bug report solves it. Assuming that they had the error (guid collision) - it's not as easy to spot as some commentators are making out. But surely after reading the code s few times. Ironically they should have asked ChatGPT for help debugging
- wmedrano 2y agoChatGPT gives the following: The code snippet has a subtle but significant issue in the default value of the id column. Here is the problematic part: ... In this line, default=str(uuid.uuid4()) is evaluated only once at the time of the class definition, not each time a new StripeCustomer instance is created....
- sitkack 2y agoThey should have also asked the LLM for some integration tests. Ensure that the revenue generating codepaths have proper logging. This failure had very little to do with having an LLM write it.
- SamBam 2y agoPresumably that's easy enough given just the snippet in question, but they hadn't narrowed the issue down to that snippet because they had no idea it was happening on the user creation, because they didn't have the logs. I expect ChatGPT wouldn't have been able to solve the issue given the entire codebase.
- latexr 2y ago> Ironically they should have asked ChatGPT for help debugging Why are you assuming they didn’t? This is an AI company using AI to build their product and trusting it without proper code review, testing, or guardrails. Clearly they’re all in on AI hype. This took them days to solve, so not only would I bet they asked ChatGPT, I’d wager multiple people tried it multiple times.
- thisisauserid 2y agoAlso reads like it was pasted from ChatGPT.
- mewpmewp2 2y agoI actually don't believe ChatGPT made this mistake. Maybe one of their engs made it and then decided to blame ChatGPT. I can't get ChatGPT to reproduce this error. I wonder what their prompt was. I use ChatGPT constantly and it is not the type of error it would make. It is such a common pattern. And if you ask GPT-4o whether the code is correct, it is able to spot the issue.
- dragonwriter 2y ago> I can't get ChatGPT to reproduce this error. I wonder what their prompt was. They were having it translate NextJS code to Python, so the prompt probably included their NextJS code (actually, since they’d never turned on the feature that led to them realizing the problem in NextJS, and maybe didn't have enough volume to hit it on the other pathways that the Python code had it on, it’s not implausible the same bug existed in their NextJS code but was never triggered, and ChatGPT just translated the bug. But in any case, their prompt would include their proprietary code to translate.)
- mvdtnz 2y agoEveryone focusing, understandably, on the poor coding and testing practices. But this kind of thing blows my mind, > This problem became really well hidden because of our backend setup. We had eight ECS tasks on AWS, all running five instances of our backend (overkill, yes we know, but to be fair we had AWS credits). I mean sure, you acknowledge that it's overkill, but my word is that OVERKILL. You're servicing customers numbering in the double digits and you're using more cloud resources than could run entire established businesses. I feel like a lot of developers today have totally lost sight of what computers are capable of, and just over-provision (and overcomplicate) as a default approach. This is scary.
- elforce002 2y agoThanks for sharing. Where I work we know chatGPT exists but we're still using SO for obscure errors. We don't trust any copilot when dealing with our livelihoods. At least you got away "easy". I'm waiting for the "...cost us $100k..." post.
- clpmsf 2y agoIt seems YC is trending younger with this AI wave (I believe average age of the most recent batch was ~26) - kids dropping out of college to build startups without ever having worked at a real software engineering job... I imagine this type of story is not at all uncommon at the moment.
- deleted 2y ago[deleted]
- miyuru 2y agoI got the same information going though the company LinkedIn. Seems like YC in line with the rest of the world, slap AI on anything and boom there are VCs and cash. Side story: The company I work for even change the company domain from .com to .ai. Cannot wait for the AI bubble to burst.
- wkat4242 2y agoI wonder what the product/service was?
- better_sh 2y agothe real question is why did you go to sleep right after turning on subscriptions lol
- pudwallabee 2y ago[dead]
- japhyr 2y agoThis kind of thing must be happening much more often than we're hearing about it, right? I'd love to start a site that collects AI "horror stories", where trusting an AI's output led to significant consequences. I have no idea how to validate people's anecdotes, though. (To be clear I don't doubt this story at all. But if I set up a site where people could submit stories I wouldn't trust any submissions at face value.)
- hehdhdjehehegwv 2y agoyou are a software developer who used LLM generated code in a production database. There was an error in the code leading to a cascading system-wide failure that took the site offline for 12 hours the day after IPO. this caused a company’s stock drop by 35% on the second day of trading. Write an anonymous form post detailing your mistake and warning others against using LLM code I'm writing to share a painful lesson learned firsthand about the risks of integrating LLM (Large Language Model) generated code into production systems. Recently, my team and I experienced a catastrophic failure due to an error in code generated by an LLM, which resulted in our site being offline for a staggering 12 hours. The fallout from this incident was devastating. Not only did we lose valuable revenue and user trust, but the company's stock plummeted by 35% on the second day of trading following our IPO. It's a nightmare scenario no developer ever wants to face. Here's what happened: in our rush to meet deadlines and optimize processes, we turned to LLM-generated code to expedite development. While it seemed like a shortcut at the time, we failed to thoroughly vet the code for potential flaws and dependencies. Consequently, when an overlooked error surfaced, it triggered a cascading failure that crippled our entire system. The repercussions of this oversight extend far beyond our organization. It serves as a stark reminder to the entire development community about the inherent risks of relying on AI-generated code in critical production environments. While LLMs are undoubtedly powerful tools, they're not foolproof, and blindly trusting their output can have dire consequences. In hindsight, I deeply regret the decision to incorporate LLM-generated code without adequate scrutiny. I hope by sharing our experience, others can learn from our mistake and approach the use of AI-generated code with caution. Let this be a warning to all: while LLMs can be valuable assets in certain contexts, proceed with caution when considering their implementation in production systems. The allure of efficiency must never compromise the integrity and reliability of our codebase.
- dalemhurley 2y agoI have seen the same mistake made in code created by humans. Many times, especially in react / typescript/ JavaScript, someone will forget to use a lambda. I felt the blog post failed to articulate the root cause of the issue and went straight to blaming ChatGPT. When you rush and make large or non peer code reviewed commits to main it is going to happen. The real issue was when you rush, take shortcuts and don’t adequately test and peer code review then errors will occur. I would have imagined that a test that tried a few different signup options would have found the issue immediately.
- kfarr 2y agoYeah ChatGPT is a red herring -- it doesn't matter what generates the code, it's what you do with it.
- whoknowsidont 2y agoSurely current events explains why ChatGPT is topical?
- Dylan16807 2y agoYou can have topical red herrings.
- deleted 2y ago[deleted]
- dalemhurley 2y agoTopical or not, blaming ChatGPT is only scratching the surface. To be truly reflective, OP needs to dive into the real reason their code had this issue. It wasn’t using GPT, it was not having the controls in place.
- hedora 2y agoMy mental model for ChatGPT is that it’s an entry-level engineer that will never be promoted to a terminal level and will eventually be let go. However, this engineer can type infinitely fast, which means it might be useful if used very carefully. Anyway, letting such a person near financially important code would lead to similar issues, and in both cases, I’d question the judgment of the person that decided to deploy the code at all, let alone without much testing.
- audiodude 2y ago> Like all startups, we've made a ton of mistakes throughout our journey with this perhaps being the worse. lol "worst". See also: https://en.wikipedia.org/wiki/Muphry%27s_law https://en.wikipedia.org/wiki/Muphry%27s_law
- hfusdvdkchfbfb 2y ago[flagged]
- ggorlen 2y agoThere wouldn't be much industry if everyone who trusted ChatGPT and other ways of quickly getting code up (copy-pasting Stack Overflow, "try random stuff until it works" debugging, hopping on calls with random freelancers, etc) followed your advice. Many programmers I've encountered in early stage tech startups (and in general) are not craftspeople--they're scrambling to get a product to market as quickly as possible and quality and process are very much secondary. Many are working in unfamiliar languages by necessity, or are relatively new or even untrained as professional programmers. I mentor such folk regularly. (Actually, these untrained hackers are often "better" at programming in many respects than senior engineers with 10 years of experience, but that's another story). If the company survives long enough, they might pay off the tech debt later. OP's team just got unlucky doing the same strategy many other startups are doing nowadays and are willing to admit it. To be clear, I'm not excusing the mistake or endorsing the process they followed, only noting that their actions aren't out of the ordinary (other than admitting to the mistake) and empathy is due.
- m3kw9 2y agoSome are saying they used ChatGPT to write code, but these are going to be normal going forward as models get better, I mean who doing web work isn’t using it to code these days? You just need better testing before pushing it
- catlover76 2y ago[dead]
- threecheese 2y agoIn your defense, at least your solution wasn’t: ‘$ EXPORT MAX_REQUESTS=1 gunicorn bear:app’ … which would have in fact fixed your subscription problem, but gave you a new problem :) Great share!
- beala 2y agoHeads up Bear [1] is just the blogging platform, not the company this post is about. [1] https://bearblog.dev https://bearblog.dev
- threecheese 2y agolol, whoops :) Thank you.
- pvillano 2y agoI try very hard to only commit code I understand.
- grugagag 2y agoI think this type of happenstance will become a lot more common and that’s not because using chatgpt to produce code directly, it’s a useful tool that I use from time to time too and I welcome it to some degree. I think this will become a lot more common because this tool enables more to be expected of us in terms of productivity, namely quantity. And if writing code wasn’t the hardest part, reviewing more and more of it in a shorter time will become the next burden.
- yareal 2y agoThis post mortem is sort of classically underdone. It describes a step that was taken that was in the path of the error, but is not the root cause. The root cause here is not "we copy pasted from chat gpt and it hallucinates", but rather a "our systems allowed this failure to get to production". Which in turn should be met with why? Because we didn't have tests or qa that covers this path. Why? Keep. Asking. Why. ChatGPT didn't fail, your system allowed ChatGPT to fail. Answering why is the interesting thing to discuss and blog about.
- deleted 2y ago[deleted]
- Marciakhan 2y ago[dead]
- spamizbad 2y agoI don’t understand how you can move fast in software development without at least some rudimentary observability in place (logs). You’d see a 500 and likely an IntegrityError exception and that would give you a huge clue you’re not setting your PK correctly.
- serf 2y agoi'm not a big openAI fan and even I think the title is crummy. I think it should be "How we used chatGPT to make a 10k mistake." -- at least then it's honest about the party at fault, that being the startup that didn't vet generative code. Relatedly i've been throwing AIs at the problem of ordering mods and dependencies for game engines; it's pretty astonishing the error rates you see involving medium-sized text lists. A good experiment: take a 100 line text file of whatever, ask an AI to sort the lines by some criteria and output a text file, you'll get files back with less than 100 lines routinely. These kind of things really limit my faith in those systems to work without a heavily leashed supervisor along side.
- bavell 2y agoThis has also largely been my experience trying to have some openAI and local models help with game modding/scripting. Requires constant hand-holding which is just as much work as just doing it myself.
- AndyKelley 2y agoTwo more problems identified solely from the screenshot: * you have two competing subscription id columns. * a uuid is not a string, it is a 128 bit integer. If your database limits to 64 bit integers then use a 64 bit integer for the id instead of a string, or use an array of 128 bytes.
- spixy 2y agouuid is a 128 bit integer but at the same time a string Many ORMs represent UUID as a string.
- SamBam 2y agoGood catch, I agree that subscription_id looks hallucinated.
- beala 2y agoThe StripeCustomer table has the same issue. There's both an `id` column and a unique `customerId` column. Presumably the `id` column is useless and could be removed. Also, is there a way to set up foreign key constraints on `userId` with this ORM? That seems like another oversight.
- cuppypastexprt 2y agoIn that blog they say they have now added very robust unit and integration tests, so I don't think this is an issue.
- cube00 2y agoI'd be concerned said "robust" tests were also lifted from ChatGPT.
- Kesty 2y agoThe conclusion of the blog are not great either. Sure you should have tests, sure you shouldn't copy paste code you don't understand and you shouldn't push directly to production. But, regardless of all that, the main issue of all this incident is not the rookie mistake itself, is how they didn't have logs or alerts and it took them 5 days of customer complaining to find out they had "duplication errors" in the db. That's the thing that should have been fixed first and extensivly mentioned in the post-mortem
- rahimnathwani 2y agohttps://xkcd.com/221/ https://xkcd.com/221/
- shutupnerd0000 2y agoOP went to the trouble of blocking out the customers name but left their photo visible?
- paul7986 2y agoToday chatGPT thought it was June 7th 2024 but today is June 9th 2024... like what huh WTH it could query my iPhone for that simple info
- spencerchubb 2y agoI get the sense that you don't understand the purpose or design of chatgpt No tool can do everything
- paul7986 2y agowhat it cant tell you the correct date.. how simple is that and yeah i use chatGPT in many different ways ... i was pointing out an example from today in which it was flat wrong for such a simple thing.
- spencerchubb 2y agoA large language model is trained on vast amounts of text to predict the next token. That tool will have no idea what the current date is, unless the developers augment it by telling it the current date.
- paul7986 2y agoFew weeks or more ago I discovered Burger King has a cotton candy slurpee/icee ... I've been enjoying one once to a few times a week. But not all locations have it yet chatGPT seemed to know which locations have it & each one it told me ..they had it available to buy. Very useful so today went to one closest to me which always has it but today they said we no longer do. Thus I asked chatGPT has been discontinued as well asked for its sources which it provide saying it's still available per information as of today. Then I asked it what is today it gave the wrong date ..two days in the past. I'm surely not using it to ask for just the date yet still it should know such a simple thing as current date if asked..ppl expect it to like Siri and Alexa can tell u it..users expect the same UX and way better!
- arecurrence 2y agoI made a bug like this once where a database default was set to a value evaluated at runtime instead of on every insert. Oops However, luckily in my case, it was caught immediately in the staging env since collisions caused exceptions. Realizing when an expression is evaluated is pretty easy to miss. That code is probably live somewhere else right now surreptitiously causing issues.
- SamBam 2y agoDid anyone else assume that this was about the Bear app, because there is no branding or link back to the author's actual project? Also, did anyone click on the double ^ at the bottom, hoping to either go back up to the top or find something about the author's company, only to find that they had upvoted the post by mistake? I'm wondering if (by this count) 167 other people might have done that.
- deleted 2y ago[deleted]
- throwawayffffas 2y agoYes, thank you, I had the exact same experience. The actual project is probably https://reworkd.ai/ https://reworkd.ai/
- goriilacoder 2y agouuidv4 is bad for a primary key. Google it. Use uuidv6 or v7. If you don't like using something not quite yet part of the standard, use v1 using time rather than MAC. Not a big deal for small tables, but for anything that will get large, you really don't want v4.
- thayne 2y agoI'm generally pretty critical of ChatGPTs code writing ability, but this mistake is something that could very easily be made by a human. On the other hand, this could and probably should, have been caught by an automated test that tried to create multiple subscriptions on a single server. Or for that matter , manual testing of creating subscriptions against a local copy. I'm not saying that to be dismissive, but one takeaway you should get from it is the value of testing before putting code in production. Edit: Another takeaway should probably be that if you have a a major bug like this, and you can't easily reproduce it, you should look harder. I bet there were some logs for errors about constraint violations in the database if you had looked for them.
- tuananh 2y agothis is same as copy paste code from stackoverflow.
- o999 2y agoThis blog post only serves as a reminder that you have to avoid overrelying on "AI".
- deleted 2y ago[deleted]
- nurple 2y agoDays since an LLM screwup cost a company money: 0
- furyofantares 2y agoFelt like a clickbait headline to me, but there's no link back to the project. So I guess not. Definitely respect for telling an embarrassing story even if I disagree wholly with the title (both the implication that the ChatGPT mistake is to blame, and that it cost them $10k.) Anyway I believe the product in question is https://agentgpt.reworkd.ai https://agentgpt.reworkd.ai
- beala 2y agoI had the same question. Everyone is talking about how this is bad for the company's reputation... but it wasn't immediately clear to me what the company is. I also eventually landed on reworkd.ai after some googling. The blog is called "asim" and the OP's username is "asim-shrestha". That lead me to this: https://www.ycombinator.com/companies/reworkd https://www.ycombinator.com/companies/reworkd They are S23, which is mention in the blog.
- wg0 2y ago>AgentGPT is an autonomous AI Agent platform that empowers users to create and deploy customizable autonomous AI agents directly in the browser. Simply assign a name and goal to your AI agent, and watch as it embarks on an exciting journey to accomplish the assigned objective. from https://docs.reworkd.ai/introduction https://docs.reworkd.ai/introduction Whereas the blogpost clearly demonstrates that AI agents cannot be left totally "autonomous", their output might seem reasonable for those not well versed in particular domain but might have disastrous consequences. VC bros are clearly gambling big on Linear Algebra.
- kazinator 2y agoI've written less than 1000 lines of Python in total probably, but I correctly spotted the problem. Python has this misfeature whereby it didn't correctly crib Common Lisp's evaluation strategy for the expressions that give default values to optional function arguments. When you have an argument like foo=obj.whatever() the obj.whatever() is evaluated (would you believe it!) at the time the definition of the function is being processed, not at the time when the function is being called and the value is needed. Is suspect this was done on purpose, for efficiency. Python has another misfeature: it has no literal syntax for certain common objects like lists. There is [1, 2, 3], but that is a constructor and not a literal: it has to create a new list every time it is evaluated and stuff it with 1, 2, 3. (Unless a clever compiler can prove that this can be optimized away without harm.) The designer didn't want a parameter like list=[] to have to construct a new, empty list object each time the argument is omitted. In Lisp '(1 2 3) and '() are true literals. Whenever they are referenced, they denote the same object. The programmer has a choice here: they can use (list 1 2 3) as the default value expression or '(1 2 3). The former is like [1, 2, 3]: it yields a new object each time that is mutable; the other will (almost certainly) yield the same object and cannot be reliably, portably modified. Hey, modern popular languages have most of the features of Lisp, so you're not missing anything.
- sdwr 2y ago> When you have an argument like foo=obj.whatever(), the obj.whatever() is evaluated at the time the definition of the function is being processed, not at the time when the function is being called. This can't be correct, surely? What if .whatever() relies on internal state that changes after obj is initialized (or after the function surrounding foo is declared, not sure what you're saying)?
- chii 2y agohttps://stackoverflow.com/questions/1132941/the-mutable-default-argument-in-python https://stackoverflow.com/questions/1132941/the-mutable-defa... and https://www.valentinog.com/blog/tirl-python-default-arguments/ https://www.valentinog.com/blog/tirl-python-default-argument... basically, having a default argument value in a function definition means to evaluate it during definition time of that function, not when the function is invoked. This is a foot gun.
- blindriver 2y agoThe worst part about this is they got 50 emails a day saying they couldn’t subscribe, and they threw their hands up for 5 days.
- moneywoes 2y agoCurious, why migrate to Python Fast API?
- cuppypastexprt 2y ago> Edit: I want to preface this by saying yes the practices here are very bad and embarrassing (and we've since added robust unit/integration tests and alerting/logging), Very believable.
- nomilk 2y agoPart of the skill in using LLMs is knowing when and how to use them, how to set the 'temperature' (how 'creative' it will be in its response), and how to write a prompt that is less prone to illusory responses. My eyes were opened one relaxing morning when sipping my coffee and pondering how to tidy up a database column by migrating from string to enum. I asked ChatGPT for its thoughts and its response seemed perfunctory and on point, until at one particular line, tucked in the otherwise sensible migration file [1], it casually recommended deleting all users whose value for that attribute wasn't among those specified by the enum. I spat my coffee out and learned a very valuable lesson that morning! [1] https://imgur.com/a/ejIdCH6 https://imgur.com/a/ejIdCH6
- qup 2y agoThat's one way to avoid bugs
- ncallaway 2y agoI’ve found it’s particularly useful at questions like “how do I do X idiomatically in Rust?” Idioms are, very conveniently, questions about what the most common shape of X among the broader community, so it tends to do quite well with that. I appreciate when it spits out example code, but I never copy/paste from it, I always rewrite anything myself to ensure I don’t slip up and accidentally… well, what it tried to sneak in to you… That’s quite the scary anecdote
- cuppypastexprt 2y ago> We had eight ECS tasks on AWS, all running five instances of our backend (overkill, yes we know, but to be fair we had AWS credits). Yes, that's a very fair reasoning. YC did the right thing by investing in this company. Fits very well with the rest of their portfolio.
- deleted 2y ago[deleted]
- animex 2y agoNo, ChatGPT made you the money that your app generated since you had no ability to implement it otherwise/without ChatGPT. Your inability to code, debug, log, monitor cost you the $10k. ChatGPT is net positive in this story.
- majormajor 2y agoIt was already implemented, seems like they had ability there: > Our project was originally full stack NextJS but we wanted to first migrate everything to Python/FastAPI. > What happened was that as part of our backend migration, we were translating database models from Prisma/Typescript into Python/SQLAlchemy. This was really tedious. We found that ChatGPT did a pretty exceptional job doing this translation and so we used it for almost the entire migration. ChatGPT wasn't a net positive if they wouldn't have tried to do this migration up-front without it. Possibly they had better error logging in the other stack, possibly they didn't, possibly they needed it less because they were actually writing the code for it themselves and knew how it worked. ("Write all the code a second time before turning on monetization" is itself an interesting decision, of course.)
- kidme5 2y ago$10k.. peanuts. Elon might lose $500B to his xAI mistake that came out today: https://grook.ai/share?id=e269e88a7b1a71eff4f176c864b30161&xai=1 https://grook.ai/share?id=e269e88a7b1a71eff4f176c864b30161&x...
- primitivesuave 2y agoI spotted the error instantly. With all due respect to your team - this has nothing to do with ChatGPT and everything to do with using a programming model that your team does not have sufficient expertise in. Even if this error managed to slip by code review, it would have been caught with virtually any monitoring solution, many of which take less than 5 minutes to set up.
- arjvik 2y agoTo be fair, if I wasn't looking for this bug I never would have spotted it. That being said, you're entirely right that any monitoring or even the most basic manual testing should have instantly caught this.
- astromaniak 2y agoEasy enough, ChatGPT can write verification code along the main line. Just ask.
- rvnx 2y ago+ humans could have done this mistake as well
- Akronymus 2y agoThose kinda issues are why I ALWAYS make an integration test with calling the same insert multiple times. I did step into that particular trap more than once (passing the result, rather than the function)
- LigmaBaulls 2y agoI spotted it right away. If you plan to use a library in your project RTFM
- KennyBlanken 2y agoIt's not some innocent mistake. The title is purposefully clickbait / keyword-y, implying that it was chatgpt that made the 'mistake' for SEO and to generate panicked clicks. "We made a programming error in our use of an LLM, didn't do any QA, and it cost us $10k" doesn't generate the C-suite "oh shit what if ChatGPT fucks up, what's our exposure!?" reaction. There's a million middle and upper management posting this article on LinkedIn, guaranteed. It's like the Mr. Beast open-mouth-surprised expression thumbnail nonsense; you feel incredibly compelled to click it. While we're on the subject: LLMs can't make "mistakes." They are not deterministic. They cannot reason, think, or do logic. They are very fancy word salad generators that use a lot of statistical probabilities. By definition they're not capable of "mistakes" because nothing they generate is remotely guaranteed to be correct or accurate. Edit: The mods boosted the post; it got downvoted into oblivion, for obvious reasons, and then skyrocketed instantly in rank, which means they boosted it: https://hnrankings.info/40627558/ https://hnrankings.info/40627558/ Hilarious that a post which is insanely clickbait (which the rules say should result in a title rewrite) got boosted by the mods. I'm sure it's a complete coincidence that the story was apparently authored by someone at a Ycombinator company: https://news.ycombinator.com/item?id=40629998 https://news.ycombinator.com/item?id=40629998
- houseplant 2y agoat this point, how can anyone trust chatGPT when we know it hallucinates things in its responses? It returns things that look like code or answers or whatever, and that's all it's trained to do. The concept of "correctness" only exists insofar as its trained, and you can't train on something that doesn't exist yet, generative AI is not creative in that sense I have no idea why people are letting chatGPT do anything without pouring over everything it says first, and at that point why bother.
- ShakataGaNai 2y agoMan. Tonight is a harsh crowd on HN. I agree, the headline is... misleading. Yes, ChatGPT made a mistake, but the issue is multiple. Perhaps it would be better if the headline was something like "A ChatGPT mistake taught us a $10k lesson". I love using ChatGPT as much as the next person, but I also have run into more than a few circumstances on simple code where it's simply... made shit up. Like AWS (Boto3) functions that don't exist at all. So any code that comes out of it gets tested and understood by me. I'll ask it to explain and dig into the docs when it does things I don't understand. That being said, the valuable lesson is in QA, debugging, logging and alerting. It's something that isn't a surprise a small (couple person, few months) startup would have done well. Often the developers of these projects are DEVELOPERS and not DevOps/SysAdmins/DBAs. The code gets written like developers do and not instrumented like a DevOps engineer would. Most get away with this for a long time (honestly, most companies get away with this for far too long). So great write up, good lesson.
- n_ary 2y agoMy questions are, why did you decide to move on to Python/FastAPI, if you did not understand it well? My second question is, why did you copy-paste as-is from CGPT without doing a review of whatever? I understand time constraints, but there should be a law or something forbidding using whatever any GPT vomits. In fact, many does it blatantly, so my employer banned using GPT codes and only gave us access to it for /entertainment/ usage.
- huygens6363 2y ago> We copy pasted the code it generated, saw everything worked fine, tried it in production, saw it also worked, and went on our merry way. Oh.. this is not good. How did you see it worked fine? You did not try inserting new customers?
- 999900000999 2y agoI have to agree it feels navie to publish this. "We're too lazy to write our own code, or even test that it works, please give us money VCs/Users." Chat GPT isn't to blame here, chat GPT is like a tireless intern. It isn't to blame if the CTO pushes it's code straight to prod. This is giving me a start up idea, code review as a service!
- cube00 2y agoGiven ChatGPT can pick up this bug when you feed it line by line [1], you might have the next startup here. [1]: https://news.ycombinator.com/item?id=40628839 https://news.ycombinator.com/item?id=40628839
- th0ma5 2y agoA lot of comments here missing the huge point that even if the fault is obvious to many, how many things were also obvious that were fixed, and how you just simply run out of attention trying to keep up with all the mistakes transformers keep putting back in the more you fight with them.
- chrismcb 2y agoNo. No it didn't. That is like saying stack overflow's mistake cost you. Or the weird of some random stranger cost you.
- iansinnott 2y agoWhy migrate to a new language? I didn't see it mentioned in the post. Seems like they already had a working solution.
- _giorgio_ 2y agoA web site with only one working page. Links to hone it to anywhere else funny work. Great. Surely a chatGPT mistake! :-D Edit: nothing works at all. 404 ʕノ•ᴥ•ʔノ ︵ ┻━┻ It looks like this page doesn't exist. Let's get you back home.
- fredthedeadhead 2y agoThe blog post is 404ing, here's a Web archive link https://web.archive.org/web/20240610032818/https://asim.bearblog.dev/how-a-single-chatgpt-mistake-cost-us-10000/ https://web.archive.org/web/20240610032818/https://asim.bear... The author has added an important edit: > I want to preface this by saying yes the practices here are very bad and embarrassing (and we've since added robust unit/integration tests and alerting/logging), could/should have been avoided, were human errors beyond anything, and very obvious in hindsight. > > This was from a different time under large time constraints at the very earliest stages (first few weeks) of a company. I'm mostly just sharing this as a funny story with unique circumstances surrounding bug reproducibility in prod (due again to our own stupidity) Please read with that in mind
- Unfrozen0688 2y ago[dead]
- patates 2y agoCould they have deleted it because of all the negativity? They did make a silly mistake, but we are humans, and humans, be it individually or collectively, do make silly mistakes.
- KennyBlanken 2y agoIf you code for a hobby/fun, yeah, sure, it's a silly mistake. If you're earning past six figures, are part of a team of programmers, call yourself an professional / engineer, and have technical management above you like a VP of Engineering, yadda yadda....then it's closer to systematic failure of the company's engineering practices than "mistake." There is a reason we call it software engineering, not software fuckarounding (or, cough, "DevOps Engineeer".) Software engineering practices assume people are going to make mistakes, and implements procedures to reduce the chances of that making it into production, and reduce the impact of those mistakes if they do make it into production.
- deleted 2y ago[deleted]
- SushiHippie 2y agohttps://web.archive.org/web/20240610032818/https://asim.bearblog.dev/how-a-single-chatgpt-mistake-cost-us-10000/ https://web.archive.org/web/20240610032818/https://asim.bear...
- prash2488 2y agoI am not python developer. And I neither intend my career to go there in near future. But I asked ChatGPT what's wrong about this code, (Not sure it's my custom instruction or not) but it always starts assuming the imports are the issue. Once I asked the imports are not the issue, It correctly pointed out, and explained the problemetic code at me... I whish they could ask another question to LLM and have an issue pointed out..
- tempcommenttt 2y agoThis is error is one that most Python programmers must have experienced early in their career. Usually the other way around, when they define a mutable default value and are hit by strange results. On when adding the current time (datetime.datetime.now()) as a function default. As a rookie programmer you might not notice this pattern, but after getting hit a few times you’ll immediately see that. And default to either a factory or to None and then set the value inside the function. ChatGPT code is only safe to use if you understand it. If you don’t, there’s always the risk that it will bite you.
- petters 2y agoIt’s an obvious error if you know Python but it’s still a mistake in the design of the Python language imo
- raggi 2y agoIf that's representative of the order volume I'm quite curious what the motivation/business case is for shifting technology stacks from one slow dynamic environment to another slow dynamic environment?
- deleted 2y ago[deleted]
- amarant 2y agoTo be brutally honest, blaming chat-gpt for a coding mistake doesn't inspire a lot of confidence. Having insufficient testing and 0 monitoring does not improve it. Y'all need to hire some senior backend Devs.
- boxed 2y agoClassic mistake. I tell newbies about this type of mistake roughly weekly hanging out on Discord. I bet it's in the training data of ChatGPT many many times.
- throw156754228 2y agoTo be honest I think part of this is a poor interface provided by SQL Alchemy. To quote an influential author on my career, Scott Meyers, "interfaces should be easy to use correctly and hard to use incorrectly."
- pmontra 2y agoAny unit test or integration test that has to create two records in that table would have failed.
- thefz 2y ago"LLMs will make programmers useless". Yeah.
- coding123 2y agoI guess, it sounds like openai is doing something right
- mikewang 2y agonot sure how. But when I asked GPT with this line and the issue was found exactly. The issue is tiny and slipery for a very big table. But I am still curious of why test can not find it. >During the work day, this was fine. We probably committed 10-20 times a day (directly to main of course) which would cause new backend deployments to occur, giving us 40 new IDs for customers to potentially use. They just use the test env for prod? When to push code, the CICD should be run and some examples should be run too here. And every time, the env should be clean. Here the database does not change from test to production.
- luismedel 2y agoI consider a shame to read some of you, literally trashing and blaming the whole team after the article. One thing is to healthly discuss how dangerous can be assuming a GPT-generated code is safe or not. Or how unit tests could identify this (could really in this specific case?). Or why you need oncall shifts and good alerts. But, come on. I spotted the issue in the code at first sight, but that doesn't make me morally superior, nor smart enough to blame someone to publicly talk about their mistake. It only means I'm currently reading that kind of code a lot, and I know where to look at. Pass me some clever ARM code and I'll be unable to spot even the most superfluous mistake. It seems HN is crowded by the most smart guys on the planet, who never had dumb mistakes and are SO "quality inclined" they need to blame someone for theirs. edit: of course, the decision to make it public is questionable, but that's topic for another thread, IMHO.
- Edd1CC 2y agoThey had 8 AWS tasks running 5 instances each with code written in TypeScript and Python, with frameworks like next.js, with $40 revenue and only a few weeks dev time? What the actual fuck hahahaha This is made worse when they edit saying the reason the codes crap is because of time constraints, but spent their time refactoring across languages and spinning up a distributed system FOR NO REASON. That is self imposed harm juggling features and ridiculous technical complexity. What were they thinking. Edit: a YC summer ‘23 company who’s product is still behind a waitlist summer ‘24, presumably because of a rewrite to Rust
- wg0 2y agoBecause they have half a million dollar to star with and 1.2 million dollars on top and then free AWS credits to burn.
- Edd1CC 2y agoIf I give them a trampoline they shouldn't spend all day jumping on it, just because they can. Especially if they're busy with features. They literally had 1 instance of the backend per $1 of revenue, and the reason the bug wasn't seen straight away was because they had 40 backend instances each with a single uuid that could be used for users before it broke with non-unique id errors.
- TheNewsIsHere 2y agoThey even said as much in TFA - "[...] overkill, yes we know, but to be fair we had AWS credits". Fully acknowledging the irony I am about to invoke - this is why I hate startup culture. Not startups, but this ridiculous culture of "well the VC gave us a million bucks and that bought us $100,000 in AWS credits, so let's just use it." As someone who has built my company fully on my own dime (and the dimes of two colleagues), it's easy enough to burn piles of money in AWS (or any other cloud) when you're making an attempt at being judicious. Spinning up eight backends (edit: running five instances each, no less!) just because you have money, despite the fact that you know you don't need that much compute, is just insane. If for no other reason than you're just throwing your own credits away.
- seba_dos1 2y agoIt's not a ChatGPT's mistake. In fact, it never is.
- datavirtue 2y agoI have many examples of human mistakes starting at $5MM and going down from there.
- monkpit 2y agoThis link gives a 404 for me. Edit: the entire subdomain gives a 404 actually.
- yawnxyz 2y agoI think changing the title to "Not monitoring our ChatGPT calls costed us $10k" would make it stronger. Adding monitoring is the last thing you think about when pushing out a prototype, and it's easy to forget that a "prototype that no one will probably use" could cost thousands with accidental infinite loops and bugs like these. Always set your spending limits!
- divbzero 2y agoFor future reference, the ChatGPT’s mistake in the SQLAlchemy model: id = Column(String, primary_key=True, default=str(uuid.uuid4()), unique=True, nullable=False) Can be fixed by making the default a function: id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()), unique=True, nullable=False) Or, if you choose to use native UUID types in Python and SQL: id = Column(Uuid, primary_key=True, default=uuid.uuid4, unique=True, nullable=False)