17 ms·
March 20 ChatGPT outage: Here’s what happened
- ajhai 4y ago> In the hours before we took ChatGPT offline on Monday, it was possible for some users to see another active user’s first and last name, email address, payment address, the last four digits (only) of a credit card number, and credit card expiration date This is a lot of sensitive data. It says 1.2% of ChatGPT Plus subscribers active during a 9 hour window, which considering their user base must be a lot.
- mach1ne 4y agoIt’s a bit unclear if this means that 1.2% of all chatGPT Plus subscribers were active during that 9-hour window
- pixl97 4y agoThere are 2 hard problems in computer science: cache invalidation, naming things, and off-by-1 errors.
- deathanatos 4y ago… in this case this variant seems more appropriate: There are 3 hard problems in Computer Science: 1. naming things 2. cache invalidation 3. 4. off-by-one errors concurrency
- sergiotapia 4y agoIt sounds like their redis key was not unique enough and yada yada yada it returned sensitive info the wrong people.
- Jabrov 4y agoDid you read the article? That’s not at all what happened.
- fintechie 4y agoI reported this race condition via ChatGPT's internal feedback system after I saw other user's chat titles loading on my sidebar a couple of times (around 7-8 weeks ago). Didn't get a response, so I assumed it was fixed... Hopefully they'll start a bug bounty program soon, and prioritise bug reports over features.
- totallyunknown 4y agosame to my. actually only the summary of the history was from a different user. the content itself was mine.
- sebzim4500 4y agoThe claim made at the time was that the titles were not from other people and were in fact caused by the model hallucinating after the input query timed out (or something like that). Obviously that sounds a little suspect now, but it might be true.
- nwienert 4y agoThat's a lie if so, if you look at the Reddit threads there's no way those were not specific other users histories as they had the logical history of reading browser history. Eg, one I saw had stuff like "what is X", then the next would be "How to X" or something. Some were all in Japanese, others all in Chinese. If it was random you wouldn't see clear logical consistency across the list.
- jetrink 4y agoThe explanation at the time was that unavailable chat data (due to, e.g. high load) resulted in a null input sometimes being presented to the chat summary system, which in turn caused the system to hallucinate believable chat titles. It's possible that they misdiagnosed the issue or that both bugs were present and they caught the benign one before the serious one.
- fintechie 4y ago
- _5hxt 4y agoThe bug: https://github.com/redis/redis-py/issues/2624 https://github.com/redis/redis-py/issues/2624
- braindead_in 4y agoWas this written by ChatGPT? Maybe it found the bug as well, who knows.
- _5hxt 4y ago.... "I am asking for this ticket to be re-opened, since I can still reproduce the problem in the latest 4.5.3. version"
- chatmasta 4y agoThe PR: https://github.com/redis/redis-py/pull/2641 https://github.com/redis/redis-py/pull/2641 According to the latest comments there, the bug is only partially fixed.
- photochemsyn 4y ago> "If a request is canceled after the request is pushed onto the incoming queue, but before the response popped from the outgoing queue, we see our bug: the connection thus becomes corrupted and the next response that’s dequeued for an unrelated request can receive data left behind in the connection." The OpenAI API was incredibly slow and lots of requests probably got cancelled (I certainly was doing that) for some days. I imagine someone could write a whole blog post about how that worked, it would be interesting reading.
- stygiansonic 4y agoThe key part: If a request is canceled after the request is pushed onto the incoming queue, but before the response popped from the outgoing queue, we see our bug: the connection thus becomes corrupted and the next response that’s dequeued for an unrelated request can receive data left behind in the connection.
- deleted 4y ago[deleted]
- nvartolomei 4y agoI wonder how much time passed between the first case of corruptions leading to exceptions (and they ignored it as “eh, not great not terrible we’ll look at it later) and users reporting seeing other’s users data?
- jchw 4y agoDoes anyone else find it a bit off-putting how much emphasis they keep putting on "open source library"? I don't think I've read about this without the word open source appearing more than once in their own messaging about it. Why is it so important to emphasize that the library with the bug is open source? The cynic in me wants to believe that it's a way of deflecting blame somehow, to make it seem like they did their due diligence but were thwarted by something outside of their control. I don't think it holds. If you use an open source library with no warranty, you are responsible (legally and otherwise) to ensure that it is sufficient. For example, if you break HIPAA compliance due to an open source library, it is still you who is responsible for that. But of course, they're not claiming it's anyone else's fault anywhere explicitly, so it's uncharitable to just assume that's what they meant. Still, it rubs me the wrong way. I can't fight the feeling that it's a wink wink nudge nudge to give them more slack than they'd otherwise get. It feels like it's inviting you to just criticize redis-py and give them a break. The open postmortem and whatnot is appreciated and everything, but sometimes it's important to be mindful of what you emphasize in your postmortems. People read things even if you don't write them, sometimes.
- bitxbitxbitcoin 4y agoNot surprising from a company that calls itself openai. The “open source” keyword stuffing is so people associate the open from openai with open source. Psyops I mean marketing 101.
- amtamt 4y agoWas postmortem generated by chatGPT?
- dilap 4y agoI half agree, but I also half-sympathize with them, because it really wasn't their fault -- it was a quite-bad bug in a very fundamental library. Bugs happen, though. Especially in Python.
- airstrike 4y ago> Especially in Python. as opposed to...?
- kristianpaul 4y agoThis more a data leak than an outage…
- sebzim4500 4y agoIt was down for quite a while, so I would call it an outage.
- killerstorm 4y agoSerious question: Why do people feel it's necessary to use a redis cluster? I understand in early 2000s we were using spinning disks and it was the only way. Well, we don't use spinning disks any more, do we? A modern server can easily have terabytes of RAM and petabytes of NVMe, so what's stopping people from just using postgres? A cluster of radishes is an anti-pattern.
- amtamt 4y agoFor caching somewhat larger objects based on ETag?
- lofaszvanitt 4y agoPeople know it, that's all.
- cplli 4y agoFor caching the query results you get from your database. Also it's easier to spin up Redis and replicate it closer to your user than doing that with your main database. From my experience anyway.
- mike_hearn 4y agoI think the idea is that if your db can hold the working set in RAM and you're using a good db + prepared queries, you can just let it absorb the full workload because the act of fetching the data from the db is nearly as cheap as fetching it from redis.
- killerstorm 4y ago> For caching the query results you get from your database. This only makes sense if queries are computationally intensive. If you're fetching a single row by index you aren't winning much (or anything).
- dpkirchner 4y agoOf course? I'm not really sure what the original question actually is if you know that users benefit from caching the results of computationally intensive queries.
- qwerty3344 4y agousing chatGPT to summarize: This section summarizes the issue with the ChatGPT service which caused some users to be able to see another user's chat history and payment information. OpenAI took the ChatGPT service offline and published technical details of the problem. The bug was found in the Redis client open-source library and has been patched. OpenAI has reached out to affected users and improved their systems to prevent this from happening again. They have also expressed their commitment to protecting user privacy and their dedication to the open-source community.
- zamnos 4y agoInteresting. I pasted in the outage report, and then asked ChatGPT to give me a single sentence summary and it gave me an answer way longer than a sentence. I pointed out that it wasn't a single sentence and it gave me: > ChatGPT experienced an outage due to a bug in an open-source library that allowed some users to see titles from another active user's chat history and payment-related information of 1.2% of ChatGPT Plus subscribers was exposed; the bug has since been patched and actions have been taken to prevent future incidents. I then asked it why there was the outage and it spit out: > The exact cause of the Facebook outage is unknown, but it is believed to be related to a configuration change in the company's routers. It's likely because I ran out of tokens because the OpenAI outage report is long. Pasting in the text of the outage report, and then re-asking about why, it was able to give a much better answer: > There was an outage due to a bug in an open-source library that allowed some users to see titles from another active user's chat history and also unintentionally exposed payment-related information of 1.2% of ChatGPT Plus subscribers who were active during a specific nine-hour window. Querying it further, again having to repeat the whole OpenAI outage report, and asking it a few different ways I eventually managed to get this succinct answer: > The bug was caused by the redis-py library's shared pool of connections becoming corrupted and returning cached data belonging to another user when a request was cancelled before the corresponding response was received, due to a spike in Redis request cancellations caused by a server change on March 20. It did take me more than a few minutes to get to there, so just actually reading the report would have been faster, and I ended up having to read the report to verify that answer was correct and not a hallucination anyway, so our jobs are safe for now.
- m00dy 4y agomaybe they've just scrolled over issue lists of popular tech stacks and cherry-picked the most compelling one to bury the dirt.
- jkern 4y agoFunnily enough I've had a very similar bug occur in an entirely separate redis library. It was a pretty troubling failure mode to suddenly start getting back unrelated data
- ketchupdebugger 4y agoIt's surprising that openai seems to be the only one being affected. If the issue is with redis-py reusing connections then wouldn't more companies/products be affected by this?
- zzzeek 4y agotheir description of the problem seemed kind of obtuse, in practice, these connection-pool related issues have to do with 1. request is interrupted 2. exception is thrown 3. catch exception, return connection to pool, move on. The thing that has to be implemented is 2a. clean up the state of the connection when the interrupted exception is caught, then return to the pool. that is, this seems like a very basic programming mistake and not some deep issue in Redis. the strange way it was described makes it seem like they're trying to conceal that a bit.
- roberttod 4y agoIt's an open source library, I assume that logic is abstracted within it and that the "basic mistake" was one of the maintainer's.
- YetAnotherNick 4y agoThere is a 1 year old autoclosed issue which is very similar to OpenAI's issue: https://github.com/redis/redis-py/issues/2028 https://github.com/redis/redis-py/issues/2028
- neurostimulant 4y agoI think most app using redis-py rarely cancel async redis command. Python async web frameworks is gaining popularity, but the majority of people using python for their web application is not using an async framework. And of those people that do use them, not many of them canceling async redis requests often enough to trigger the bug.
- 19h 4y agoIt boggles my mind how they're not absolutely checking the user & conversation id for EVERY message in the queue given the possible sensitivity of the requests. How is this even remotely acceptable? In the one reddit post first surfacing this the user saw conversations related to politics in china and other rather sensitive topics related to CCP. This can absolutely get people hurt and they absolutely must take this serious.
- zaroth 4y agoIt doesn’t boggle my mind at all. Session data appears, and is used to render the page. Do you verify every time the actual cookie and go back to the DB to see what user it pointed to? No, everyone assumes their session object is instantiated with the right values at that level of the code.
- rvz 4y agoFirst GitHub, then OpenAI. Two of Microsoft finest(!) (majority owned and acquired) companies on the top of HN announcing a serious security incident. It's quite unsettling to see this leak of highly sensitive information and a private key exposure as well. Doesn't look good and seems like they don't take security seriously.
- skybrian 4y agoIn the case of OpenAI, the product is more of a research demo that had to be drastically scaled up, though. From an operations point of view it’s more like a startup.
- deltree7 4y agoNobody cares and yet another case study of HN being out-of-touch with reality
- rvz 4y ago> "Nobody cares" Yet another case study of absolutism, which can be simply dismissed. People paying for ChatGPT care once it goes down and getting their details and chats leaked that is certainly outside of HN. Same with GitHub. Both having ~100M users between them. That's the reality.
- sebzim4500 4y agoI'm paying for ChatGPT and I don't care about this any more than the many, many other services I use that have at some point had an embarassing security issue.
- deltree7 4y agoI'm paying and I don't care. If I write perfect bug-free code, lead a perfect life, live in a perfect world, I'd be upset. But, I know that shit happens and the reliability meter should be flexible for different things (bridges, heart surgery and chat agent). If I train my brain to bitch, whine, moan about every thing, I'd not have resources to care about really important things.
- picodguyo 4y agoIf you're subscribed to their status page, you'll know it's actually unusual for a day to go by without an outage alert from OpenAI. They don't usually write them up like this but I guess this counts as PII leak disclosure for them? For having raised billions of dollars the are comically immature from a reliability and support perspective.
- thequadehunter 4y agoTo be fair, they accidently made a game-changing breakthrough that gained millions of users overnight, and I don't think they were ready for it. Before chatgpt, most normal people had never heard of OpenAI. Their flagship product was basically an API that only programmers could make useful. Team leaders at OpenAI have stated that they were not expecting the success, let alone the highest adoption rate for any product in history. In their minds, it was just a cleaned-up version of a 2-year old product. It was billed as a research preview. So, all of a sudden you go from hiring mostly researchers because you only have to maintain an API and some mid-traffic web infra, to suddenly having the fastest growing web product in history and having to scale up as fast as you can. Keep in mind that they didn't get backing from Microsoft until January 23, 2023-- that was only 2 months ago. I'd say we should cut them some slack.
- picodguyo 4y agoThese problems predate ChatGPT. Their API has been on the market for nearly 3 years. And they raised their first $1B in 2019. That's plenty of money and time to hire capable leadership.
- deleted 4y ago[deleted]
- thequadehunter 4y agoYeah but again, this is the fastest growing app in history and it uses way more compute than your standard webapp, and basically delivers all functionality from a single service that handles that load. I can see why there would be some growing pains.
- deleted 4y ago[deleted]
- layer8 4y agoThat sounds like the kind of bug that could be prevented by modeling with TLA+.
- w10-1 4y agoIt's interesting (read: wrong) for an AI company to bother writing the user interface for their web application. This was a failure of integration testing and defensive design, whether the component was open-source or not. There's no reason to believe that an AI company would have the diligence and experience to do the grunt work of hardening a site. But management obviously understood the level and character of interest. Actual users include probably 10,000 curiosity seekers for every actual AI researcher, with 1,000 of those being commercial prospects -- people who might buy their service. This is a clear sign that the managers who've made technical breakthrough's in AI are not capable even of deploying the service at scale -- no less managing the societal consequences of AI. The difficulty with the board getting adults in the room is that leaders today give the appearance of humility and cooperation, with transparent disclosures and incorporation of influencers into advisory committees. The leaders may believe their own abilities because their underlings don't challenge them. So there's no obvious domineering friction, but the risk is still there, because of inability to manage. Delegation is the key to scaling, code and organizations. "Know thyself" is about knowing your limits, and having the humility to get help instead of basking in the puffery of being in control. This isn't a PR problem. It's the Achilles' heel of capitalism, and the capitalists in OpenAI's board should nip this incipient Musk in the bud or risk losing 2-3 orders of magnitude return on their investment.
- chatmasta 4y agoWhy did it take them 9 hours to notice? The problem was immediately obvious to anyone who used the web interface, as evidenced by the many threads on Reddit and HN. > between 1 a.m. and 10 a.m. Pacific time. Oh... so it was because they're based in San Francisco. Do they really not have a 24/7 SRE on-call rotation? Given the size of their funding, and the number of users they have, there is really no excuse not to at least have some basic monitoring system in place for this (although it's true that, ironically, this particular class of bug is difficult to detect in a monitoring system that doesn't explicitly check for it, despite being immediately obvious to a human observer). Perhaps they should consider opening an office in Europe, or hiring remotely, at least for security roles. Or maybe they could have GPT-4 keep an eye on the site!
- eep_social 4y agoStaffing an actual 24x7 rotation of SREs costs about a million dollars a year in base salary as a floor and there are few SREs for hire. A metrics-based monitor probably would have triggered on the increased error rate but it wouldn’t have been immediately obvious that there was also a leaking cache. The most plausible way to detect the problem from the user perspective would be a synthetic test running some affected workflow, built to check that the data coming back matches specific, expected strings (not just well-formed). All possible but none of this sounds easy to me. Absolutely none of this is plausible when your startup business is at the top of the news cycle every single day for the past several months.
- sosodev 4y ago"there are few SREs for hire" How do you figure? If you mean there are few SRE with several years of experience you might be right. SRE is a fairly new title so that's not too surprising. However, my experience with a recent job search is that most companies aren't hiring SRE right now because they consider reliability a luxury. In fact, I was search of a new SRE position because I was laid off for that very reason.
- chatmasta 4y ago
- ElijahLynn 4y agoThat is a pretty good disclosure that creates trust.
- m_0x 4y agoDid they use chat-gpt to fix the bug?
- deleted 4y ago[deleted]
- qwertox 4y agoNice writeup, it's fair in the content presented to us. Yet I'm wondering why there is no checking if the response does actually belong to the issued query. The client issuing a query can pass a token and verify upon answer that this answer contains the token. TBH as a user of the client I would kind of expect the library to have this feature built-in, and if I'm starting to use the library to solve a problem, handling this edge-case would be of a somewhat low priority to me if the library wouldn't implement it, probably because I'm lazy. I hope that the fix they offered to Redis Labs does contain a solution to this problem and that everyone of us using this library will be able to profit from the effort put into resolving the issue. It doesn't [0], so the burden is still on the developer using the library. [0] https://github.com/redis/redis-py/commit/66a4d6b2a493dd3a20cc299ab5fef3c14baad965 https://github.com/redis/redis-py/commit/66a4d6b2a493dd3a20c... --- Edit: Now I'm confused, this issue [1] was raised on March 17 and fixed on March 22, was this a regression? Or did OpenAI start using this library on March 19-20? Interesing comment: > drago-balto commented 3 hours ago > Yep, that's the one, and the #2641 has not fixed it fully, as I already commented here: #2641 (comment) > I am asking for this ticket to be re-oped, since I can still reproduce the problem in the latest 4.5.3. version [1] https://github.com/redis/redis-py/issues/2624#issue-1629335140 https://github.com/redis/redis-py/issues/2624#issue-16293351...
- deleted 4y ago[deleted]
- menzoic 4y agoThat sounds more like a hindsight thing. In most systems authorization doesn't happen at the storage layer. Most queries fetch data by an identifier which is only assumed to be valid based on authorization that typically happens at the edge and then everything below relies on that result. It's not the safest design but I wouldn't say the client should be expected to implement it. That security concern is at the application layer and the actual needs of the implementation can be wildly different depending on the application. You can imagine use cases for redis where this isn't even relevant, like if it's being used to store price data for stocks that update every 30 seconds. There's no private data involved there. It's out of scope for a storage client to implement.
- deleted 4y ago
- kkarpkkarp 4y agois this only me who don't see any chat history since yesterday and generally chat does not work (you can type the message, but clicking button or hitting enter / ctrl+enter) does not give any effect? in chat history there is a button to "retry" but clicking it and inspecting the result, you see "internal server error"
- giza182 4y agoIve had the exact same issue since last week. it never fixed itself if that’s what you’re waiting for. I had to resort to creating a new account with a different email to get access again. Contacted support but yet to hear back. Not sure if it was the cause the issue happened when I try to buy Plus and that failed.
- kkarpkkarp 4y agoI found it was issue with Firefox and its default privacy settings. Clicking shield icon in Firefox address bar and setiting this privacy guard off made Chat working again
- qwertox 4y agoThis reminds me of a comment I made 1.5 months ago [0]: I was logging in during heavy load, and after typing the question I started getting responses to questions which I didn't ask. gdb answered on that comment "these are not actually messages from other users, but instead the model generating something ~random due to hitting a bug on our backend where, rather than submitting your question, we submitted an empty query to the model." I wonder if it was the same redis-py issue back then, but just at another point in the backend. His answer didn't really convince me back then. [0] https://news.ycombinator.com/item?id=34614796&p=2#34615875 https://news.ycombinator.com/item?id=34614796&p=2#34615875
- semiquaver 4y agoI'd put money on it.
- galnagli 4y agoWell - they have had more bugs and will have more bugs to worry from. https://twitter.com/naglinagli/status/1639343866313601024 https://twitter.com/naglinagli/status/1639343866313601024
- lopkeny12ko 4y agoThe original issue report is here: https://github.com/redis/redis-py/issues/2624 https://github.com/redis/redis-py/issues/2624 This bit is particularly interesting: > I am asking for this ticket to be re-oped, since I can still reproduce the problem in the latest 4.5.3. version Sounds like the bug has not actually been fixed, per drago-balto.
- benmmurphy 4y agoThis is a common bug with a lot of software. For example some HTTP clients that do pooling won’t invalidate the connection after timing out waiting for the response.
- DeathArrow 4y agoI'm the only one terrible bored by the assault of the trivial AI news last months? Every fart some AI related person makes becomes a huge news. And it's followed by tens of random blog postings all posted to HN.
- celdon25 4y agoAt least it isn't about the Rust language this time grumbles
- DeathArrow 4y agoBecause Rust hasn't conquered AI the way it conquered crypto. But we will see AI stuff rewritten in Rust quite soon.
- jpeter 4y agoI bet againt it
- spprashant 4y agoFor some reason I liked reading about Rust (or any other technology) a lot more that the AI. Part of it is that, the average engineer could understand and grok what those articles were talking about, and I could appreciate, relate, and if applicable criticize it. The AI news just seems to swing between hype and doomsday prophecies, and little discussion about the technical aspects of it. Obviously OpenAI choosing to keep it closed source makes any in-depth discussion close to impossible, but also some of this is so beyond the capabilities of an average engineer with a laptop. It can be frustrating.
- Kuinox 4y agoI managed to manually produce this bug 2 months ago. As they don't have any bug bounty, I didn't submitted it. By starting a conversation and refreshing before ChatGPT has time to answer, I managed to reproduce this bug 2-3 times in January.
- breckenedge 4y agodid you reach out via https://openai.com/security.txt https://openai.com/security.txt?
- Kuinox 4y agoNo, as I said, their disclosure page says there is no and I'm not a professional security researcher so I was not very interested in helping them. I find even it funny now, they could write that they will provide free API creds, but now they had a very bad moment due to their greed.
- Kuinox 4y ago*there is no bug bounty Missed a word.
- MacroChip 4y agoYou won't fire off a quick email nor warn others because there's no bug bounty?
- catmanjan 4y agoNot everyone has the privilege of working for free
- capableweb 4y agoAs I understand, the "work" was already done, the only thing missing was sending a heads-up email with "hey, this seems iffy, maybe you ought to look into it". I dunno, I generally report issues I find in software, paid or not, as I've always done. Takes usually ~10 minutes and 1% of the time, they ask for more details and I spend maybe 20 minutes more to fill out some more details. Never been paid for it ever, most I gotten was a free yearly subscription. But in general I do it because I want what I use to be less buggy.
- davedx 4y agoThey were/are storing payment data in redis? LOL!
- taxman22 4y agoThe postmortem doesn’t say that. It just says they were caching “user information”. Maybe that includes a Stripe customer or subscription ID that they look up before sending an email, for example.
- tmpz22 4y agoYeah probably the session id and when the wrong session id is returned other operations like GET User details would pull its data from relational storage.
- doodlesdev 4y agoYes they were. It says they stored billing address and some pieces of credit card data.
- polyrand 4y agoCommit fixing the bug: https://github.com/redis/redis-py/commit/66a4d6b2a493dd3a20cc299ab5fef3c14baad965 https://github.com/redis/redis-py/commit/66a4d6b2a493dd3a20c...
- abujazar 4y agoThe disclosure is provides valuable information, but the introduction suggests someone else or «open-source» is to blame: >We took ChatGPT offline earlier this week due to a bug in an open-source library which allowed some users to see titles from another active user’s chat history. Blaming an open-source library for a fault in closed-source product is simply unfair. The MIT licensed dependency explicitly comes without any warranties. After all, the bug went unnoticed until ChatGPT put it under pressure, and it was ChatGPT that failed to rule out the bug in their release QA.
- voidfunc 4y agoThey're not blaming anyone. Objectively there was a bug in that library that caused the problem.
- Matl 4y agoWhy mention it is open source then? What does that add?
- fsckboy 4y ago"there was a bug in an outside library that we used" does not mention open source but has the same meaning, and would probably provoke the same complaints ("they're trying to blame somebody else for their problem"). In that case, though, they could say "look, we used a popular open source library because we had more faith that it would be better tested and correct" which would be a compliment to open source. That's essentially the information that we have. In today's world, who builds anything from anything close to stratch? embedded developers, probably come closest. It's no worse or better to say "there was a bug that our release uncovered." If they continue to announce as many details as possible, we as the audience can develop a sense whether they're creating bugs or just uncovering bugs we're glad to know about.
- abujazar 4y agoI think it's fair to mention the bug originated in redis-py, but I don't find it relevant at all to mention «open-source» in the opening line of the public statement about the outage. Or «outside library» for that matter. It was ChatGPT release QA that failed, and then they failed to admit it.
- LarsDu88 4y agoI called it: https://news.ycombinator.com/item?id=35267569#35270165 https://news.ycombinator.com/item?id=35267569#35270165
- syspec 4y agoChance this was not 90% written by ChatGPT is 10%
- daguava 4y ago[dead]
- daguava 4y ago[dead]
- hackerbrother 4y agoI don’t fault them for having an outage. It’s hard to think of any other recent site with a comparable popularity spike.
- taf2 4y agoSounds like they need to use a lua function to ensure that pop and push operation remains atomic within the redis instance? For example: spopsadd local src_list = ARGV[1] local dst_list = ARGV[2] local value = redis.call('spop', src_list) if value then -- avoid pushing nils redis.call('sadd', dst_list, value) end return value
- AaronFriel 4y agoI suspect the fatal error OpenAI saw was an "SSL error: decryption failed or bad mac"? Or if SSL were disabled, they likely would see the sort of parsing error vaguely described as an "unrecoverable server error" due to streams of data from one request being swapped with another, with incorrect data structures, bad alignment, etc. I can see how if the stars aligned and SSL were disabled this data race would manifest in viewing another user's request, so long as the sockets were receiving similar responses when they were swapped. I suspect the issue is deeper than the bug recently fixed in just redis-asyncio. Libraries written prior to asyncio/green threads often have that functionality enabled by means of monkey patches or shims that juggle file handles or other shared state, and there are data races. When SSL is used hopefully those data races reading from sockets result in MAC errors and the connection is terminated. It's easy to imagine the same mistake happening one level up, in message passing code that manages queues or other shared data structures. Search for "python django OR celery OR redis OR postgres OR psycopg2 decryption failed or bad mac" and you'll see the scale of the issue, it's fairly widespread. I don't have a high degree of confidence in this ecosystem. I've written about this before on Hacker News[1], and I'm not confident in the handling of shared data structures in Python libraries. I don't think I can blame any maintainers here, there's a huge number of people asking for these concurrency features and the way that it's often implemented - monkey patching especially - makes it extremely difficult to do correctly. [1] https://news.ycombinator.com/item?id=31065472 https://news.ycombinator.com/item?id=31065472
- richdougherty 4y ago> Actions we’ve taken > > - [testing, assertions, log alerting, debugging] I reported this bug in the forums (no response), because I couldn't find an official way to report bugs - at least it wasn't documented anywhere that I could find. The actions they've taken fix that bug, but not the next one. OpenAI should take an action to improve their bug reporting channels, so the next bug gets found more quickly. https://community.openai.com/t/bug-incorrect-chatgpt-chat-session-titles-seeing-other-users-sessions/101921 https://community.openai.com/t/bug-incorrect-chatgpt-chat-se... Other messages indicate issues with bug reporting too: * https://news.ycombinator.com/item?id=35291943 https://news.ycombinator.com/item?id=35291943 * https://news.ycombinator.com/item?id=35295747 https://news.ycombinator.com/item?id=35295747
- wz366 4y agoUnbelievable, have many times have you seen AWS explain an outage (or PII leak like this kind) as open source library bug? Have they asked 5 whys? Why the bug was deployed into production? Why there was not enough testing before deployment? Why integration testing was only done with low concurrency? Why standard release testing procedure is missing? Why there's no synthetic traffic testing and gates before rolling to 100% in production?
- kaustyap 4y agoNot sure why no one is talking about serious data breach of personal and credit card information in this case. On the contrary, everyone is very concerned about compromise of github ssh key in another thread.
- grogers 4y agoI've long thought that it is often better to return a bit of extra data in internal API responses to validate that the response matches the request sent. That can be fairly simple like parroting a request ID, or including some extra metadata (e.g. part of the request) to validate the response is valid. It's not the most efficient, but it can safe your bacon sometimes. Mixing up deployment stacks (e.g. thinking you are talking to staging but actually it's prod) and mixing user data are pretty scary, so any defense in depth seems useful.