4 ms·
> we had a significant issue in ChatGPT due to a bug in an open source library, for which a fix has now been released and we have just finished validating. WHI
by doodlesdev 4y ago
> we had a significant issue in ChatGPT due to a bug in an open source library, for which a fix has now been released and we have just finished validating.
WHICH open source library? Seems like they fucked up and are looking for a scapegoat. There's no excuses for issues like this. Also, I believe this is yet again one more reason for companies and developers to consider row-level security [0], it simply erases the possibility of a malformed query returning data that shouldn't be accessible altogether.
[0]: https://www.postgresql.org/docs/current/ddl-rowsecurity.html https://www.postgresql.org/docs/current/ddl-rowsecurity.html
- dilippkumar 4y agoFrom their blog [0] The bug was discovered in the Redis client open-source library, redis-py. As soon as we identified the bug, we reached out to the Redis maintainers with a patch to resolve the issue. Here’s how the bug worked: - We use Redis to cache user information in our server so we don’t need to check our database for every request. - We use Redis Cluster to distribute this load over multiple Redis instances. - We use the redis-py library to interface with Redis from our Python server, which runs with Asyncio. - The library maintains a shared pool of connections between the server and the cluster, and recycles a connection to be used for another request once done. - When using Asyncio, requests and responses with redis-py behave as two queues: the caller pushes a request onto the incoming queue, and will pop a response from the outgoing queue, and then return the connection to the pool. - If a request is canceled after the request is pushed onto the incoming queue, but before the response popped from the outgoing queue, we see our bug: the connection thus becomes corrupted and the next response that’s dequeued for an unrelated request can receive data left behind in the connection. - In most cases, this results in an unrecoverable server error, and the user will have to try their request again. - But in some cases the corrupted data happens to match the data type the requester was expecting, and so what gets returned from the cache appears valid, even if it belongs to another user. At 1 a.m. Pacific time on Monday, March 20, we inadvertently introduced a change to our server that caused a spike in Redis request cancellations. This created a small probability for each connection to return bad data. This bug only appeared in the Asyncio redis-py client for Redis Cluster, and has now been fixed. [0] https://openai.com/blog/march-20-chatgpt-outage https://openai.com/blog/march-20-chatgpt-outage
- merek 4y agoThanks for the summary.
- doodlesdev 4y agoA few questions this still doesn't answer: why was there credit card data in their Redis cluster, and why were they the only ones affected that I know of? Still, this response seems much more reasonable than what I was expecting from OpenAI considering the seriousness of the issue. Anyways, thank you very much for the summary, I didn't find this blog post in the publications that were talking about the issue.
- deleted 4y ago[deleted]
- jonny_eh 4y agoA think a reasonable lesson from this is to double check that the returned data belongs to the user you're about to serve it to. So include some way to associate the owning user with each cache entry, so it can be verified later when it's recalled.
- bspammer 4y agoWouldn’t this require creating a database role for every user? Then your backend needs to be able to assume every user role, which still seems to allow for bugs.
- jansommer 4y agoThis seems to be a common misconception. You do not need a role for every user. You can assign session variables, at least in Postgres, and use them in your policies.
- adql 4y ago> Also, I believe this is yet again one more reason for companies and developers to consider row-level security [0], it simply erases the possibility of a malformed query returning data that shouldn't be accessible altogether. That requires 1:1 DB user to real user mapping + one connection to DB per users, it is entirely horrible performance-wise, which is why it is so rarely used for anything more than separating apps from eachother.
- jansommer 4y agoNo it doesn't. No need for 1:1 db user mapping.