3 ms·
From their blog [0] The bug was discovered in the Redis client open-source library, redis-py. As soon as we identified the bug, we reached out to the Redis mai
by dilippkumar 4y ago
From their blog [0]
The bug was discovered in the Redis client open-source library, redis-py. As soon as we identified the bug, we reached out to the Redis maintainers with a patch to resolve the issue. Here’s how the bug worked:
- We use Redis to cache user information in our server so we don’t need to check our database for every request.
- We use Redis Cluster to distribute this load over multiple Redis instances.
- We use the redis-py library to interface with Redis from our Python server, which runs with Asyncio.
- The library maintains a shared pool of connections between the server and the cluster, and recycles a connection to be used for another request once done.
- When using Asyncio, requests and responses with redis-py behave as two queues: the caller pushes a request onto the incoming queue, and will pop a response from the outgoing queue, and then return the connection to the pool.
- If a request is canceled after the request is pushed onto the incoming queue, but before the response popped from the outgoing queue, we see our bug: the connection thus becomes corrupted and the next response that’s dequeued for an unrelated request can receive data left behind in the connection.
- In most cases, this results in an unrecoverable server error, and the user will have to try their request again.
- But in some cases the corrupted data happens to match the data type the requester was expecting, and so what gets returned from the cache appears valid, even if it belongs to another user.
At 1 a.m. Pacific time on Monday, March 20, we inadvertently introduced a change to our server that caused a spike in Redis request cancellations. This created a small probability for each connection to return bad data.
This bug only appeared in the Asyncio redis-py client for Redis Cluster, and has now been fixed.
[0] https://openai.com/blog/march-20-chatgpt-outage https://openai.com/blog/march-20-chatgpt-outage
- merek 4y agoThanks for the summary.
- doodlesdev 4y agoA few questions this still doesn't answer: why was there credit card data in their Redis cluster, and why were they the only ones affected that I know of? Still, this response seems much more reasonable than what I was expecting from OpenAI considering the seriousness of the issue. Anyways, thank you very much for the summary, I didn't find this blog post in the publications that were talking about the issue.
- deleted 4y ago[deleted]
- jonny_eh 4y agoA think a reasonable lesson from this is to double check that the returned data belongs to the user you're about to serve it to. So include some way to associate the owning user with each cache entry, so it can be verified later when it's recalled.