5 ms·
Really interesting write-up. Curious how they do logging, and how they were able to get such a detailed look at the previous requests, including the HTTP respon
by thecodemonkey 6y ago
Really interesting write-up. Curious how they do logging, and how they were able to get such a detailed look at the previous requests, including the HTTP response. Wouldn’t it be a massive amount of data if all HTTP requests/responses were logged? (Not to mention the security implications of that)
- bob1029 6y agoIt would indeed be a massive amount of data, but bear in mind that only a small fraction of HTTP requests hitting GitHub are actually both authenticated and attempting to change the state of the system. Most requests are unauthenticated read-only and cause no state changes in GitHub.
- vlovich123 6y agoIt’s a reasonable theory but step one required to trigger this bug is an unauthenticated request. It’s unclear if their logs indicated that or they got lucky trying to repro that someone thought to try with an unauthenticated request first when the back to back session didn’t repro.
- whimsicalism 6y agoFrom the writeup, it sounds like the latter, but we obviously can't be sure.
- aidos 6y agoCould be that there’s some tell tale in the logs. We recently fixed a bug with incomplete requests by noticing in nginx logs that the size of the headers + the size of the body was the size of the buffer of a linux socket.
- firebaze 6y agoAs expensive as it is, datadog provides the required level of insight via, among other drill-down methods, detailed flamegraphs, to get to the bottom of problems like this one. No, not affiliated with datadog, and not convinced the cost/benefit ratio is positive for us in the long run.
- nullsense 6y agoWe've implemented it recently and it's helped us tremendously.
- ckuehl 6y agoPurely a guess, but I've sometimes been able to work backwards from the content length (very often included in HTTP server logs) to figure out what the body had to be.
- xvector 6y agoDistributed tracing with a small sample size or small persistence duration?
- dmlittle 6y agoGiven the rarity of this bug I doubt a small sample sized would have captured the right requests to be able to debug this. The probability would be the probability of this bug occurring * (sample percentage)^3 [ you'd need to have sampled the correct 3 consecutive requests].
- deleted 6y ago[deleted]
- SilverRed 6y agoYes it is a large amount of data but its worth paying to store it because it lets you go back in time and work out wtf happened which will save you far more hours and solve more problems making it worth it.
- kevingadd 6y agoIt's pretty common to do exhaustive logging even if the data is enormous. It's valuable for debugging and security investigations. As far as the security of storing log data, you need log messages to not contain PII regardless of how long you retain them. It's not uncommon to see companies talk about processing terabytes or even petabytes of log data per day. In the end the storage isn't that expensive [1], and you can discard everything after a little while. [1] Assuming of course that any company generating a petabyte of logs every day is doing so because they have an enormous number of customers, AFAIK the cost of storing 1PB of logs on something like Glacier is in the 100-300k USD/yr range. Absolutely a tolerable expense for the peace of mind you get by knowing you can go back in 2 months and find every single request that touched a security vulnerability or hunt down the path an attacker took to breach your infrastructure.
- marshmellman 6y agoAny idea how GitHub could tell that the cookies in the response are wrong, without logging the session cookie?
- londons_explore 6y agoJust log the session cookie...
- andix 6y agoI guess the incident happened more often, then they wanted to tell us. So as a first step they probably turned up the logging and waited for new reports...