7 ms·
Incident Report: Inadvertent Private Repository Disclosure
- kozak 10y agoThis honest report is a good example of transparency.
- jorge_leria 10y agoGithub takes security seriously, this disclosure post is a proof of that.
- wojcech 10y agoProbs to github for the disclosure. And congratulations to gitlab for probably getting a nice boost in on premise support contracts:)
- OJFord 10y agoGithub Enterprise is on-premise too. I don't know that this would make you necessarily want to make both the change to self-hosting, and the change of platform.
- stonogo 10y agoBecause there is one critical characteristic in a private repository, and they failed to execute. Moving on-prem doesn't fix that failure, it just mitigates fallout.
- iancarroll 10y agoIt seems highly unlikely this commit made it into a GitHub Enterprise release.
- stonogo 10y agoWe'll never know, which is a problem unto itself.
- bsder 10y agoI approve of the handling, but this just underscores why you want self-hosted instances.
- ifhs 10y agoDon't you mean on-premise?
- Normal_gaussian 10y agoDatacentre's have outstanding track records, and if you secure your box correctly there are few ways to compromise it. On-premise will either be incredibly costly or missing key protections or infrastructure. Self-hosted git (through the many installable git servers or raw git) running on a correctly sized box is almost certainly the way to go
- newsat13 10y agoThis github vulnerability has nothing to do with insecure box. It has to do with a bad application logic. This can happen anywhere - self-hosted or not.
- bsder 10y agoThat is true. But, if I'm running my own instance, the probability is that it doesn't matter if someone else gets access via bad application logic. Everybody is probably employed by the same company. It's a difference of degree: compromising my self-hosted or on-premise server means that somebody already in my employ has more access than they should. If I'm a small organization, that probably doesn't matter. If I'm a big organization, I probably have an IT staff to deal with this and the people involved are still "nominally" under my control. The github mistake means that people completely unrelated to my repository can get access.
- ifhs 10y agoWell, the code can simply be made public by bad application logic. Which is why I thought you where talking of on premise where the intranet will seal off outsiders
- Hovertruck 10y agoWe received an email from Github yesterday informing us that one of our repositories had been accessed by a third party due to this issue. While it's not a fun notification to receive, it definitely made our general security paranoia feel justified – we're lucky that from the get-go we've held best practices around keeping secrets out of the codebase. Obviously we still dedicated time as a team to prune through our repository history with a fine-toothed comb for anything that could potentially be a vulnerability, as we take this very seriously. One of our engineers came up with a useful script to grab all unique lines from the history of the repository and sort them according to entropy. This helps to lift any access keys or passwords which may have been committed at any point to the top. I think this is a great example to illustrate the tough edges of security to less experienced engineers. Github will most likely never let something like this happen to you, but on the off-chance that they do it's great to be prepared. Additionally, the response from Github was very well received. No excuses, just a thorough explanation of what happened. I also can't help but mention that we're hiring, if you'd like to work at an organization that values security and data privacy very highly. :) usebutton.com/join-us
- foota 10y agoI'm curious, how did they calculate entropy? My first thought was to do something with Huffman encoding.
- jasonmoo 10y agoI wrote the script in question and actually used a simple shannon entropy value. (http://codereview.stackexchange.com/questions/868/calculating-entropy-of-a-string/909#909 http://codereview.stackexchange.com/questions/868/calculatin...). It worked well enough help rule out several problem spaces.
- pavel_lishin 10y agoWould you mind posting the script? I'd love to run it against our codebase and see what it comes up with. It might be a fun thing to open source as part of a "I've inherited a project, what now?" toolkit that helps you decide what to fix.
- webmaven 10y agoInteresting that they don't mention expanding the information being logged to make the multiple joins they had to do unnecessary or more deterministic.
- revelation 10y agoNext step: setup development system ?! Surely they do some end-to-end testing?
- uxp 10y agoThey state in the post that of 17 million requests to their git-proxy server, only 230 of those requests could be identified as successful responses to incorrect data/repos, at a percent of 0.0013%. I don't know of anyone that would recommend creating tests, even integration tests, that hammers a service to check to see if something like one hundredths of one percent of requests returns invalid data. If anything, the fact that a script is hammering a service that probably (in a Dev or QA environment) has much less data in it's database and file stores, and much less protection (like load balancing and caching) than it would in production would generate more false positives than it would generate in substantial data disclosure regression defects.
- smashed 10y agoBut the overwhelming majority of requests failed with errors. The happy-path was not tested either.
- manquer 10y agoNormally perhaps not, but if you host other people's IP . The risk of a leak like this can have major economic consequences to other organisations which trusted you with security for their code. To me it looks like poor design , I would expect private repos to be hosted completely independently and in isolation with more secure and throughly audited code with longer release cycle (LTS ?) after the code has been well tested in the public free repos. It is not excessive if you consider the potential value of the private repos that github has control over. They already do something similar for enterprise edition. It leaves bad taste that smaller customers are not treated with similar caution
- Silhouette 10y agoOne of the most striking things about this report is the scale that GitHub has now reached: the whole incident apparently lasted only 10 minutes, but during that time 17 million requests were sent to their git proxy. It's obviously unfortunate in this case, since even a relatively small and quickly fixed bug affecting a tiny proportion of requests still had serious consequences. However, it's a remarkable achievement (if also a little terrifying for the software development industry from a single-point-of-failure perspective).
- faitswulff 10y agoHow did they become aware of the bug so quickly (<10 minutes)? Unless I'm missing something from the report, it doesn't say.
- cheiVia0 10y agoThe only private repository is one you created on your own computer and didn't push to github.
- 0x0 10y agoLeaking private repositories is one thing, but if you have a private build server that pulls and runs scripts, you could be in for a bad time even if you ended up pulling a random public repository, if the build script is malicious... hmmm...
- deleted 10y ago[deleted]
- vemv 10y agoIn retrospect of course it's always easy to criticise, but still, the diff is really cringeworthy. The deleted code is very specific-looking. Nobody writes that just casually or out of ignorance. Also it is what was at use in production. It's very naive to just go and replace that with nice-looking, shorter code. Key lessons: - Understand what you are deleting - Treat production code as sacred - Add reasonably extensive comments for delicate code (as the original one). Git commit messages aren't enough. - Try out infrastructure changes in production-like staging servers. I really doubt they properly did, as they say the "majority" of 17M requests failed.