8 ms·
What will happen when you commit secrets to a public Git repo?
- _the_inflator 6y agoNice honey pot experiment.
- 0xad 6y agoThanks!
- weejewel 6y agoThanks for sharing. Did you also investigate what they actually did with the keys?
- 0xad 6y agoYou mean adversaries? No. For token generation I used https://canarytokens.org/ https://canarytokens.org/ so the only information I got was abou triggering the token, but not the context in which it was triggered. BTW. GitHub (apart from GitGuardian) also has Secret Scanning feature [1] that basically allows the provider to act on the leaked secret. Amazon is integrated and it should invalidate and inform the owner but this also went to Thinkst, not me, so I don't know if it was actually invalidated and alerted. [1] https://developer.github.com/partnerships/secret-scanning/ https://developer.github.com/partnerships/secret-scanning/
- deleted 6y ago[deleted]
- greysteil 6y agoCool experiment! I PM the secret scanning team at GitHub and wanted to mention what GitHub did behind the scenes here. GitHub scans every commit to a public repo for secrets one of our secret scanning partners may have issued. We forward those candidate secrets to the issuing partner, and they take action. In some cases they auto-revoke the secret (AWS normally does this, I believe), in some cases they notify the user, and in some cases the response is configurable. I checked that GitHub detects these tokens myself - within 1 second of the commit GitHub had notified AWS and Slack of the leak. AWS and Slack will then have taken action and informed the token owner, which in this case is Thinkst Canary, rather than Andrezj (the OP). I believe AWS normally auto-revoke, but they may have a custom setup with Thinkst Canary's tokens that allows Thinkst to continue to monitor them even once compromised. Finally, GitHub actually delays the indexing of our search by a couple of seconds to ensure that, for normal cases, our secret scanning partners have time to take action before anyone else can find the tokens. We're always looking to make secret scanning at GitHub better, so feedback as always welcome. It's also fascinating (and validating!) to see what happens to exposed tokens. * List of GitHub secret scanning partners: https://docs.github.com/en/free-pro-team@latest/github/administering-a-repository/about-secret-scanning#about-secret-scanning-for-public-repositories https://docs.github.com/en/free-pro-team@latest/github/admin... * Thnkst Canary tokens: https://www.canarytokens.org/ https://www.canarytokens.org/
- 0xad 6y agoAwesome, thanks for the background information!
- tobr 6y agoWhy not refuse to publish a detected secret at all until the repo owner takes an action to allow it?
- watt 6y agoThink about it in terms of incentives and nudges.
- greysteil 6y agoThere are a few considerations on that one, but one very practical reason is the developer experience of dealing with false positives. False positives are one of the big problems in secret scanning. Some partners issue credentials with patterns that make them very hard to distinguish from innocuous strings. For example, a Datadog token looks identical to a commit SHA. We would never block developers from pushing commits to GitHub just because they had 40 character hexadecimal strings in them! GitHub's partnership approach works around the false positive problem by having the token issuer check whether a token is real and take action only if it is. However, this is a one-way communication from GitHub - the token issuer doesn't need to tell us whether the candidate secret we sent them was real or not, and in most cases we never know. As a result, we can't replicate the zero false positive experience in a pre-receive hook (i.e., before the commit is pushed to GitHub). There would also be performance considerations from making 30+ http requests as part of a pre-receive hook. In future, we are looking at creating a pre-receive hook solution that focuses on patterns that have a very low false positive rate. There are already some open source solutions that do this (links below) - in fact the OP linked to one from his Twitter thread. If/when GitHub offer is, it will definitely be opt-in, rather than opt-out! * https://github.com/thoughtworks/talisman/ https://github.com/thoughtworks/talisman/ * https://github.com/awslabs/git-secrets https://github.com/awslabs/git-secrets
- 0xad 6y ago
- rwmj 6y agoIs there a way (outside Github) that adversaries can get access to the "full feed" of commits? I don't understand how the attackers can find a new key from all the changes that must go into github across millions of repos, within 11 minutes.
- bostik 6y agoThere are bots (some even run by security and threat intel companies) feeding off of the firehose. For a public display of one type of scanning functionality, take a look at shhgit[0,1]. 0: https://www.shhgit.com/ https://www.shhgit.com/ 1: https://github.com/eth0izzle/shhgit https://github.com/eth0izzle/shhgit
- rwmj 6y agoIs the firehose public or do these companies have a relationship with github? If the latter, I assume github doesn't give the firehose feed to attackers who are only looking for AWS keys.
- greysteil 6y agoThe author of sshgit wrote a great post on how it works using the public GitHub API: https://darkport.co.uk/blog/ahh-shhgit!/ https://darkport.co.uk/blog/ahh-shhgit!/
- deleted 6y ago[deleted]
- oefrha 6y agoThe firehose is simply the /events endpoint of GitHub API v3 off all public events. It’s delayed by 5 minutes. Anyone has access (subject to rate limits of course, which is 5000/hr when authenticated?). You can even have a look at the response in your browser, without any authentication: https://api.github.com/events https://api.github.com/events Docs: https://developer.github.com/v3/activity/events/#list-public-events https://developer.github.com/v3/activity/events/#list-public... https://docs.github.com/en/free-pro-team@latest/rest/reference/activity#list-public-events https://docs.github.com/en/free-pro-team@latest/rest/referen...
- notRobot 6y agoJust use a fucking blog, man. I'm so sick of threads like this. Edit: I'm sorry, this came off as way more aggressive than I intended. I get why people use twitter to share stuff like this, but it's much harder to archive, find or reference in the future, not to mention it being much less readable than a simple webpage. To anyone reading this, please consider publishing your findings on a blog as well as on twitter.
- sofixa 6y agoOr post it on Reddit or Medium or whatever if you can't be bothered with a blog. Twitter "threads" need to die.
- forgotmypw17 6y agoReddit and Medium are no better in terms of weight and complexity.
- kaszanka 6y agoold.reddit.com is better than the abomination that is the most recent Twitter redesign. Though new Reddit is truly terrible, probably even worse than Twitter.
- xeyownt 6y ago> Twitter "threads" need to die. I would even remove "threads". Sick of all the hate and fakedom.
- forgotmypw17 6y agoYou can use Nitter to make it a bit more readable and a bit less bloated. https://nitter.net/andrzejdyjak/status/1324360905237372929 https://nitter.net/andrzejdyjak/status/1324360905237372929
- 0xad 6y agoHey, OP here. I agree that a blog post would be more readable. In this particular case I just didn't expect that it will catch fire. If I would then I would spend more time on the form. I won't make that mistake again (i.e. in the future I will use a blog post as main driver of such twitter thread).
- yyyk 6y agoIn some cases, github will require you to remove the offending file from the commit history - or make the repo private. e.g. https://github.com/aliostad/deep-learning-lang-detection https://github.com/aliostad/deep-learning-lang-detection
- forgotmypw17 6y agoMore readable version: https://nitter.net/andrzejdyjak/status/1324360905237372929 https://nitter.net/andrzejdyjak/status/1324360905237372929
- 0xad 6y agoGreetings fellow Hackers! OP here. I see that my experiment got some traction which means more awareness should be spread about this class of bugs. For starters I recommend reading "How Bad Can It Git" [1] and "Detecting and Mitigating Secret-Key Leaks inSource Code Repositories" [2] papers. After that you can read "How I made $10K in bug bounties from GitHub secret leaks" [3] and some notable reports on HackerOne Hacktivity [4] [5] and [6]. This last one is interesting - leaking secrets is not only about code repository! Actually it's about entire toolset used for software development, hence secret scanning could (should?) be performed for other places such as CICD logs or even Slack messages [7]. Anyhow, back to code repositories. GitHub and GitLab both recognized secrets as a problem, so they came up with solutions. If you use GitHub you can easily integrate GitGuardian [8] into your workflow ($$$) but even if you don't GitHub provides you with Secret Scanning feature [9] (both are mentioned within the Twitter and HN threads). If you use GitLab you have a Secret Detection feature [10] at your disposal BUT in order to use it you need to setup Auto DevOps (that's why in my experiment GitLab didn't alert me - I just pushed commits to my public repo but didn't setup anything). Apart from built-in solutions provided by GitHub and GitLab, one can use tooling of their own choice. For this I'd recommend two types of solutions: proactive and reactive. For proactive security, as mentioned in the Twitter thread, you can use Talisman [11] as pre-commit hook. For reactive security you can use GitLeaks [12] (used by GitLab) or similar tools - there are many of them but one stands out, namely truffleHog [13] which can sniff each and every commit across all branches (also used by GitLab). What if you already commited a secret into the public repository? Start with revoking and continue with this tutorial [14] gl, hf. [1] https://www.ndss-symposium.org/ndss-paper/how-bad-can-it-git-characterizing-secret-leakage-in-public-github-repositories/ https://www.ndss-symposium.org/ndss-paper/how-bad-can-it-git... [2] https://people.eecs.berkeley.edu/~rohanpadhye/files/key_leaks-msr15.pdf https://people.eecs.berkeley.edu/~rohanpadhye/files/key_leak... [3] https://tillsongalloway.com/finding-sensitive-information-on-github/index.html https://tillsongalloway.com/finding-sensitive-information-on... [4] https://hackerone.com/reports/716292 https://hackerone.com/reports/716292 [5] https://hackerone.com/reports/396467 https://hackerone.com/reports/396467 [6] https://hackerone.com/reports/496937 https://hackerone.com/reports/496937 [7] https://github.com/PaperMtn/slack-watchman https://github.com/PaperMtn/slack-watchman [8] https://www.gitguardian.com/ https://www.gitguardian.com/ [9] https://developer.github.com/partnerships/secret-scanning/ https://developer.github.com/partnerships/secret-scanning/ [10] https://docs.gitlab.com/ee/user/application_security/sast/#secret-detection https://docs.gitlab.com/ee/user/application_security/sast/#s... [11] https://github.com/thoughtworks/talisman https://github.com/thoughtworks/talisman [12] https://github.com/zricethezav/gitleaks https://github.com/zricethezav/gitleaks [13] https://github.com/dxa4481/truffleHog https://github.com/dxa4481/truffleHog [14] https://docs.github.com/en/free-pro-team@latest/github/authenticating-to-github/removing-sensitive-data-from-a-repository https://docs.github.com/en/free-pro-team@latest/github/authe...
- thdrdt 6y agoIt is amazing how fast and effective those bots are. I remember one time I installed Windows 95/98. I wanted the PC to be on internet but did not have a firewall for Windows. But I knew the internet address where I could get one. So after installing Windows I took my chances, connected to the internet, downloaded the firewall asap, installed it, and was already too late. The PC was compromised within 10 minutes and I had to reinstall it.
- peterwwillis 6y agoThis would make a good blog post. Maybe they should consider making one so we can have the article [in an easily readable/shareable/updateable form] after it's deleted from Twitter
- 0xad 6y agoOP here. I'm planning to do so, however it will require more work (better description of the problem, wider description of viable solutions, additional case studies). Most probably it will land on Medium and Dev.to.
- jefftk 6y agoI think secret detection is great overall, but the only times I've run into it are false positives with client side API keys that are by their nature public. For example, I recently configured something to use the Google calendar API from JavaScript on the client. It's fully safe to check in this key, since it is intended to be run in client-side JavaScript anyway, but I was still nagged about it.
- mackenzie-gg 6y agoIt's a difficult challenge. Secrets detection is probabilistic, without checking the credentials it's nearly impossible to determine, with 100% accuracy, a true vs a false positive. But it has made big improvements. What detection solutions have you been using?
- jefftk 6y agoI get automated emails that I didn't sign up for from GitGuardian: "GitGuardian has detected the following Google Key exposed within your GitHub account." My understanding was that they could use an API to check whether it was a real key, but perhaps that doesn't say whether it is a client-side or server-side key?
- halfjoking 6y agoYou'll probably get an email from AWS that your account is compromised and you have 5 days to rotate your keys or your account could be terminated. Then everyday they email you to see if you made any progress rotating the keys. I made this meme about it that my boss didn't find funny. https://imgur.com/ZCUu9rr https://imgur.com/ZCUu9rr
- 0xad 6y agoYes you will, but only because GitHub already recognised this class of problems and came up with their own solution [1]. Bear in mind that it works only for vendors that integrated, so while it's true for AWS it might not be for your FOO API. I giggled at meme. [1] https://developer.github.com/partnerships/secret-scanning/ https://developer.github.com/partnerships/secret-scanning/
- remram 6y agoCouldn't we come up with a standard format for secret keys, that would make it obvious they are a secret and which service they're from? This would make scanners easier to implement, and would remove the requirement to partner with GitHub to get your key format supported. AWS uses an `AKIA` prefix for access keys (but none for secrets), SendGrid uses an `SG.` prefix on API keys, etc.