26 ms·
Entropy, a CLI that scans files to find high entropy lines (might be secrets)
- deleted 2y ago[deleted]
- DLA 2y agoThis looks like a very handy CLI tool. Nice Go code also. Thanks.
- lanfeust 2y agothanks!
- trulyhnh 2y agoGoDaddy open sourced something similar https://github.com/godaddy/tartufo https://github.com/godaddy/tartufo
- thsksbd 2y agoThis is very cool, but I have a thought - I see this as a last line of defense, and I am concerned that this would this give a false sense of security leading people to be more reckless with secrets.
- 0cf8612b2e1e 2y agoEhhh considering how low the security bar is, I think it is better than nothing. If you inherit a code base, make it a quick initial action to see how much pain you can expect. In practice, I expect a tool like this has so many false positives you cannot keep it as an always running action. More a manual review you run occasionally. I hope that more secrets adopt a GitHub like convention where they are prefaced with an identifier string so that you do not require heuristics to detect them.
- lanfeust 2y agoIndeed. I open-sourced `entropy` after we discovered an actual secret leak in our client codebase
- alexchantavy 2y agoThe pie in the sky goal for any security org is to have a cred rotation process that is so smooth that you’re able to not worry about leaked creds because it’s so fast and easy to rotate them. If the rotation is automated and if it’s cheap and frictionless to do so, heck why not just rotate them multiple times a day.
- bongodongobob 2y agoNo, it's a way to audit and see the modes that your security policy is failing. At least that's how I look at it.
- kmoser 2y agoYou could make the same argument for any tool that does not provide high security. In fact security is layered, and no single tool should be relied upon to be your one security tool. You said as much yourself: "I see this as a last line of defense," but I don't see how you conclude that this would inherently cause people to be more reckless with secrets.
- lm411 2y agoWelcome to information security :)
- coppsilgold 2y agoNote that in an adversarial setting this will only be effective against careless opponents. If you properly encode your secret it will have the entropy of its surroundings. For example you can hide a string of entropy (presumably something encrypted) in text as a biased output of an LLM. To recover it you would use the same LLM and measure deviations from next-token probabilities. This will also fool humans examining it as the sentence will be coherent.
- textninja 2y agoWhat you described sounds like a very cool idea - LLM-driven text steganography, basically - but intentional obfuscation is not the problem this tool is trying to solve. To your point about secrets with entropy similar to the surrounding text, however, I wonder if this can pick up BIP39 Seed Phrases or if whole word entropy fades into the background.
- buildbot 2y agoIn general (for those unaware) this is called stenography. You can hide an image in the lower bits of another image for example too.
- dragonwriter 2y agoSteganography; stenography is completely different.
- BeefWellington 2y agoSee also: - trufflehog: https://github.com/trufflesecurity/trufflehog https://github.com/trufflesecurity/trufflehog - detect-secrets: https://github.com/Yelp/detect-secrets https://github.com/Yelp/detect-secrets - semgrep secrets: https://semgrep.dev/products/semgrep-secrets https://semgrep.dev/products/semgrep-secrets -- (Paid, but may be included in existing licenses in some cases
- bbno4 2y agoAlso see PyWhat for both interesting strings and secrets https://github.com/bee-san/pyWhat https://github.com/bee-san/pyWhat
- jonstewart 2y agonoseyparker is another good one: https://github.com/praetorian-inc/noseyparker https://github.com/praetorian-inc/noseyparker I think these solutions are all much better for finding secrets than something naive based on entropy. Yes, entropy is more general but these are well established tools that have been through the fire of many, many data sets.
- upg1979 2y agoSee also: https://github.com/gitleaks/gitleaks https://github.com/gitleaks/gitleaks
- xedeon 2y agoggshield from GitGuardian has been great for us. Their free service can also auto detect and notify you of leaked secrets, passwords or high entropy lines from your online repos. https://github.com/GitGuardian/ggshield https://github.com/GitGuardian/ggshield
- icapybara 2y agoGonna have to explain how a “high entropy line” is calculated and why it might be secrets.
- daemonologist 2y agoEntropy of information is basically how well it can be compressed. Random noise usually doesn't compress much at all and thus has high entropy, whereas written natural language can usually be compressed quite a bit. Since many passwords and tokens will be randomly generated or at least nonsense, looking for high entropy might pick up on them. This package seems to be measuring entropy by counting the occurrences of each character in each line, and ranking lines with a high proportion of repeated characters as having low entropy. I don't know how closely this corresponds with the precise definition. Source: https://github.com/EwenQuim/entropy/blob/f7543efe130cfbb5f0afa9eb98abffe7a87241ba/main.go#L166 https://github.com/EwenQuim/entropy/blob/f7543efe130cfbb5f0a... More: https://en.wikipedia.org/wiki/Entropy_(information_theory) https://en.wikipedia.org/wiki/Entropy_(information_theory)
- eru 2y agoOf course, this heuristic fails for weak passwords. And it fails for passphrases like 'correct battery horse staple', which have a large enough total entropy to be good passwords, but have a low entropy per character.
- dumbo-octopus 2y ago4 diceware words is hardly a good password. It's ~51 bits of entropy, about the same as 8 random ascii symbols. It could be trivially cracked in less than an hour. Your average variable name assigned to the result of an object name with a method name called with a couple parameter names has much more entropy.
- conradludgate 2y agoIf you can crack a single 52bit password in an hour, that's suggesting you can crack a 40bit password every second. That's 1 trillion hashes per second.
- hamasho 2y agoI didn't know what entropy means in software, so here's the definition[0]: ---- Software entropy is a measure of the disorder or complexity of a software system. It is a natural tendency for software entropy to increase over time, as new features are added and the codebase becomes more complex. High entropy in software development means that the code is difficult to understand, maintain, and extend. It is often characterized by: Duplicated code: The same code or functionality is repeated in multiple places, which can make it difficult to find and fix bugs. Complex logic: The code is difficult to follow and understand, which can make it difficult to add new features or fix bugs without introducing new ones. Poor documentation: The code is not well-documented, which can make it difficult for new developers to understand and contribute to the codebase. Technical debt: The code has been patched and modified over time without proper refactoring, which can lead to a tangled and cluttered codebase. Low entropy in software development means that the code is well-organized, easy to understand, and maintain. It is often characterized by: Well-designed architecture: The code is structured in a logical way, with clear separation of concerns. Consistent coding style: The code follows a consistent coding style, which makes it easy to read and understand. Comprehensive documentation: The code is well-documented, with clear explanations of the code's purpose and functionality. Minimal technical debt: The code has been refactored regularly to remove technical debt, which makes it easy to add new features and fix bugs without introducing new ones. [0] https://www.kisphp.com/python/high-and-low-entropy-in-software-development https://www.kisphp.com/python/high-and-low-entropy-in-softwa...
- itemize 2y agothanks for the search. this is textual entropy however, I am not sure if definition is applicable
- tonyabracadabra 2y agointeresting! can the similar measurement be applied to finding redundant code (like low entropy) with extra works?
- cowsaymoo 2y agoI transcend this problem by making all my database passwords 'abcd'
- kgeist 2y agoThe tool found "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz1234567890" in our codebase as a high entropy line :)
- g15jv2dp 2y agoWell, it is...
- saurik 2y agoI mean, it certainly has a low Kolmogorov complexity (which is what I would really want to be measuring somehow for this tool... note that I am not claiming that is possible: just an ideal); I am unsure whether how that affects the related bounds on Shannon entropy, though.
- jraph 2y ago…a very verbose way to match alphanumeric characters :-)
- ngneer 2y agoThen use it as your password ;)
- josephg 2y agoYou can use LLMs as compressors, and I wonder how it would go with that. The approach is simple: Turn the file into a stream of tokens. For each token, ask a language model to generate the full set of predictions based on context, and sort based on likelihood. Look where the actual token appears in the sorted list. Low entropy symbols will be near the start of the list, and high entropy tokens near the end. I suspect most language models would deal with your alphabet example just fine, while still correctly spotting passwords and API keys. It would be a fun experiment to try!
- krick 2y agoIs there any good posts about the use of entropy for tasks like that? I am wondering for quite some time of how do people actually use it and if it is any effective, but never actually got to investigating the problem myself. First of all, how to define "entropy" for text is a bit unclear in the first place. Here it's as simple as `-Sum(x log(x))` where x = countOccurences(char) / len(text). And that raises a lot of questions about how good this actually works. How long string needs to be for this to work? Is there a ≈constant entropy for natural languages? Is there a better approach? I mean, it seems there must be: "obviously" "vorpal" must have lower "entropy" than "hJ6&:a". You and I both "know" that because 1) the latter "seems" to use much larger character set than natural language; 2) even if it didn't, the ordering of characters matters, the former just "sounds" like a real word, despite being made up by Carroll. Yet this "entropy" everybody seems to use has no idea about any of it. Both will have exactly the same "entropy". So, ok, maybe this does work good enough for yet-another-github-password-searcher. But is there anything better? Is there more meaningful metric of randomness for text? Dozens of projects like this, everybody using "entropy" as if it's something obvious, but I've never seen a proper research on the subject.
- hackinthebochs 2y agoEntropy is a measure of complexity or disorder of a signal. The interesting part is that the disorder is with respect to the proper basis or dictionary. Something can look complex in one encoding but be low entropy in the right encoding. You need to know the right basis, or figure it out from the context, to accurately determine the entropy of a signal. A much stronger way of building a tool like the OP is to have a few pre-computed dictionaries for a range of typical source texts (source code, natural language), then encode the string against each dictionary, comparing the compressibility of the string. A high entropy string like a secret will compress poorly against all available dictionaries.
- jazzyjackson 2y agobookmarking to think about later... does this hold for representing numbers as one base compared to another? Regarding a prime as having higher entropy / less structure than say a perfect square or highly divisible number a prime is a prime in any base, but the number of divisors will differ in non-primes, if the number is divisible by the base then it may appear to have more structure (smaller function necessary to derive, kolmogorov style), does prime factorization have anything to do with this? i can almost imagine choosing a large non-prime whose divisibity is only obvious with a particular base such that the base becomes the secret key - base of a number is basically specifying your dictionary, no?
- janniehater 2y ago[dead]
- g15jv2dp 2y agoWhy would I need to install go to run this tool? I thought one advantage of go was that devs could just distribute a single binary file that works...
- benterix 2y agoBecause it's a security tool so trusting a binary upfront defeats the purpose. With source you at least have the option to inspect what it really does.
- menacingly 2y agodoes the stated purpose of the tool influence whether or not you can trust it?
- spoonjim 2y agoIf you're trying to improve the security of your product by running random binaries from the Internet you're going to have a bad time
- saagarjha 2y agoThat's how most people run compilers
- benterix 2y agoThis is argumentum ad absurdum - there is a reason why trusting your kernel and compiler is a reasonable compromise, even though there might be security issues in them, but random pieces of software downloaded from the Internet is not.
- Ensorceled 2y agoWait ... you download random compilers from the internet? Or are you asserting equivalence between getting go from Google or Xcode from Apple and an random home brew install?
- kqr 2y agoInteresting. If I had to do this, I would have done something like perl -lne 'next unless $_; $z = qx(echo "$_" | gzip | wc -c); printf "%5.2f %s\n", $z/length($_), $_' on the principle that high entropy means it compresses badly. However, that uses each line as the dictionary, rather than the entire file, so it has a little trouble with very short lines which compress badly. It did react to this line return map { $_ > 1 ? 1 : ($_ < 0 ? 0 : $_) } @vs; which is valid code but indeed seems kind of high in entropy. I was also able to fool it to not detect a high-entropy line by adding a comment of natural English to it. I'm on the go but it would be interesting to see comparisons between the Perl command and this tool. The benefit of the Perl command is that it would run out of the box on any non-Windows machine so it might not need to be as powerful to gain adoption.
- blixt 2y agoI guess you could take all lines in the file except the one you're testing and measure the filesize, then add the line and measure again. The delta should then be more fair. You could even do this by concatenating all code files and then testing line by line across the entire repo, but that would probably be too slow.
- GuB-42 2y agoI would use a better compressor than gzip but I have done this trick several times. xz or zstd may be better choices, or you can look at Hutter Prize [1] winners for best compression and therefore best entropy estimate. [1] http://prize.hutter1.net/ http://prize.hutter1.net/
- nequo 2y ago> best compression and therefore best entropy estimate That's a good point. But the Hutter Prize is for compressing a 1 GB file. On inputs as short as a line of code, gzip doesn't do so badly. For a longer line: $ INPUT=' bool isRegPair() const { return kind() == RegisterPair || kind() == LateRegisterPair || kind() == SomeLateRegisterPair; }' $ echo "$INPUT" | gzip | wc -c 95 $ echo "$INPUT" | bzip2 | wc -c 118 $ echo "$INPUT" | xz -F xz | wc -c 140 $ echo "$INPUT" | xz -F lzma | wc -c 97 $ echo "$INPUT" | zstd | wc -c 92 For a shorter line: $ INPUT=' ASSERT(regHi().isGPR());' $ echo "$INPUT" | gzip | wc -c 48 $ echo "$INPUT" | bzip2 | wc -c 73 $ echo "$INPUT" | xz -F xz | wc -c 92 $ echo "$INPUT" | xz -F lzma | wc -c 51 $ echo "$INPUT" | zstd | wc -c 46
- crazypython 2y agoIt would be interesting to see a variant of this that used a small language model to measure entropy.
- saagarjha 2y agoWhy would you do that when measuring entropy is easy to do with a normal program
- saagarjha 2y agoI assume this will have a bad time on compressed files?
- lanfeust 2y ago.zip extension is ignored by default along with other binary formats :)
- saagarjha 2y agoRight but like .tar.gz, etc. are also a thing
- lanfeust 2y agoYou can just add your extensions to ignore with --ignore-ext. But I'll add .tar.gz and .tar.bz2 since they are widely used.
- frumiousirc 2y agoOr, have the tool recursively read the .tar files' contents.
- seethishat 2y agoThis reminds me of the program 'ent' (which I have used for a very long time) https://fourmilab.ch/random/ https://fourmilab.ch/random/
- MarkMarine 2y agoAnother way to do this would be to compress the file and compare the compressed size to the uncompressed size. Encrypted files do not compress well compared to code, I saw a phd thesis that postulated an inverse ratio of compression efficiency to performance data mining, this would be the opposite
- p0w3n3d 2y agoxkcd.com/936/
- blixt 2y agoI guess a language model like Llama 3 could model surprise on a token-by-token basis and detect the areas that are most surprising, i.e. highest entropy. Because as one example mentioned, the entire alphabet may have high entropy in some regards, but it should be very unsurprising to a code-aware language model that in a codebase you have the Base62 alphabet as a constant.
- weipe-af 2y agoIt would be useful if it also trawled through the full git history of the project - a secret could have been checked in and later removed, but still exist in the history.
- thomascountz 2y agoThank you DrJones for asking what a high entropy string is several years ago[0] and linking to a good article on it.[1] [0] https://news.ycombinator.com/item?id=13304641 https://news.ycombinator.com/item?id=13304641 [1] https://www.splunk.com/en_us/blog/security/random-words-on-entropy-and-dns.html?301=/blog/2015/10/01/random-words-on-entropy-and-dns.html https://www.splunk.com/en_us/blog/security/random-words-on-e...
- baryphonic 2y agoNeat tool. Would be cool if this CLI could have a flag to read .gitignore and exclude all of the contents automatically. Also it might be cool to have different strategies for detecting secrets, e.g. Kolmogorov complexity as other comments have noted.