3 ms·
I transcend this problem by making all my database passwords 'abcd'
by cowsaymoo 2y ago
I transcend this problem by making all my database passwords 'abcd'
- kgeist 2y agoThe tool found "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz1234567890" in our codebase as a high entropy line :)
- g15jv2dp 2y agoWell, it is...
- saurik 2y agoI mean, it certainly has a low Kolmogorov complexity (which is what I would really want to be measuring somehow for this tool... note that I am not claiming that is possible: just an ideal); I am unsure whether how that affects the related bounds on Shannon entropy, though.
- jraph 2y ago…a very verbose way to match alphanumeric characters :-)
- ngneer 2y agoThen use it as your password ;)
- josephg 2y agoYou can use LLMs as compressors, and I wonder how it would go with that. The approach is simple: Turn the file into a stream of tokens. For each token, ask a language model to generate the full set of predictions based on context, and sort based on likelihood. Look where the actual token appears in the sorted list. Low entropy symbols will be near the start of the list, and high entropy tokens near the end. I suspect most language models would deal with your alphabet example just fine, while still correctly spotting passwords and API keys. It would be a fun experiment to try!
- randomtoast 2y agoReminds me of https://xkcd.com/936/ https://xkcd.com/936/ I think "correct horse battery staple" has a low entropy, since it is just ordinary looking words (strings).
- josephg 2y agoA quick Google search suggests English has about 10 bits of entropy per word. Having a long password like that can still have high total entropy I suppose, but it has a low entropy density.
- kqr 2y agoMaybe 10 bits is the average over the dictionary – which is what matters here, but over normal text it is significantly less. Our best current estimation for relatively high-level text (texts published by the EU) is 6 bits per word[1]. However, as our methods of predicting text improve, this number is revised down. LLMs ought to have made a serious dent in it, but I haven't looked up any newer results. Anyway, all of this to say is that which words are chosen matters, but how they are put together matters perhaps more. [1]: http://arxiv.org/pdf/1606.06996 http://arxiv.org/pdf/1606.06996
- soraminazuki 2y agoThe diceware method is supposed to generate totally random words, so it should fundamentally be unpredictable unless there's a flaw in the source of randomness.
- nvy 2y agoUsername: postgres Password: postgres