4 ms·
Nice idea! Can anyone point the data. May be we can try RNN to generate the passwords.
by matrix2596 10y ago
Nice idea! Can anyone point the data. May be we can try RNN to generate the passwords.
- Faizann20 10y agoYes, we definitely can. The reason I didn't do this was because I did not have enough cpu/gpu power. The results will be better if RNN is applied.
- matrix2596 10y agoi have a gpu. may be i can help. let me know or point me where the whole data is available. thanks
- Faizann20 10y agoThe link is given.
- Faizann20 10y agoHere's the data. https://muslimmatch.thecthulhu.com/ https://muslimmatch.thecthulhu.com/
- matrix2596 10y agothanks
- gwern 10y agoMarkov chains can do amazing things in password cracking: https://arstechnica.com/security/2013/05/how-crackers-make-minced-meat-out-of-your-passwords/ https://arstechnica.com/security/2013/05/how-crackers-make-m... But an RNN isn't necessarily going to help as much as you think. An RNN has two problems compared to a Markov chain: 1. Markov chains memorize strings very very easily, accurately, and scalably; it's easy to memorize phrases, words, suffixes, and prefixes from the existing corpuses of billions of passwords. That's all a Markov chain does, memorize & count. On the other hand, an RNN will struggle to do so because there's no 'place' for it to put all of that, everything has to be encoded into the fixed set of neural net weights, otherwise, it just doesn't know about it; and the more you ask it to learn, the more the competing demands fight each other. RNNs augmented with external memories might help fix this but are still cutting edge research. 2. Markov chains are also very fast, far faster than an RNN. Multiple orders of magnitude difference are possible, unless you use a RNN so small as to be irrelevant (since then it can't memorize anything). For cracking hashes, a small gain in plausibility of guesses is not worth being able to make hundreds or thousands times fewer guesses (unless perhaps the hash are something proper like bcrypt/scrypt where it takes seconds to check, in which case the guessing phase takes up a much smaller fraction of runtime and better guesses may be worthwhile).
- matrix2596 10y agoNice explanation. Thanks Its more from a theoretical point of view. I want to try similar (conditioned on user info) to https://github.com/thoppe/5baa61e4c9b93f3f0682250b6cf8331b7ee68fd8 https://github.com/thoppe/5baa61e4c9b93f3f0682250b6cf8331b7e...
- ma2rten 10y agoActually RNNs/LSTMs are surprisingly good at memorizing in addition to generalization. Have a look at the famous blog post "The Unreasonable Effectiveness of Recurrent Neural Networks" [1] for instance and notice how many words it's able to generate from characters. However, your second point is valid. [1] http://karpathy.github.io/2015/05/21/rnn-effectiveness/ http://karpathy.github.io/2015/05/21/rnn-effectiveness/
- gwern 10y agoThat's not 'surprisingly' good, that's effective only on a small corpus. You can see for yourself that if you try to feed it multiple corpuses with many vocabulary words or proper names, which a Markov chain wouldn't break a sweat on memorizing them all, the RNN has limited memorization ability and what tends to happen is that the less common ones get overwritten in favor of the general grammar of English and the vocabulary of the largest corpus: https://www.gwern.net/RNN%20metadata https://www.gwern.net/RNN%20metadata
- ma2rten 10y agoIf I understand you correctly you are comparing Markov Chains on words to a RNN over characters. That's not fair. This paper shows a large LSTM outperform n-gram models: "In this paper we have shown that RNN LMs can be trained on large amounts of data, and outperform competing models including carefully tuned N-grams. [...] Unlike previous work, we do not require to interpolate both the RNN LM and the N-gram, and the gains of doing so are rather marginal." https://arxiv.org/abs/1602.02410 https://arxiv.org/abs/1602.02410
- 10y ago