5 ms·
As a corpus to download for password research, this is indeed useful. But for providing a blacklist -- his stated purpose - it is not. The crucial tell: his AP
by royce 9y ago
As a corpus to download for password research, this is indeed useful. But for providing a blacklist -- his stated purpose - it is not.
The crucial tell: his API does not allow the implementor to specify a frequency threshold (by top X in the list, or by Y number of unique uses of the password or higher).
By both API and explicit language in the announcement, he is promulgating the idea that checking the entire blacklist is useful, and "the larger the blacklist, the better." This is exactly what I'm arguing against.
- kbenson 9y ago> his API does not allow the implementor to specify a frequency threshold Yes it does. The output contains the number of matching passwords. It's just client side instead of server side. The reason for not doing so on the server is also obvious taking into account his explanation of cost and caching, which informed much of the API design itself. > By both API and explicit language in the announcement, he is promulgating the idea that checking the entire blacklist is useful Because the entire blacklist is useful. He's given all the relevant information to the client to do with as they may. It's up to them to choose how to utilize it. I'm not sure why you seem to think some narrower use case is necessarily better, using end use cases as arguments, given it's an API and needs a client implementation fore being usable anyway.
- royce 9y agoIt's a fair point that raw password count is available. But that value is an absolute number, without any in-API context of the total size of the corpus. This makes expressing relative rarity only possible by hard-coding the total size of the corpus into a calculation. Put another way: the 20,000th position has a frequency value of "7889". But what does that mean? Where is that in the distribution of password frequency? It's impossible to tell, without manually constructed context that will change over time as the total number of passwords in his corpus expands. But more crucially, there is no way to tell relative rank ("is this password in the top 20k?") using the API that I can see. That would make using the top X much easier. But with the K-anonymity "feature", there's no way to do that that I can see.
- jasonpeacock 9y agoI don't follow - how is the relative rarity better than absolute frequency? What really matters is how common your password is - not how highly it's ranked in a compromised password list, which has no relevance to how common it may be. You want to filter on users choosing a password that's been re-used across all compromised more than N times. Filtering users on choosing a password that ranks N of M on a list of compromised passwords doesn't tell the user how bad that password is. In fact, once you get to the rail, the ranking is basically based on sort order and become irrelevant?
- royce 9y agoThe ranking in Troy's list is based entirely on how common the words are. Here are the top 10, with their relative frequency: c4a8d09ca3762af61e59520943dc26494f8941b:123456 (20760336) f7c3bc1d808e04732adf679965ccc34ca7ae3441:123456789 (7016669) b1b3773a05c0ed0176787a4f1574ff0075f7521e:qwerty (3599486) 5baa61e4c9b93f3f0682250b6cf8331b7ee68fd8:password (3303003) 3d4f2bf07dc1be38b20cd6e46949a1071f9d0e3d:111111 (2900049) 7c222fb2927d828af22f592134e8932480637c0d:12345678 (2680521) 6367c48dd193d56ea7b0baad25b19455e529f5ee:abc123 (2670319) e38ad214943daad1d64c102faec29de4afe9da3d:password1 (2310111) 20eabe5d64b0e216796e834f52d61fd0b70332fc:1234567 (2298084) 8cb2237d0679ca88db6464eac60da96345513964:12345 (2088998) So ... what is the "right" threshold for N? $ for topx in 1 100 1000 5000 10000 20000 50000 100000 200000 500000 1000000; do \ echo -n "$topx: "; head -n ${topx} pwned-passwords-2.0.txt | tail -1; done 1: 7C4A8D09CA3762AF61E59520943DC26494F8941B:20760336 100: 482FA19D5C487CB69ACDA19EEE861CC69D82CC94:272371 1000: 5B9FE558F673D63309BEB13BFA5DA6C30A3CA1BF:64912 5000: FE648FC459A6F6EF6CD347BEE3D494766239BBB5:19860 10000: 2682A3DBA7A1452EE7EE9980F195C6A768055DA6:11055 20000: 53490A3C8567342B57B6A4FF24908DF73182B357:6309 50000: 7517CD23A308BBCD05E5AD24AA6AD054237ED470:3153 100000: BA6D6A41B9548C523833627A8B0E5170558BE1EA:1752 200000: E50E6893264519636E90E95B6B1A85D0A691E0B1:931 500000: AF8DF653177BBB3FEE2DA68D314B94CB5281B4F3:381 1000000: BDD57A4CAA691A3441C1190C6F087B58B2EE3EF6:186 2000000: C824AF24AA8F2FD99AD6842DC0E4B49100D96161:93 10000000: 352DB7177AB7848DF1C102234401097FE40EB87D:22 The third field indicates how common the password is in the corpus (for example, the single most common password - "123456" - appears in the corpus 20,760,366 times). So ... based on this data ... what is a reasonable value for that count, such that if the value is exceeded, the user should be disallowed from using the password? How much real-world online or offline resistance is provided by disallowing, say, passwords used at least 186 times in the corpus (roughly a million passwords, though 5201 passwords are at the 186 mark)? (The answer should be self-evident; if it isn't, I can provide more background). Put another way ... if the corpus was only 1M in size, those right-hand values would be much smaller. How could you determine the threshold then? What I'm trying to illustrate here is that it's not the absolute value of that commonality number that matters; it's the relative rank. But that relative rank can't be determined via the API; you must analyze the entire corpus directly - and then discard the vast majority of it for blacklisting purposes. I totally get that the threshold might vary per implementation. But it varies much less once the hash is slow enough, and the authentication service is suitable rate-limited. In other words, any system that would get real benefit from a 1-million-word blacklist is one that needs to be improved elsewhere instead. But Troy didn't provide any guidance about that, or even how to judge for yourself what the threshold might be. He just provided an API to blacklist a corpus of passwords that is three orders of magnitude larger than a properly designed system would ever need. 1. https://blogs.dropbox.com/tech/2012/04/zxcvbn-realistic-password-strength-estimation/ https://blogs.dropbox.com/tech/2012/04/zxcvbn-realistic-pass...
- rspeer 9y agoThis list is small for the purposes of password-cracking. Enumerating 500 million things is something a computer can do very quickly. Consider this: if you store the hash of one of these passwords in your login database, you have stored something that can quickly be turned back into the plaintext password, just by enumerating the list. Passwords that have been leaked can't become good passwords again.
- prewett 9y ago> Passwords that have been leaked can't become good passwords again. This doesn't make any sense. We already know all possible passwords: the set of all the permutations of the set of legal password characters of a given length. Your argument applies equally well to these. You can't just use these passwords and check the hash, since the database (hopefully) at least salted their hash to prevent rainbow attacks like that. But even if they didn't salt their passwords, nothing prevents you from hashing all possible passwords and checking. I think I read somewhere that making a rainbow table for all possible 8 character passwords took little time and space; thus all 8 character passwords are already broken. Need to crack a password? Generate the rainbow table and just look it up. However, your assumption that computers can check passwords quickly is not necessarily correct: that's why bcrypt exists.
- rspeer 9y agoOkay, I may be overstating just how quick it is -- you might have to spend a few days of CPU time to go through the whole list, based on estimates I'm seeing. (Parallelize it however you want.) But don't act like I don't know there are a finite number of passwords. The number of possible passwords grows exponentially with length, and the number of leaked passwords grows linearly with leaks. It's about the number of possible passwords you have to check. There's a huge difference in magnitude between having to hash "all 10-character passwords" and "all 10-character passwords that are definitely someone's actual password".
- tialaramex 9y agoNo. Let me illustrate with a different security situation where we use blacklisting. Many years ago Debian mistakenly shipped a version of openssl that didn't use good entropy to pick RSA keys. As a result, everybody with that Debian would get one from a relatively small pool of keys when they asked for a new one. There's nothing special about these RSA keys, other than the fact that Debian systems from a particular era would always pick them. A good CA (e.g. Let's Encrypt) blacklists the public halves of those key pairs. Again, there's nothing special about them, no reason they're worse than any other random key _except_ Debian always picked those, and since it did bad guys can trivially find out the corresponding _private_ key for each value and so they're useless. If you propose to use one of these blacklisted public keys, there is a near certainty that it's because you have a broken Debian system making the keys, and so refusing you keeps you safe. Even though there's nothing special about these keys. Now, if I have a system that generates RSA keys in a known secure way, I needn't check for Debian weak keys myself. Why not? Because there is statistically no chance I'd ever pick one at random, it's a total waste of engineering effort to check. But if I ask somebody _else_ to make a key pair and send me the public half, I should check against the Debian weak keys, because I shouldn't trust that they're smart enough not to use the broken code. These passwords are crap. They wouldn't necessarily be crap if nobody had ever known what they were, but now they do, so they're crap now. Pick a different password.