5 ms·
The UX of a blacklist with a half billion entries would be so crippling that it would cause a user revolt. Most people's password-selection strategies are simi
by royce 9y ago
The UX of a blacklist with a half billion entries would be so crippling that it would cause a user revolt.
Most people's password-selection strategies are similar enough to other people's (like kbenson's 4000+ hit) that they could spend hours trying to come up with a password that has never been leaked before.
I tried to encourage Troy to suggest to implementors that blacklisting all passwords was a Bad Idea. Instead, he doubled down:
https://twitter.com/TychoTithonus/status/966400790221930496 https://twitter.com/TychoTithonus/status/966400790221930496
Please don't use the entire list for blacklisting unless you actively also guide the user in how to generate a random passphrase (if a human must remember it) or a random password (if it will be stored in a password manager).
Instead:
1. Use a high-level password-strength assessment widget like zxcvbn:
https://github.com/dropbox/zxcvbn https://github.com/dropbox/zxcvbn
2. Configure a blacklist with, say, 10K or 20K of the most common passwords.
3. Hash passwords with bcrypt cost 12 (adjusted to your platform's hashrate capabilities), scrypt, or the appropriate method from the Argon2 family.
But for all that is holy, please don't use Troy's entire corpus - or even the first million - as a blacklist. To quote the old NANOG saw, I encourage all of my competitors to use it. ;)
Edit: And since Troy's API at this writing does not support only querying the top X passwords, there's no way to use the API while avoiding the UX nightmare. So if you want to use this data in a professional manner, but don't want to download the entire corpus, here are the first 20K from his list (of which I've only personally cracked 19965 so far, interestingly; gist will be updated once I get all 20K):
https://gist.github.com/roycewilliams/281ce539915a947a23db17137d91aeb7 https://gist.github.com/roycewilliams/281ce539915a947a23db17...
Edit 2: Preliminary results indicate that this data may be dirty. The 273rd most common password, according to Troy, is '$HEX'. This is almost certainly an import/conversion artifact, since the '$HEX' prefix is how most cracking suites escape non-ASCII or passwords that contain colons. I expect that there will be more artifacts. Use the data with caution.
- kbenson 9y ago> I tried to encourage Troy to suggest to implementors that blacklisting all passwords was a Bad Idea. Instead, he doubled down > Please don't use the entire list for blacklisting unless you actively also guide the user in how to generate a random passphrase (if a human must remember it) or a random password (if it will be stored in a password manager). I think he did the right thing, and think you are correct as well. I think we have the best of both worlds with this, in that it includes the count, so API users can determine what the correct cut-off is for them. Once you get into the thousands (or maybe less) might be a good indicator that your password is not only relatively common, but also likely to be on (and maybe even fairly high on) many dictionary lists. More secure services that cater to more technically savvy users (or security conscious companies) may decide to blacklist any password on the list period, and that may be okay because those sites either trust their users to deal with it or can dictate conditions for a captive audience.
- royce 9y agoAs a corpus to download for password research, this is indeed useful. But for providing a blacklist -- his stated purpose - it is not. The crucial tell: his API does not allow the implementor to specify a frequency threshold (by top X in the list, or by Y number of unique uses of the password or higher). By both API and explicit language in the announcement, he is promulgating the idea that checking the entire blacklist is useful, and "the larger the blacklist, the better." This is exactly what I'm arguing against.
- kbenson 9y ago> his API does not allow the implementor to specify a frequency threshold Yes it does. The output contains the number of matching passwords. It's just client side instead of server side. The reason for not doing so on the server is also obvious taking into account his explanation of cost and caching, which informed much of the API design itself. > By both API and explicit language in the announcement, he is promulgating the idea that checking the entire blacklist is useful Because the entire blacklist is useful. He's given all the relevant information to the client to do with as they may. It's up to them to choose how to utilize it. I'm not sure why you seem to think some narrower use case is necessarily better, using end use cases as arguments, given it's an API and needs a client implementation fore being usable anyway.
- royce 9y agoIt's a fair point that raw password count is available. But that value is an absolute number, without any in-API context of the total size of the corpus. This makes expressing relative rarity only possible by hard-coding the total size of the corpus into a calculation. Put another way: the 20,000th position has a frequency value of "7889". But what does that mean? Where is that in the distribution of password frequency? It's impossible to tell, without manually constructed context that will change over time as the total number of passwords in his corpus expands. But more crucially, there is no way to tell relative rank ("is this password in the top 20k?") using the API that I can see. That would make using the top X much easier. But with the K-anonymity "feature", there's no way to do that that I can see.
- halflings 9y agoThen add NIST to the list of people you should be reaching out to (report linked from the homepage of pwnedpasswords): https://www.nist.gov/itl/tig/projects/special-publication-800-63 https://www.nist.gov/itl/tig/projects/special-publication-80...
- royce 9y agoIndeed. One of the authors of 800-63B is actively involved in the password-research community, and is already aware that the guidance places no restrictions on blacklist size.
- varenc 9y agoDropbox's zxcvbn password-strength estimator already incorporates a list of 250k+ common passwords and words. No need to make a separate blacklist check unless you want to check more than what zxcvbn already does. (and you could enforce it more strongly on the server instead of locally in JS like zxcvbn) For reference, these are the passwords and common words zxcvbn already checks against: https://github.com/dropbox/zxcvbn/tree/master/data https://github.com/dropbox/zxcvbn/tree/master/data And with zxcvbn it'll still flag a password as low entropy/low security if the password it's checking is just a simple modification of something on the common password list.
- always_good 9y agoAll depends on the threat model. Reusing username/email/password can cost your users hundreds per day on a gambling site. And users don't care about a password-gen guide. For example, in that case, you'd want to consider just generating passwords for them. But of course this would be silly for the run of the mill website.
- chatmasta 9y agoThat’s funny because coral.co.uk posts its login over HTTP. In fact if you try to login via https it redirects you to http. It would be fun to setup a hotspot called _The_Cloud outside a Coral and see what you find on the wire!
- Ajedi32 9y agoKeep in mind a lot of these passwords are associated with email addresses in the actual dumps. By allowing a user to use one of these passwords, there's a non-negligible chance you're knowingly allowing them to use a username/password combo that is publicly available, and that any hacker who wanted to compromise their account could do so _on their first try_. Besides, when it comes to passwords... half a billion really isn't all that much. There are (26*2+10)^8 = ~218 _trillion_ possible 8-character alphanumeric passwords. So even if every single one of these half a billion passwords were 8-characters long, you'd still only be disallowing ~0.0002% of them.
- royce 9y agoSome large services do use the actual dumps, and correlate them with the email address associated with the current user, in order to give users a personalized warning that they're reusing a leaked password that's already associated with that specific email address. This is a much different proposition from forbidding that specific user from using a half a billion passwords. A full 80% of the v1 corpus can be avoided by simply requiring a minimum password length of 12. As Troy has pointed out elsewhere, this wouldn't be great UX, either. While it would dramatically increase the chances that they came up with a word that would be A) not in the existing blacklist, and B) harder to attack offline ... it would still be significantly bad UX compared to the best-practice alternative that I lay out in a separate thread branch. But it would still be much better UX than use of the full blacklist.
- caf 9y agoRather than "I am right and Troy is wrong", this seems to be a case of "reasonable people may disagree".
- mysterypie 9y ago> The UX of a blacklist with a half billion entries would be so crippling that it would cause a user revolt. The situation doesn't need to get as bad as you think. If you suggest XKCD's four random common word method[1] to your users as part of the user interface, you'll be fine. As a test, I tried putting together 2 random, common, unrelated words together: - yak elephant -> yakelephant - crowd brown -> crowdbrown - plastic envy -> plasticenvy - colon spanish -> colonspanish - jogging adhesive -> joggingadhesive Even with two words, none of the above appear in the Pwned Passwords database of a half a billion passwords. It's not difficult to choose memorable passwords and avoid entries from Pwned Passwords if you suggest this method to your users (preferably recommending four words, but three might be OK depending on the threat model). [1] https://xkcd.com/936/ https://xkcd.com/936/