7 ms·
This is not a good idea, as implemented. You shouldn't disallow a user from using an otherwise strong password just because it's detected in a breach unless you
by throwawaymath 8y ago
This is not a good idea, as implemented. You shouldn't disallow a user from using an otherwise strong password just because it's detected in a breach unless you can definitively see that it's already associated with their email address or username.
The logical conclusion of a password checking system like this is that this password:
ZBjHWJd$8XbJhY7LQvkmARBW)p7xgiDzDw}iMLLw
can no longer be used by anyone, because I've just "breached it" by posting it on Hacker News.
- bijection 8y agoAny known password is no longer a particularly strong one.
- throwawaymath 8y agoThat doesn't make sense. If I publish a list of 20 trillion alphanumeric passwords, each of which is 20 characters long, your thesis is that no one should ever use any of those passwords again?
- simcop2387 8y agoIf it's published as a list of known passwords, yes. That's roughly 1.50463276905253e-21 percent of the potential passwords for that character space (assuming 64 possible characters). If it's know that those are passwords then they're much much easier to test against than the 1329227995784915872903807060280344576 possibilities.
- throwawaymath 8y agoFirst of all, in an online brute forcing scenario attackers will never get through the entire 20 trillion. Even if they knew for certain that the target victim's password satisfied the precise constraints of the passwords published in that list, they'd need around 2 years (with reasonable assumptions of millisecond latency) of constant, 24/7 attempts to run through them all. This is assuming there is no rate limiting. In an offline brute forcing scenario, either the passwords have been hashed with a strong key derivation function and a randomized salt or they haven't. If they have, it doesn't matter if the password is in that list. If they haven't, you're most likely screwed either way, because attackers can get up to 1 trillion password attempts per second in real world cracking setups now.
- geofft 8y agoNo, you're not screwed either way, because there are trillions of trillions of trillions of possible 20-character passwords. So even if they can get one trillion attempts per second, it will still take an attacker trillions of trillions of seconds to brute-force all possible 20-character passwords, which is longer than the lifetime of the universe.
- throwawaymath 8y agoNow keep in mind that 1) it is possible to choose a strong password with fewer than 20 characters, and 2) most people will not choose a password that long. Extrapolating further, this "policy" would disallow people from choosing passwords below a certain length, because it's theoretically plausible someone has published the set of all possible strings fewer than n characters long.
- deleted 8y ago[deleted]
- geofft 8y ago> Extrapolating further, this "policy" would disallow people from choosing passwords below a certain length, because it's theoretically plausible someone has published the set of all possible strings fewer than n characters long. Yes, that seems like a good password policy. A list of possible alphanumeric strings that is actually reasonable to physically publish (i.e., not 20 trillion) is a list of extremely short alphanumeric strings. 5 alphanumeric characters is 380 million possible strings. 6 is about 2 billion. You should absolutely ban passwords that are 6 characters or shorter! In fact, I would go so far as to say that the questions of "Is this password too short because someone could brute-force all the possibilities, even if we're using a good password hash and a previously-unknown salt" and "Can someone physically enumerate all passwords of this size and put them on Pastebin or otherwise get them in the HIBP database" are equivalent.
- throwawaymath 8y ago
- geofft 8y agoYes, that is the correct conclusion. There are about 0.7 trillion trillion trillion alphanumeric passwords 20 characters long (62^20). Banning 20 trillion of them is a drop in the oceans, and nobody using a password generator is statistically likely to generate them within the expected lifetime of the universe, let alone of any given website. So, if you see one of those passwords, it is overwhelmingly likely that someone took a shortcut and used your list.
- throwawaymath 8y agoNo, it's not the correct conclusion. Virtually no passwords are safe under an offline brute force attempt if they haven't been protected by a robust key derivation function and a randomized salt. If the password has been protected like that, you not only need to try all 20 trillion of those hypothetical passwords; you also need to try them with the correct salt. And this is aside from the fact that you won't even rip through those 20 trillion in an online brute force attempt. Eventually you are only constrained by computational resources. If the passwords are cryptographically secured, it's fine if any one of them is published on the internet if you cannot associate it with any given user. If they're not cryptographically secured, this won't meaningfully reduce your already poor security anyway.
- geofft 8y agoBut what is the harm in blocking all 20 trillion such passwords? Do you expect to have any false positives?
- throwawaymath 8y agoThe harm is the principle of it. You should not design a system that greps through e.g. every single breach dump for arbitrary passwords every single time someone tries to sign up for your service. That's maddeningly inefficient.
- geofft 8y agoWhy is it inefficient? Is the HIBP API too slow? How slow is too slow? My principle is that you should not let people sign up with breached passwords at all - don't make judgment calls about which breaches matter, and whether you think it's the same user or not, or the password is strong enough or not. Just ban the passwords. (Remember that no actual data breach contains 20 trillion passwords.)
- rmtech 8y ago> publish a list of 20 trillion alphanumeric passwords, each of which is 20 characters long, your thesis is that no one should ever use any of those passwords again? No because they each still have a very low probability. You have to be Bayesian about this: a list of one trillion passwords that have no further distinguishing information about each one of them cannot be assigned a probability of > 1/(1 trillion) In a data breach, a given password appears next to a particular username or email, which means it has a very high probability of being the password for that account.
- tialaramex 8y agoYes, it's just safer. We do this _all the time_ in the Web PKI. Remember the "Debian weak keys"? What's weak about those particular keys? Nothing. Nothing whatsoever. Those keys aren't special in any way. Except, Debian shipped releases that always picked one of these key pairs. So anyone with a mind to can go find the list of private keys that corresponds to these particular public keys and thus we don't let you use those public keys any more. (You can go try this if you don't believe me, submit a CSR to your preferred public CA asking for a certificate for one of the Debian Weak Keys, it will be rejected and there may or may not be an explanation attached saying your keys are crap and to get new ones) Whole swathes of keys are blacklisted. ROCA is another example, somebody took one mathematical short-cut too many in their optimised design for RSA key generation, and so the resulting keys all have this very obvious structure that's exploitable (not easily, but enough that a sovereign entity could definitely break them). So we just blacklisted all those keys. If you pick truly random keys you'll never notice this in a lifetime because of statistics, it's just some code on the issuer's systems that you never need to care about.
- echelon 8y agoIf they use an already compromised password, they're prone to a dictionary attack. edit: I did try logging into your account with the password you posted. :P
- stormbrew 8y agoThat said, I know of one event where someone got into a bunch of accounts on a site and did some real damage by using known username password combinations and preemptively forcing specifically those accounts to change their passwords and blocking them from being used would have prevented the outcome that the site invalidated literally every user's password instead.
- rmtech 8y agoIf the password actually has a lot of entropy but it appears in a breach then that's some fairly strong evidence that the user is reusing it. Specifically if it appears n times in the HIBP database you should assign at least roughly 1/n probability that the user is reusing it. So if you assign disutility -V to letting a user have a known username + password combo and utility U to letting a user sign up with a known password but unknown username, the utility is (n-1)/n×U - 1/n×V Reasonable values of U and V for a given site will be different depending on the application, but for online banking -V would be maybe -20 and U might be negative as well. You wanna bank with a public password lol? For something like gmail or Facebook it would be the same story. On the other hand if the password is quite weak then it's vulnerable to credential stuffing. If it appears, say, 10,000 times in the HIBP database then most likely it's as good as public whether or not the user account name is known. Maybe there's a sweet spot around 50 instances where you can't really credential stuff it, and you also aren't that sure that it's a reuse. In terms of usability you could tell the user to change it up a bit, add some words. For example, r0bbiewilliams appears 5 times in the database. luvrobbiewilliams appears 0 times AND IS PROBABLY EASIER TO REMEMBER! You can almost always get away from a breached password by adding a small amount of text.
- throwawaymath 8y ago> If the password actually has a lot of entropy but it appears in a breach then that's some fairly strong evidence that the user is reusing it. I'm not talking about scenarios where you can associate the password with a specific user.
- greglindahl 8y agoMany people use the known passwords list with offline cracking tools.
- rmtech 8y agowell when the password got breached it is associated with a particular user. And HIBP will tell you how many times a given password appears, but not which account it appears with. See: https://haveibeenpwned.com/Passwords https://haveibeenpwned.com/Passwords
- Buttons840 8y agoThe odds of a user randomly selecting that password, rounded to 20 decimal places, is 0.
- throwawaymath 8y agoThat's a fair point. My original problem with this is that you'd likely not have the entire copy of breached passwords available for local comparisons. But Troy Hunt provides that freely so you don't need to use his API. That wouldn't have much latency at all...that fact in combination with the obvious probabilities involved here, yeah I'll concede the point.