22 ms·
This is remarkable. I always find it interesting when bugs like this occur. It reminds me of a hackathon I attended where a food ordering startup (I forget the
by grardb 11y ago
This is remarkable. I always find it interesting when bugs like this occur.
It reminds me of a hackathon I attended where a food ordering startup (I forget the name, but they were chosen to feed us dinner that night) had a similar bug, which baffled me beyond belief. Without going into crazy detail about my password, it typically follows a certain pattern but is never the same across websites. For some reason, the website kept saying my password was invalid. It met all the password requirements that the website asked for (length, capital letter, etc.).
I forget the exact details, but it ended up being the exact location of a capital letter, the location of a number, or some combination of both. I could never figure out how a bug like that could even be coded up. My best guess is that it was some poorly-formed regex.
> Some people, when confronted with a problem, think "I know, I'll use regular expressions." Now they have two problems.
- akerl_ 11y agoIt should be common knowledge at this point, but just in case: If you're doing regex or any other text manipulation on user input when you ask them to set a password, you're doing it wrong.
- chucky_z 11y agoWhat's the best way to deal with this problem, and what's the correct way to deal with this problem?
- akerl_ 11y agoThe input the user puts into the password prompt should by taken exactly as is, pushed into bcrypt/scrypt/etc, then stored as the user's password hash. I'm not entirely opposed to requiring a minimum length, but imposing max lengths / character class rules / etc ends up hurting people who want to pick strong passwords more than it helps people who would pick weak ones (enforcing character classes just gets us lots of password1A! and similar)
- cookiecaper 11y agoI agree that the only real restriction that makes sense is a minimum character count. The others just tend to get in the way. I haven't seen anyone implement it in the wild, but it'd also be cool if there was a wordlist of the 25 most common passwords that the site matched against and refused to accept. I think those two policies, minimum length and no super common passwords, would do a lot to minimize the effectiveness of dictionary attacks.
- mbreese 11y agoI personally like the Stanford policy[1]. It basically boils down to: the longer the password, the fewer restrictions. Each password needs to be at least 8 characters, but if you only have 8 characters, you might need uppercase, lowercase, a symbol and a digit. If you have 12 characters, you only need upper/lower/digit. Once you hit 20 characters, you can have whatever you want. I think that this is a good balance between security for short passwords, while still allowing ridiculously long ones (pass-phrases). [1] http://arstechnica.com/security/2014/04/stanfords-password-policy-shuns-one-size-fits-all-security/ http://arstechnica.com/security/2014/04/stanfords-password-p...
- warfangle 11y agoAlmost like you're validating on entropy and not specific rules..........
- Dylan16807 11y agoThe concept is nice, but their numbers are horribly, unforgivably wrong. 8 random upper/lower/digit/symbol characters are equal to 9 random mixed-case letters. Not 16. Their cutoffs for different mixes are 8, 12, 16, 20. Realistic cutoffs would look more like 10, 11, 11, 14. Even worse, they encourage counting the individual letters in words. Never do that. Random words are only as good as two random characters.
- technion 11y ago$ dd if=/dev/urandom bs=32 count=1 | xxd -p That nearly every site on the Internet will refer to that output as "not strong enough" and instead suggest P@ssword1 as a better alternative definitely speaks to the issue.
- bjt 11y agoUse https://github.com/dropbox/zxcvbn https://github.com/dropbox/zxcvbn.
- lmm 11y agoThe correct way is not to use passwords. Use X.509 client certificates, and let the user secure theirs whatever way makes sense to them (whether that's a password, a smartcard, both, or something else). Unfortunately the browser UX for them is terrible.
- 0xcde4c3db 11y agoBesides the usual regex aches and pains, the grammar for email addresses is far more complex than most people realize. According to a highly-voted Stack Overflow answer [1], the current RFC-specified grammar for addresses can't even be matched with regex alone. Combining the edge cases of the grammar with (say) Unicode normalization sounds like a recipe for hours of fun. [1] https://stackoverflow.com/questions/201323/using-a-regular-expression-to-validate-an-email-address https://stackoverflow.com/questions/201323/using-a-regular-e...
- kccqzy 11y agoI find that quite unbelievable. When I had a similar problem last year, the first resource I found was a W3C specification[1] about <input type=email>. The specification clearly states that email addresses should match: /^[a-zA-Z0-9.!#$%&’*+/=?^_`{|}~-]+@[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)*$/ Since this is an official W3C doc, I see no reason why people shouldn't use this. Edit: There is also a version by WHATWG[2] here: /^[a-zA-Z0-9.!#$%&'*+\/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/ It apparently does a more thorough validation than the W3C one in the domain part, but the difference between the two is not apparent in practice. [1]: http://www.w3.org/TR/html-markup/input.email.html http://www.w3.org/TR/html-markup/input.email.html [2]: https://html.spec.whatwg.org/multipage/forms.html#e-mail-state-(type=email) https://html.spec.whatwg.org/multipage/forms.html#e-mail-sta...
- wyldfire 11y ago>Since this is an official W3C doc, I see no reason why people shouldn't use this. Well, as far as authority/canon goes, it's typically dictated by the IETF and not W3. And, sure, they could cooperate with one another but, really -- if there's an authority on the protocols that describe email (SMTP, POP, IMAP, etc) -- it shouldn't be the World Wide Web Consortium. That said, the IETF does tend to draft RFCs that reflect actual implementations (at least their intended design), but since they often bias towards interoperability, it's unlikely they'd narrow the scope of the email address grammar.
- oinksoft 11y ago
- cookiecaper 11y agoI find a lot of websites will eat passwords that contain special characters. I don't mean that they'll tell you it doesn't match the password policy, I mean that they'll accept the password and then tell you the password is wrong when you come back to sign in. I eventually had to teach my password generator to use only a few usually-properly-handled special characters when generating to avoid the hassle of having to reset the password every time. The same thing is often true of long passwords -- websites will accept the password at the UI level, but it probably gets truncated somewhere in the processing, and you don't know which character you got cut off at, so you have to reset to something shorter.
- ars 11y agoIf it's going to eat a character it should at least be consistent and also eat it at the validation stage.
- Animats 11y ago"Some people, when confronted with a problem, think "I know, I'll use regular expressions." Now they have two problems." Yes. My favorite was a Coyote Point load balancer bug. If the last character of the HTTP options is "m", the connection will not get past the load balancer.[1] I found this because a web crawler was having trouble with one site. Fortunately, I knew someone with their own Coyote Point load balancer, and was able to establish that the connection went into the load balancer and never came out. The load balancer has a big file of rules which contain regular expressions. Somewhere, I think there's a "\m" where they meant "\n". Reporting this to the vendor, along with a Python program to demonstrate the problem, was of course futile; they suggested "upgrading the software". I demonstrated that the bug existed on their own load balancer on their own site. I finally added a completely useless field to the HTTP header so that the last character was not "m". [1] https://www.webmasterworld.com/webmaster_hardware/3312997.htm https://www.webmasterworld.com/webmaster_hardware/3312997.ht...
- chris_wot 11y agoTell them they've violated their contract and you are migrating from them ASAP.
- rdancer 11y agoThat doesn't work when it's software on your client's computer. Remember when all the web devs revolted and ditched IE6 back in 2001, because MS was just taking the Mickey? Yeah, me neither.
- Animats 11y agoRight. I didn't have a Coyote Point unit. Many other sites did, and they all appeared to be down to my web crawler until I figured out the problem. (Current web crawler problem: sites that won't let you read their robots.txt file if they don't like your user-agent string.)
- ThisIs_MyName 11y ago
- raverbashing 11y agoFrom what I've seen from "regular people" writing regular expressions, they seem to not have the slightest clue on how to do it And then putting it into the program without testing it properly So, sorry, the issue is not regexes, but people just going for it at an "trial and error" fashion (and sometimes just trial)
- TeMPOraL 11y agoFor people here that may ask themselves just how exactly one could test regular expressions, I recommend a visual tool like [0]. Having the regex structure and meaning drawn in front of you helps tremendously. That, and for the love of God please comment any non-trivial regexp. Either like this, with 'x' option: preg_match('/^ .* # Match any number of characters... (?=.{6,}) # ... AND match at least 6 characters (lookahead) ... (?=.*\d) # ... AND match one digit after any number of characters (lookahead) ... (?=.*[a-zA-Z]) # ... AND match one letter after any number of characters (lookahead) ... .* # ... AND allow any number of characters later. $/x', $password); ... or just with normal programming language comments and stitching regexp from multiple strings in multiple lines. Also give some semantic meaning to groups if you use them, e.g. tag them with constants so that your code isn't full of stuff "result.get(3)", which makes you waste time on trying to recall what was that group 3 in the code from last month. I know it's pretty much software engineering 101. It's the basics of basics. But from my experience, even the brightest of engineers in most serious projects suddenly forget how to write code when they touch regular expressions. [0] - https://www.debuggex.com https://www.debuggex.com
- raverbashing 11y agoUsing tools is good, but also people should test them in their unit test, for what it should/should not be accepted If you're forbidding items beginning with numbers, just have a test try passing '1a' and failing the match
- TeMPOraL 11y ago
- noonespecial 11y agoRegex's in perl5 are what introduced me to test automation. Hard.
- annnnd 11y ago> Some people, when confronted with a problem, think "I know, I'll use regular expressions." Now they have two problems. Great quote. However, I think regexes got a bad reputation just because of the way people use them. In essence they are a pretty reliable way of parsing because the parsing engine is well tested. But the expression should be kept as simple as possible and developer should avoid using any nonstandard / nonexplicit extensions. I even avoid using \w because, well, what IS a word character? I am sure it is defined somewhere... but I'll always use explicit form (like "[a-zA-Z]" when I want ASCII chars) instead. Anyway, if you use the form as used in the regex puzzle [0], you'll be fine. As long as you use regex only for what it was meant for, of course [1]... [0] https://news.ycombinator.com/item?id=10787509 https://news.ycombinator.com/item?id=10787509 [1] http://blog.codinghorror.com/parsing-html-the-cthulhu-way/ http://blog.codinghorror.com/parsing-html-the-cthulhu-way/
- ifdefdebug 11y ago> Some people, when confronted with a problem, think "I know, I'll use regular expressions." Now they have two problems. Yeah sure. I think I heard that quote before... just about a million times? People expect regex to be an easy-to-use tool. Well it's not, and it's a foot gun if you don't take your time to learn it right. But no, people hack up some expressions, hit their feet and blame... the tool of course, not themselves. Just learn it right, it's a great tool if you know how to make it work for you :)
- pg_is_a_butt 11y agoyeah, very remarkable... only 1 of your contacts is fucked up in google contacts? and you were able to figure out which one? whoa. remarkable. at least 20 of mine are like this, and i swear it isn't the same 20. google creates terrible software that breaks with every other release. deal with it.