6 ms·
Yes, it is, without a doubt. It's one of the most universal tricks of the trade that you'll literally never regret learning, mainly because just about any envir
by bermanoid 14y ago
Yes, it is, without a doubt. It's one of the most universal tricks of the trade that you'll literally never regret learning, mainly because just about any environment you'll ever work in will by necessity support regexes, and in many, it will be the primary way you interact with text.
But it's also a must that you realize that despite the fact that they're exceptionally useful and widely supported, regexes are a disgusting abomination, one that we should be absolutely mortified to be associated with. It's one of the worst syntaxes to ever be invented, and every one of us should feel the cold stink of UX failure wash over us every time we write a regex. If we ever catch ourselves writing a DSL that in any way, shape, or form resembles regular expressions, we should stop immediately and ask what the fuck is wrong with us and why we're being so opaque and random. Regular expressions are quite literally one of the worst syntaxes to ever be introduced in our field.
I worry a lot when someone doesn't know regular expressions at all. But I worry far more when someone thinks they're beautiful. That person has far too high a tolerance for unintuitive syntax and code, and will cause vastly more damage to my codebase than even the rank amateur that still uses "goto" on a regular basis.
Which is not to say that we shouldn't lean on regexes heavily anyways when they're appropriate - as programmers our primary job description is that we're paid good money to work with shitty interfaces in order to express simple ideas and algorithms.
- itmag 14y agoI could imagine a Jquery-like DSL where something like /^[A-Z]+[0-9]{2}$/ could be expressed like Match(str).BeginsWith().BigCaseAlpha().AtLeastOne().Numeric().FixedLength(2).EndsWith() Of course, you would also have to be able to nest these for more advanced matching... Is there something like this already in existence? :)
- eurleif 14y agoHow about pyparsing? http://pyparsing.wikispaces.com/ http://pyparsing.wikispaces.com/
- grovulent 14y agoHmm - you know I actually think it's not correct to look at regex as an interface - even though we use it as one. It's really more accurate to look at it as a grammar (type 3 if I remember correctly). Anything that comes out of the whole chomskian hierarchy stuff isn't going to look intuitive. But the point is that it is a particular, very rigorously defined system of representation. And various systems of representation are always more or less intuitively accessible - and come with a whole set of trade offs around what they can represent vs their ease of use etc... These things just are - written into the laws of the world. They are discovered - not invented. We pick them up and use them as we would rocks left lying around. We find better ones when we can and fashion better ones when we can too...
- arnsholt 14y ago> Anything that comes out of the whole chomskian hierarchy stuff isn't going to look intuitive. Why not? Type-0 (recursively enumerable) languages are equivalent to Turing machines, and we've managed to invent some pretty good syntaxes for that. The main problem with regexes really is the syntax. Regular languages are a lot easier to understand (IMO) if you look at the left/right-linear grammars that define the same language as the regex. Regex syntax as a representation is very close to the FSA used for matching, and that's not necessarily the representation best suited for human consumption.
- praptak 14y ago> These things just are - written into the laws of the world. They are discovered - not invented. This applies only to the theoretical computer science regexes. The practical regexes are very different in this aspect. Practical regexes are neither discovered nor invented - they are constructed. The theoretical regexes are just the basis but on top of that there's a lot of features added. Some of them are just syntactic sugar for theoretical regexes but others actually make the language non-regular. Groups that match what previous named groups have matched are definitely in this category.
- Karunamon 14y ago>It's one of the worst syntaxes to ever be invented, and every one of us should feel the cold stink of UX failure wash over us every time we write a regex. Okay, so they're ugly and hard to read at first glance. Considering the purpose of a regex, I can't think of another way to implement them that doesn't involve typing more characters needlessly (therefore even making it more hard to comprehend).
- padolsey 14y agoHow would you design a regular expression syntax more intuitively? Personally, I find beauty and simplicity in regular expressions. Sure, they can grow to hideous atrocities, but you can achieve such disastrous feats with any language/syntax. Maybe you could back up your claim of regexes being a disgusting abomination with, at the very least, anecdotal evidence.
- yen223 14y agoA typical regex looks like this: \b[A-Z0-9._%-]+@[A-Z0-9.-]+\.[A-Z]{2,4}\b Which is also what happens when a cat walks across the keyboard.
- Argorak 14y agoCall me weird, but I find that very readable.
- Auguste 14y agoI find that perfectly readable, except for the \b which I hadn't seen before. It's matching an all-uppercase email address.
- subsystem 14y agoThere's plenty of upper case e-mail addresses that won't match that expression.
- aidos 14y agoWell, if you want to be fully compliant you can go for this 6kb monster http://ex-parrot.com/~pdw/Mail-RFC822-Address.html http://ex-parrot.com/~pdw/Mail-RFC822-Address.html
- InclinedPlane 14y ago/i
- InclinedPlane 14y agoPlease name 2 concrete ways to improve the "interface" of a standard regex.
- qlkzy 14y agoI think that regular expressions have a good syntax for most typical (small) uses of regular expressions. Maybe the choice of special characters isn't ideal, and it might be nice to have 'English' versions of more special characters (similar to e.g. [[:digit:]] in POSIX), but for small regexes the (mostly) one-to-one mapping between characters in the pattern and characters in the string is a very nice and intuitive syntax. I think the real problem is that we lack (or don't learn) good tools to bridge the gap between regular expressions and 'custom parser'. We're reluctant to refactor from '1 line of just-starting-to-be-horrible regex' to tens or hundreds of lines (depending on language and libraries) to do it 'properly', and so we end up stretching regular expressions beyond the point where they make life easier. Perl has Parse::RecDescent (and probably several others), which is pretty close to the right thing, and clearly it's very doable in a lot of languages - anyone got any suggestions in other languages?