4 ms·
Claiming that regular expressions are too terse is a bit much. There are only three (!) fundamental operators in basic regular expressions (four if you include
by jmts 8y ago
Claiming that regular expressions are too terse is a bit much. There are only three (!) fundamental operators in basic regular expressions (four if you include parenthesis), with all other non-language specific operators being derived from that (ignoring precendence rules):
1. concatenation, to append regex A or regex B: AB
2. alternation, to select between A or B: A | B
3. kleene star, to repeat A zero or more times: A*
4. parenthesis allows specification of a sub-expression: (A)
The following are all derived/syntactic sugar:
[ABCD] -> (A | B | C | D)
A+ -> AA*
A{2} -> AA
A{2,4} -> AA(|A|AA) or A(A|AA|AAA)
A? -> (A|)
Just about everything else is implementation specific (if choice of special characters and available operators isn't already). That means you either need to be using the features regularly to remember them, or you have to look them up anyway.
Regular expressions are terse not because they are badly designed, but because by definition the description of regular languages is inherently minimal. It is part of their beauty. Without this minimalism, every tidy little one liner we have to perform some simple match becomes a multi-line specification in Backus-Naur form.
The world needs to get over this fear of regular expressions from ignorance and continued misinformation. They are not magic or impossible to understand. They are an elegant description of a very simple state machine which steps through a string one character at a time, nothing more.
Edit: corrected derivation of A{2,4} a la Twisol and
jbnicolai.
- Twisol 8y agoMinor nitpick: A{2,4} becomes AA(|A|AA), not AA(A|AA).
- jrochkind1 8y ago`AA(|A|AA)` is valid regexp syntax? Holy moly it is! I had no idea you could put a | right after parens like that! I guess (|A) is nothing or A? Woah.
- Twisol 8y agoYep! It's equivalent to `AA(A|AA)?`, by way of the parent's reduction for `?`, but it's a nice option when you want to emphasize some kind of symmetry.
- jmts 8y agoWell spotted. Thanks for providing the correct derivation.
- jbnicolai 8y agoAgreed. The problem is the risk of small, relatively hard to spot & nearly impossible to properly debug mistakes. > A{2,4} -> AA(A|AA)
- Twisol 8y agoAnd this is why we have the syntax sugar.
- slavik81 8y agoThat introduces problems too. If you try to use sugar like '+' with an implementation that doesn't support it, you don't get any sort of error. Instead you get a different expression. Unfortunately, there's an inherent tradeoff between encoding efficiency and error detection. Notice that with the VerbalExpressions it would be trivial to return a useful error message if the 'at_least_one' pattern did not exist.
- b2gills 8y agoPerl 6 regexes attempt improve upon this situation by making regexes more like a regular programming language. That is it errs on the side of error detection rather than encoding efficiency. (It also adds features that would be difficult to add to Perl 5/PCRE regex design) For a start if it didn't support using `+`, then any attempt to use it would generate a compiler error because it is not alphanumeric. (regex is code in Perl 6) All non-alphanumeric characters are presumed to be metasyntactic, and so must be escaped in some way to match literally. Arguably best way is to quote it like a string literal. (Uses the same domain specific sub-language that the main language uses for string literals) / "+" + / # at least one + character It really is a significant redesign. /A{2,4}/ # Perl 5/PCRE /A ** 2..4/ # Perl 6 /A (?:BA){1,3}/x /A [BA] ** 1..3/ # Perl 6: direct translation /A ** 2..4 % B/ # Perl 6: 2 to 4 A's separated by B /A (?:BA){1,3} B?/x /A ** 2..4 %% B/ # Perl 6: %% allows trailing separator /\" [^"]* \"/x # Perl 5/PCRE /\" <-["]>* \"/ # Perl 6: direct translation /「"」 ~ 「"」 <-["]>*/ # Perl 6: between two ", match anything else # (can be used to generate better error messages) --- # Perl 5 my $foo = qr/foo/; 'abfoo' =~ /ab $foo/x; # Perl 6 my $foo = /foo/; 'abfoo' ~~ /ab <$foo>/; # or my token foo {foo} # treat it as a lexical subroutine 'abfoo' ~~ /ab <&foo>/; --- # Perl 5 my $foo = 'foo'; 'abfoo' =~ /ab \Q $foo \E/x; # treat as string not regex # Perl 6 my $foo = 'foo'; 'abfoo' ~~ /ab $foo/; # that is the default in Perl 6
- shafte 8y agoTo be fair, you are not including more advanced operators, like positive/negative lookahead/behind (which is the specific example the article uses), capturing and non-capturing groups, greedy vs non-greedy kleene stars, etc. As you say, they are implementation specific, but that's part of the problem: the basic regular expression syntax is insufficient for many tasks, so people take to extending it in complicated and syntactically opaque ways. That's the sign of a bad DSL, not a good one.
- jmts 8y agoMaybe I'm arguing semantics here. To be clear, my point is that I do not agree it is reasonable to declare that regular expressions are a bad DSL simply because it is possible (however common) for people to write difficult to read, or difficult to understand regular expressions. It is the responsibility of the author of the expression to ensure that it is readable and understandable - to the extent that they should exercise restraint when possible use of an available feature would hinder readability and understandability. There is absolutely no need for the example regex of http-like strings to be written the way that it is - there is only the want of the author, because they have a hammer and they are looking for a nail. If anything, using a regex for such a thing sets a bad precedent because anybody who wishes to come along and add user@password support to it is going to extend it and make it worse. A more understandable way to process such a string would be to split it into constituent parts and use regex only for validation. Split at the :// for schema, split at the next / for path, etc. Turn these into functions, and keep the regexes simple. Regular expressions are notorious because they are abused, not because they are evil.
- repsilat 8y ago> A more understandable way to process such a string would be to split it into constituent parts and use regex only for validation Split the regex, or split the URL itself? I kinda think the split regex is pretty reasonable: protocol = "[a-z]{3,10}://" domain = "([^/?#]*)" path = "([^?#]*)" query = "(?:\?([^#]*))?" fragment = "(?:#(.*))?" url = protocol + domain + path + query + fragment Not so terse now though, probably has to be wrapped in a function now (or stored as a constant somewhere else.)
- mirimir 8y agoWhat drove me crazy at first was that any character can be used as a delimiter. That's very useful, I admit. But it complicates understanding examples found through searching ;)