5 ms·
If you want readable regexp, just use combinators and your language's variable declaration facilities. No need for more. I don't understand why people still in
by Drup 10y ago
If you want readable regexp, just use combinators and your language's variable declaration facilities. No need for more.
I don't understand why people still insist on using insane syntax for regexps instead of just ... functions (`rep` for repetition, `seq` for sequences, `opt` for optional ..).
- rmetzler 10y agoYes, this would work in most languages. I also like CoffeeScript's multi-line Regexes where you're able to add comments. Example: http://elijahmanor.com/regular-expressions-in-coffeescript-are-awesome/ http://elijahmanor.com/regular-expressions-in-coffeescript-a...
- JadeNB 10y agoLike almost everything to do with regexes, Perl did it first: http://perldoc.perl.org/perlre.html#%2fx http://perldoc.perl.org/perlre.html#%2fx .
- pygy_ 10y agoI made a JS lib just for that: https://github.com/pygy/compose-regexp.js https://github.com/pygy/compose-regexp.js Edit: real life example, a minimal lexer used to split compound CSS selectors on comas. We must skip those in strings, comments and in `:not(a, b)` pseudo-classes, so `.split(',')` doesn't cut it. var selectorTokenizer = flags('g', either( /[(),]/, sequence( '"', greedy('*', either( /\\./, /[^"\n]/ ) ), '"' ), sequence( "'", greedy('*', either( /\\./, /[^'\n]/ ) ), "'" ), sequence( '/*', /[\s\S]*?/, '*/' ) ) ) which compiles to selectorTokenizer = /[(),]|"(?:\\.|[^"\n])*"|'(?:\\.|[^'\n])*'|\/\*[\s\S]*?\*\//g Which is then used as follows function splitSelector(selector) { var indices = [], res = [], inParen = 0, match while (match = selectorTokenizer.exec(selector)) { switch (match[0]) { case '(': inParen++; break case ')': inParen--; break case ',': if (inParen) break; indices.push(match.index) } } for (var i = indices.length; i--;){ res.unshift(selector.slice(indices[i] + 1)) selector = selector.slice(0, indices[i]) } res.unshift(selector) return res }
- benley 10y agoI appreciate your library, that's pretty cool. I just want to point out that some (most?) regex libraries support a whitespace-insensitive mode, which allows you to write out the raw regex in a way that's considerably easier for humans to visually grok: (?x) [(),] | "(?: \\. | [^"\n] )*" | '(?: \\. | [^'\n] )*' | \/ \* [\s\S]*? \* \/ That works in python, at least; I don't really know javascript so I can't speak to that.
- TeMPOraL 10y agoIn Java I prefer to do something like this: "[(),]" // match the foo part... + "|\"(?:\\.|[^\"\n])*\"" //... or, match the bar part in *double* quotes, putting quoted value in capture group 1... + "|'(?:\\.|[^'\\n])*'" //... or, match the bar part in *single* quotes, putting quoted value in capture group 1... + "|\/\*[\s\S]*?\*\" //... or, match whatever the hell that is. Simple string splitting + commenting the semantic parts. Also, labeling the capture group (and creating named constants for them in your code next to your regex) is a huge win.
- pygy_ 10y agoIf the string you happen to match contains a lot of metacharacters, you end up with backslashes all over the place, which makes the result hard to read. Nested groups and captures are also often hard to parse. FWIW, you forgot to double escape `\\\\.`, and didn't close the CSS comment (last alternative). "[(),]" // match the foo part... + "|\"(?:\\\\.|[^\"\\n])*\"" //... or, match the bar part in *double* quotes + "|'(?:\\\\.|[^'\\n])*'" //... or, match the bar part in *single* quotes + "|\/\*[\s\S]*?\*\/" // or match the comment Also, you're probably not familiar with the quirks of JS regexps, but the two string alternatives use non-capturing groups, and `[\s\S]` is the true "any" matcher, `.` doesn't match new lines. At last, `*?` is a non-greedy `*`. (Edited thrice, damn you italics).
- TeMPOraL 10y agoThanks for the clarifications! I admit I just copied your example and tried to sort-of convert it into Java style. It definitely won't be a correct Java-compatibile regex. As you said, I'm not familiar with the JS regex quirks. I just couldn't invent a good example on the spot, and didn't want to post ones from the code I work on at my day job for legal reasons. And yeah, I agree about "lots of backslashes" part. It gets messy - but splitting regexps in parts makes it at least more manageable. I'm not yet angry enough at the cases I have at my day job to whip up a DSL for it though.
- mikegerwitz 10y agoIt's concise. Regexps can be documented and split onto multiple lines in many languages and commented, be it through string concatenation or formatting modifiers. I write some complicated regular expressions, and I've found that splitting groups of expressions onto multiple lines and indenting them handles most of the problems that my coworkers have with groking them, and that I have when returning to them. I prefer the concise syntax (provided that it's reasonably formatted) for the same reason that I prefer the concise syntax of sed, ed, and similar: it's easy to mentally map and reason about symbols than it is large blocks of text. I've been programming for nearly 20 years and I have found that I much prefer manipulating mathematical expressions than I do large chunks of code, because it uses a concise syntax where symbols mean something. I love such a notation. (In the case of programs, when refactoring, my mind works in blocks of code as units, as I'm sure most others' do.) I'm not saying those benefits aren't possible with verbose code---they are. But just as many prefer a concise mathematical syntax to a verbose program that does the same thing, I prefer a concise formal definition. I'm also not implying that you should try to write an entire grammar in a single regular expression.
- junke 10y agoConcise notations are great, and this is why regexp are so much used IMO. I am by the way a fan of sed, which is clever enough to give you the choice over the delimiter you use (s+/+_+g). On the other hand, there are so many additions to the core formal language, like backtracking or Larry Wall knows what, that syntax has become cryptic. Besides, building regexps out of smaller ones is generally a pain with strings, because you need to quote special regex characters, along with any character that might interfere with the host language's syntax (e.g. emacs regexes with four backslashes in a row). I prefer to read actual words, so the following is fine for me: (defvar *email-regex* '(:sequence :word-boundary ;; IDENTIFIER PART (:regex "[A-Z0-9._%+-]+") #\@ ;; DOMAIN (:regex "[A-Z0-9.-]+") #\. ;; TOP-LEVEL DOMAIN (:greedy-repetition 2 nil (:char-class (:range #\A #\Z))) :word-boundary)) After the recent discussions about Lisp, here is an actual example that can be used by CL-PPCRE to scan "\b[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}\b". The list structure allows you to compose your regular expression like any other list, with intermediate functions, etc. without having ever to thing about escaping your characters. When you need to use the string based, concise regex, you wrap it in a ":regex" form and you have the best of both worlds.
- spc476 10y agoI do a lot of text processing in Lua, primarily because of LPeg. It can parse text that regular expressions can't (or have real trouble with, like http://www.ex-parrot.com/pdw/Mail-RFC822-Address.html http://www.ex-parrot.com/pdw/Mail-RFC822-Address.html), can transform the data on the fly (convert digit characters into its value) and more importantly, they're composable. Have an LPeg expression that can parse an email address? You can then plop that into a larger LPeg expression to parse, say, a header line.
- yellowapple 10y agoYeah, PEGs are pretty sweet. Perl6 is implementing a similar system (at least capability-wise) in the form of "Grammars": http://doc.perl6.org/language/grammars http://doc.perl6.org/language/grammars. They're what regexps should've been all along, IMO.