5 ms·
What could anyone possibly do to improve Regular Expressions ?
by Graphon1 10y ago
What could anyone possibly do to improve Regular Expressions ?
- dkarapetyan 10y agoAdd types of course.
- Kinnard 10y agoHow do these improve regular expressions?
- tantalor 10y agoIt doesn't improve regexp; it improves your code that uses regexp to populate type-checked data. Instead of writing a lot of "if foo is int then ..." type checking code you can just annotate the regexp with types and it either extracts those types for you or throws an error or something. In general you have the same problem when parsing CSV, command line flags, URL query parameters, JSON, etc.
- Kinnard 10y agoI see! Thanks.
- willvarfar 10y agoThose type-safety problems for parsing tainted data are a royal pain and, often, security issues! Typed regex are a very good idea and very welcome. At the moment, when I use regex and extraction, I have a lot of boilerplate and I could easily have lots of bugs unfuzzed. (For those having the type-safety problem with JSON in Python, can I recommend my own library obiwan? https://pypi.python.org/pypi/obiwan/ https://pypi.python.org/pypi/obiwan/ :) )
- chongli 10y agoAt this point, why not just use the more powerful, maintainable, and flexible parser combinators[0]? I thought the point of regex was for quick and dirty bash/grep/perl one-liners and search/replace in vi/emacs. If you're going to actually maintain the code, then use something designed for that. [0] https://en.wikipedia.org/wiki/Parser_combinator https://en.wikipedia.org/wiki/Parser_combinator
- Drup 10y agoWell, there is the speed argument. Regular expressions (not the pcre kind with backtracking and stuff) are blazing fast. Ragel[1] is a good example of that. In general, though, I actually agree. [1]: http://www.colm.net/open-source/ragel/ http://www.colm.net/open-source/ragel/
- chongli 10y agoAmong parser combinator libraries, attoparsec[0] is no slouch. It gets very close to the speed of the hand-rolled http-parser library while being much shorter, easier to read, and maintain[1]. [0] https://hackage.haskell.org/package/attoparsec https://hackage.haskell.org/package/attoparsec [1] http://www.serpentine.com/blog/2010/03/03/whats-in-a-parser-attoparsec-rewired-2/ http://www.serpentine.com/blog/2010/03/03/whats-in-a-parser-...
- tantalor 10y agoAnother good example is RE2: https://github.com/google/re2 https://github.com/google/re2 It's pretty fast too. I don't know whether you can compare the performance of Ragel to regex libraries like RE2; they look like they solve different problems.
- gergoerdi 10y agoAs I mentioned in another comment, there are applicative parser combinator libraries for regular languages, e.g. https://hackage.haskell.org/package/regex-applicative https://hackage.haskell.org/package/regex-applicative so asymptotic speed is orthogonal to 'combinatorialness'.
- Drup 10y agoIf you look at the API of ocaml-re (and tyre), you'll see it's pretty much equivalent to regex-applicative. Regex combinators are nice. "Parser combinators", for me, denotes actual parsers, for non-contextual or contextual grammars.
- tantalor 10y ago
- Drup 10y agoIndeed, that is the exact goal. For the little history, I started this for URLs[1]. I couldn't (and still can't) figure out a proper interface for that, but someone mentioned he would like something like that for regular expressions, so here it is. We already have something similar for command line flags in OCaml[2] (and several for JSON). [1]: https://github.com/Drup/furl https://github.com/Drup/furl [2]: http://erratique.ch/software/cmdliner http://erratique.ch/software/cmdliner
- coldtea 10y agoLots of things. Make syntax consistent across RE engines for one. Improve the syntax. Introduce clear modes with different tradeoffs, so you don't get e.g. regex based DOS attacks. Make them easier to compose, not just through copy pasting...
- adrusi 10y agoPerl6 made a lot of improvements https://docs.perl6.org/language/regexes https://docs.perl6.org/language/regexes Technically not regular expressions since they are appropriate for matching much more than regular grammars, but that's true for almost any programming language's "regular expressions".
- cyphar 10y ago> Perl6 made a lot of improvements https://docs.perl6.org/language/regexes https://docs.perl6.org/language/regexes IMO the word "improvements" should be in quotes. Pathological regular expressions are an issue because of things like back references and look-behind matching. If Perl hadn't enshrined the idea of "a regular expression for every problem" people would probably be less skeptical of regular expressions in code.
- berntb 10y agoSo why not use grammars, if regexps are too complex...? https://docs.perl6.org/language/grammars https://docs.perl6.org/language/grammars (The extensions to regexps in Perl 5 (and presumably now also in PCRE?) have made it possible to write them more like grammars than anything else. But you would probably not approve. :-) )
- vmorgulis 10y agoA linq interface to visit the resulting parse tree.
- unsignedqword 10y agohttps://github.com/GuillaumeBadi/Verbal-Exprejon https://github.com/GuillaumeBadi/Verbal-Exprejon
- draegtun 10y agoI've steered more to composing grammars by using tools like Regexp::Grammars - https://metacpan.org/pod/Regexp::Grammars https://metacpan.org/pod/Regexp::Grammars The real sweet spot I've found is parse in Rebol / Red - http://blog.hostilefork.com/why-rebol-red-parse-cool/ http://blog.hostilefork.com/why-rebol-red-parse-cool/ | http://www.rebol.com/docs/core23/rebolcore-15.html http://www.rebol.com/docs/core23/rebolcore-15.html