10 ms·
Tyre – Typed regular expressions
- kukx 10y agoIt's an interesting experiment and an inspiring one. However, I don't think this solution will help too much with regexes. To me it's overly complex and less readable than regex strings.
- Drup 10y agoHi, author here! I didn't invented the combinators for regex, that was already in ocaml-re (which is used as backend for tyre). The combinator approach has many advantages: - You don't need to remember which regex syntax the library is using. Is it using the emacs one ? The perl one ? The javascript one ? The bash one ?! (shudder) ... - It's "self documenting". Your combinators are just functions, so you just expose them and give them type signatures, and the usual documentation/autocompletion/whatevertooling works. - It composes better. You don't have to mash string together to compose your regex, you can name intermediary regexs with normal variables, etc. - Related to the point above: No string quoting hell. Now, tyre doesn't really improve any of that. If anything, combinators are simpler in ocaml-re[1]. What tyre gives you is the automatic extraction and casting of matching groups into the desired datatype (and the reverse direction, but that's almost free bonus). No need to select groups manually and transform into an integer, no need to reconstruct tuples/lists/records, tyre will do that for you (and handle conversion failures cleanly). [1]: https://github.com/ocaml/ocaml-re/blob/master/lib/re.mli#L207-L277 https://github.com/ocaml/ocaml-re/blob/master/lib/re.mli#L20...
- hyperpallium 10y agoHi, your homepage "unparsing" link goes to #evaluating, but the anchor on the docs page has id="eval" Also, how flexible is the unparsing? Does it have to be the same regular expression? [Sorry, I haven't read in detail and I'm only slightly familiar with ocaml - I'm interested from an academic perspective.]
- Drup 10y agoIf I understand your question correctly, no, it doesn't. Actually, the thing that you "unparse" doesn't even have to come from a regex matching to begin with. As long as the type matches, everything is fine. If you are interested by unparsing, I would advise you to read the (famous) paper functional unparsing by Danvy[1]. The module Printf and Format of the OCaml standard library are basically this paper on steroid. I recently wrote a blog on part of the implementation technique[2]. Hope that helps. :) [1]: http://www.brics.dk/RS/98/12/BRICS-RS-98-12.pdf http://www.brics.dk/RS/98/12/BRICS-RS-98-12.pdf [2]: https://drup.github.io/2016/08/02/difflists/ https://drup.github.io/2016/08/02/difflists/ [Printf]: http://caml.inria.fr/pub/docs/manual-ocaml/libref/Printf.html http://caml.inria.fr/pub/docs/manual-ocaml/libref/Printf.htm... [Format]: http://caml.inria.fr/pub/docs/manual-ocaml/libref/Format.html http://caml.inria.fr/pub/docs/manual-ocaml/libref/Format.htm...
- hyperpallium 10y agoAh, thanks. I think one needs (structural) type transformation for the flexibility I had mind, things like swapping sum operands, re-ordering concantenation operands. Uh... did you see my bug report on your webpage (first sentence above)? That link still doesn't work properly.
- Graphon1 10y agoWhat could anyone possibly do to improve Regular Expressions ?
- dkarapetyan 10y agoAdd types of course.
- Kinnard 10y agoHow do these improve regular expressions?
- tantalor 10y agoIt doesn't improve regexp; it improves your code that uses regexp to populate type-checked data. Instead of writing a lot of "if foo is int then ..." type checking code you can just annotate the regexp with types and it either extracts those types for you or throws an error or something. In general you have the same problem when parsing CSV, command line flags, URL query parameters, JSON, etc.
- Kinnard 10y agoI see! Thanks.
- willvarfar 10y agoThose type-safety problems for parsing tainted data are a royal pain and, often, security issues! Typed regex are a very good idea and very welcome. At the moment, when I use regex and extraction, I have a lot of boilerplate and I could easily have lots of bugs unfuzzed. (For those having the type-safety problem with JSON in Python, can I recommend my own library obiwan? https://pypi.python.org/pypi/obiwan/ https://pypi.python.org/pypi/obiwan/ :) )
- chongli 10y agoAt this point, why not just use the more powerful, maintainable, and flexible parser combinators[0]? I thought the point of regex was for quick and dirty bash/grep/perl one-liners and search/replace in vi/emacs. If you're going to actually maintain the code, then use something designed for that. [0] https://en.wikipedia.org/wiki/Parser_combinator https://en.wikipedia.org/wiki/Parser_combinator
- yawaramin 10y agoSee also an interesting typed regex implementation by Gabriel Gonzalez: https://github.com/Gabriel439/slides/blob/master/regex/regex.md https://github.com/Gabriel439/slides/blob/master/regex/regex...
- mchaver 10y agoIs there any advantage to using regex over a parser combinator?
- gergoerdi 10y agoDepends on the parser combinator library. Real regexes can be matched in time linear in the input size; you'll need to use an applicative parser combinator like regex-applicative (https://hackage.haskell.org/package/regex-applicative https://hackage.haskell.org/package/regex-applicative) to get that kind of asymptotic complexity. A monadic one like Parsec is impossible to analyze statically to turn into a finite automaton.
- mchaver 10y agoInteresting. Attoparsec is usually my preferred parsing library, but regex-applicative looks very useful for smaller, less complicated parsers.
- aslakhellesoy 10y agoSlightly related, Cucumber will soon support a typed expression format called Cucumber Expressions: https://docs.cucumber.io/cucumber-expressions/ https://docs.cucumber.io/cucumber-expressions/ It's a standalone library and may well be used outside of Cucumber. It's currently implemented in JavaScript, Java and Ruby.
- 1_player 10y agoA minor nitpick: wrong syntax highlighting is worse than no syntax highlighting at all. The example in the home page is hard to parse and understand.