6 ms·
Show HN: Using functions to construct Regex in Python
- gattilorenz 9y agoWhile the idea is very cool, it seems to have most of the drawbacks of traditional regexes (i.e. the little Syntax quirks that eventually one has to learn), with none of the benefits (regex are everywhere including text editors, not just a python thing). It does make them more readable I guess, but I'd like to hear other HNers' opinions. I think for beginners a better approach would be using something like RegexBuddy (https://www.regexbuddy.com https://www.regexbuddy.com not affiliated, just found it super useful when I started writing regexes).
- pygy_ 9y agoI have a similar lib in JS that uses either regexps or plain strings for the leafs, and combinators to implement the various regexp operators. It is useful for maintainability. The regexp syntax was designed for a write-only scenario (the command line), but complex regexps quickly become unwieldy. The regular grammar is all about composition, and the regexp syntax (in JS) doesn't allow one to store sub-expressions in variables for reuse and readability. So this allows one to create regexps piece by piece, to write tests for the sub-expressions, etc... https://github.com/pygy/compose-regexp.js https://github.com/pygy/compose-regexp.js
- Drup 9y agoRegex combinators are quite common in functional languages (where it works quite better than in python). One particular advantage compared to Regex syntax is that you stay inside the language and only use function calls, so you get the documentation, typing and checking property of the language. I said a bit more about that here: https://news.ycombinator.com/item?id=12385840 https://news.ycombinator.com/item?id=12385840
- baq 9y agoso these would work great with python augmented with mypy.
- ivanbakel 9y agoMypy is pretty painful to use, though. The benefits of typing aside, patching it into Python has become too-little too-late, and everything feels like a kludge.
- rtpg 9y agoYou can get a lot of advantages of Mypy by combining it with a whitelist/blacklist (you can do this through a config file). Lots of missing stuff but with strict optionality alone you're gonna catch a lot of real world bugs
- baq 9y agoit's painful and in hindsight too late but i disagree that it's too little. i've been using it for about a year now and the amount of bugs it finds results in a (very) positive ROI; at $WORK new projects are now started with full mypy awareness from day 1 and old ones are being slowly but surely retrofitted.
- deleted 9y ago[deleted]
- stonewhite 9y agoI don't really think this is more readable than the regex itself once it starts to get complex, like they all do. Maybe if this was implemented on a prefix notation language it may have would looked/read better
- fiddlerwoaroof 9y agoI really like the way Common Lisp's ppcre library works: the functions all accept either standard strings or a s-expression version of the regular expression and then, using compiler macros, all invocations with statically-determinable regular expressions get compiled to some internal representation at compile-time rather than generating that representation at run-time. http://weitz.de/cl-ppcre/#create-scanner2 http://weitz.de/cl-ppcre/#create-scanner2
- iogf 9y agoThe example is longer than the regex itself however it is simpler to explain to a beginner what crocs's example does than explaning the regex itself. Based on that assumption using crocs's syntax should favour reasoning somehow since it is simpler to understand than obscure regex's syntax.
- dizzystar 9y agoI agree with you. At some point, you are bound to get into situations where writing the regex is shorter and easier to do the normal way. Some the examples are getting a little long, and I'm not entirely sure if they are more clear. Also, it's only for Python2.
- iogf 9y agoIt can be ported to py3 with no difficult though it is meant to be used as a tool to construct regex, it shouldnt be put in your programs for a matter of performance at all(mainly those that are critical). About some examples being long, i believe verbosity pays off in understanding complex systems, so yea, in some situations it wouldnt benefit using crocs unless you dont know regex and you dont want to spend some hours to get proficient in it.
- orf 9y agoInteresting, how does this compare to PyParsing[1]? It seems that pyparsing does a lot of this already. Not that you shouldn't write your own, mind you! Some improvements could be to use operators to reduce some boilerplate, like using * instead of Times(), i.e `(X() * 3) * 5` 1. http://pyparsing.wikispaces.com/ http://pyparsing.wikispaces.com/
- iogf 9y agoInteresting suggestion.
- agumonkey 9y agoI always liked this idea (reminiscent of emacs rx dsl). But I'm not sure I'd be using it that often...
- carlochess 9y agoHi, what's the diference between this project and a parser combinator?
- yablak 9y agoA parser combinator may be more powerful? Not sure. The python parsec package is really terribly documented, but there's a good example here: http://www.valuedlessons.com/2008/02/pysec-monadic-combinatoric-parsing-in.html http://www.valuedlessons.com/2008/02/pysec-monadic-combinato... EDIT: The name of the package may have changed; it looks like this is the pypi package being documented (if not; i'm confused because pypi's pysec package is something totally different): https://pypi.python.org/pypi/parsec https://pypi.python.org/pypi/parsec
- deleted 9y ago[deleted]
- matthberg 9y agoReally cool concept, yet not practical in my opinion. The classes and functions called to replace the regex I consider to be harder to learn than the initial regex. With regex there is a standardised, succinct way to query strings, while this system adds unneeded complexity for the sake of using English words. For quick uses integrated in other programs I can see the usefulness, yet as a standalone I'll stick with regex.
- iogf 9y agoThe idea is having a set of entities on which one could reason better to implement more complex filters(regex) as well as debugging them. It is sort of a way of reasoning on filtering results it merely uses regex as an underlying tool. One could write the filter using crocs format, compile it to regex then use in their programs. Using crocs you sort of define your data type then apply functions on those data types to fetch your required output. The crocs framework is a way to give more power to your imagination and readability to others about what your imagination produces.
- avar 9y agoIn case you aren't aware of this, what you've made is very similar to the rx.el[1] that's shipped with Emacs since 21.1 (released in 2001). Many of the comments here are speculating on what this might/could be used for, but could simply look at how it's used in various Emacs modes compared to writing raw regexes in string form. It's not fully comparable, since a large reason to use rx.el in Emacs is Emacs's nasty regex syntax coupled with having no native quoting construct for regex, making writing anything an exercise in bashing your backslash key. But it also makes it easier to programmatically generate regexes, which as your Python implementation also shows can lead to much clearer code. 1. http://doc.endlessparentheses.com/Fun%2Frx.html http://doc.endlessparentheses.com/Fun%2Frx.html
- bhrgunatha 9y ago> compared to writing raw regexes in string form ... > Emacs's nasty regex syntax It's also case-sensitive too. Writing regex by hand when there is a mixture of case sensitivity to recognise is not at all pleasant.
- nemetroid 9y agoI find this: mail = '(?P<name>[a-z][a-z0-1\_\.\-]{1,})' hostname = '(?P<hostname>python[a-z]{1,})' domain = '(?P<domain>br[a-z])' match_mail = f'{mail}\@{hostname}\.{domain}' quite more readable and easy to verify than the 50-line example in the link. The real advantage of parser combinators over regexes is the ability to parse data structures instead of just capturing regex groups.
- klenwell 9y agoOr even just code comments, as with this example[0]: import java.util.regex.Pattern; public class Main { public static void main(String[] args) { String regexStr = ""; regexStr += "\\b"; //Begin match at the word boundary(whitespace boundary) regexStr += "\\d{3}"; //Match three digits regexStr += "[-.]?"; //Optional - Match dash or dot regexStr += "\\d{3}"; //Match three digits regexStr += "[-.]?"; //Optional - Match dash or dot regexStr += "\\d{4}"; //Match four digits regexStr += "\\b"; //End match at the word boundary(whitespace boundary) if (args[0].matches(regexStr)) { System.out.println("Match!"); } else { System.out.println("No match."); } } } [0] https://codereview.stackexchange.com/questions/47432/commenting-string-matching-regex https://codereview.stackexchange.com/questions/47432/comment...
- kqr 9y agoPython even supports this in its regex library, where it's called "verbose regexes". You'd write your code as def main(args): regexStr = \ """ \\b # begin match at word boundary \\d{3} # match three digits [-.]? # optional – match dash or dot """ if args[0].match(regexStr): print("match") else: print("no match") It's even better with nested structures, because it allows indentation as well!
- blibble 9y agothose comments are no better than: i++; /* increase i by one */ sure, break up a larger regex to reflect its structure and comment those sections, but something like \\d{3} or [-.]? should really not need commenting
- reikonomusha 9y agoIn Common Lisp, the library CL-PPCRE [0] is used for regexes, and has an "AST"-syntax like this. It supports full Perl-compatible regular expressions in both the tree syntax as well as the standard string syntax. The tree syntax actually uses symbols and lists, idiomatic in symbolic programming settings. This has the hugely convenient property of being trivially serializable, just as the string representation is. This isn't true with an unadorned object representation. The documentation for the tree syntax is here [1]. Some examples from the page are reproduced below. The PARSE-STRING function isn't used by the user except for testing. All of the regex scanning and matching functions allow either a string or a tree. * (parse-string "(ab)*") (:GREEDY-REPETITION 0 NIL (:REGISTER "ab")) * (parse-string "(a(b))") (:REGISTER (:SEQUENCE #\a (:REGISTER #\b))) * (parse-string "(?:abc){3,5}") (:GREEDY-REPETITION 3 5 (:GROUP "abc")) ;; (:GREEDY-REPETITION 3 5 "abc") would also be OK * (parse-string "a(?i)b(?-i)c") (:SEQUENCE #\a (:SEQUENCE (:FLAGS :CASE-INSENSITIVE-P) (:SEQUENCE #\b (:SEQUENCE (:FLAGS :CASE-SENSITIVE-P) #\c)))) ;; same as (:SEQUENCE #\a :CASE-INSENSITIVE-P #\b :CASE-SENSITIVE-P #\c) * (parse-string "(?=a)b") (:SEQUENCE (:POSITIVE-LOOKAHEAD #\a) #\b) [0] http://weitz.de/cl-ppcre/ http://weitz.de/cl-ppcre/ [1] http://weitz.de/cl-ppcre/#create-scanner2 http://weitz.de/cl-ppcre/#create-scanner2
- kazinator 9y agoIt's worth mentioning that CL-PPCRE is not a set of bindings to a C implementation. It is written in Common Lisp, and it's fast.
- reikonomusha 9y ago(Separate comment addressing the library itself.) I find the naming of the classes off. Why not call X as Any? Why use nouns for some things, and verbs for others? I found Include and Exclude to be confusion. Inclusion and exclusion are relative to something. You include something with something else. I think Only and AnyBut would be better names. Why call it Seq when it's closer to a Python Range? Seq usually means sequentiality. Why not call Times as Repeat or Repetition? It didn't look like this had any compatibility with existing regexes. I can't parse an existing regex into this library. It looks like debugging could be a nightmare since it just passes things off to Python's re library. So if an error happens there, it will be tough to trace it to your original construction. All of this class hierarchy building seems prime to replace with algebraic data types, which would probably cut the code down to just a fraction of the size.
- iogf 9y agoI may add synonyms for those classes, so, it is just a matter of adding: Repeat = Times etc. you wouldnt use these structures in python code, you use them to build your regex then compile it and use the regex in your programs. It is possible to add some type safety and messages to help debugging which would be better than with natural regex syntax.
- iogf 9y agoSeq stands for a-z or 0-9 which can be used only in Include and Exclude. When you write Include(Seq('a', 'm')) it means you are including a character from that sequence in a given position of your pattern. So if you have Pattern(Include(Seq('a', 'm')), 'c') That is sort of building the following set of strings. ac bc cc . . . When you use Exclude, it would be like you are picking a char that is not belonging to the sequence, like in: Pattern(Exclude(Seq('0, 9')), c) then you would end up with a set like: ac bc cc . . . It means, a set whose first character of the elements arent digits.
- masklinn 9y agoSeq is odd, to me Seq('a', 'm') says it matches 'am'. Why not overload X to support filtered forms instead? E.g X() is ".", X(from_, to_) is $from-$to, X(items) is [$items] and X(exclude) is [^$items]?
- rnhmjoj 9y agoIf you are interested in different ways to write/implement regular expression this talk is worth the watch: https://begriffs.com/posts/2016-06-27-fast-haskell-regexes.html https://begriffs.com/posts/2016-06-27-fast-haskell-regexes.h...
- bane 9y agoRegexes aren't really all that hard if you avoid all the crazy back reference craziness. I've taught them to non-programmers in about a half-hour and they were able to use them reasonably well right after with a couple reminders. There's three things you need: 1 - concatenation - basically one thing next to another, 'a' goes next to 'b' to make 'ab' which matches any string with that in it. Examples: 'xzyqr2321abtwe' - matches 'zyxabc' - matches 'abcxyz' - matches 'ab' - matches 'ba' - doesn't match 'azb' - doesn't match 2 - alternation - one thing or another. The operator for this is the pipe symbol '|'. So 'a|b' is 'a or b'. 'xzyqr2321abtwe' - matches 'xzyqr2321atwe' - matches 'xzyqr2321btwe' - matches 'zyxabc' - matches 'abcxyz' - matches 'ab' - matches 'azb' - matches Here's another 'abc|cab' 'abc or cab' 'xzyqr2321abtwa' - doesn't match 'zyxabc' - matches 'zyxcba' - doesn't match 'zyxcab' - matches If it helps, you can use parentheses to group blocks of things for clarity. '(abc)|(cab)' 3 - repetition - the operator we want to care about is '* ' also called the Kleene star. This means that anything that comes before the star can repeat 0 or more times. so 'a* ' means a pattern of 'a' 0 or more times. 'xzyqr2321abtwe' - matches 'xzyqr2321btwe' - matches (because of zero or more times) 'xzyqr2321btwaaaaaaaaa' - matches so if you want to match an 'a' one or more times you can simply use rule #1 (concatenation) and put 'aa* ' 'xzyqr2321abtwe' - matches 'xzyqr2321btwe' - doesn't match 'xzyqr2321btwaaaaaaaaa' - matches It turns out this is such a common regex, that some "syntactic sugar" was invented to make it easier to work with, the '+' symbol. Which is used exactly the same way: 'aa* ' = 'a+ ' Congratulations, you now know everything you need to make a regex that can pretty much do anything. So what about all that other line noise one usually sees in a regex? That's more "syntactic sugar", designed to make certain kinds of regexes simpler to write. For example, a regex using the above rules that can match any alphabet letter from a through f is: a|b|c|d|e|f This is a lot of typing, so you can use square brackets to build what's called a "character class". '[abcdef]' = 'a|b|c|d|e|f' If you have a long string of characters, you can use some more sugar and simply put a '-' in between the first and last characters '[a-f]' = '[abcdef]' Here's a complex example, a regex that can match any lower-case alphabet character or number '[0-9a-z]' It also turns out that a special character class pattern that can match any character (except for newlines) is so common that a special operator '.' was created. '.' matches literally anything. For example, concatenating it with an "any number" regex gives you '.[0-9]' which means any character followed by a number. For repetition, there's also some syntactic sugar, here's a table, these all go after the thing you want repeated: * - 0 or more times + - 1 or more times {p} - p times {p,q} - p to q times {,q} - 0 to q times {p,} - at least p times ? - 0 or 1 times Finally, in the list of basic regex stuff, there's "capturing", this is how regexes get used to parse strings. It works basically like this, anything that is matched inside of a parentheses gets "copied" to a special variable. It depends on your language on where this variable is. The name of the variable is typically some 1, indexed variable like $1, $2, $3. The number of the variable is the number of left parentheses in the regex. Example: 'abc(.+)123' abcd123 - matches, $1 is 'd' abcf123 - matches, $1 is 'f' abcqqq123 - matches, $1 is 'qqq' abc123 - doesn't match Another example: (abc([0-9]+)) abc123 - matches $1 is abc123 $2 is 123 If you want to suppress this capturing behavior, the '(' parentheses should be written '(?:' (?:abc([0-9]+)) Try these out here, it does a nice job highlighting https://regex101.com/ https://regex101.com/ it calls the capture variable "groups" and shows them on the right under "capture information" There, that's about it. Other advanced techniques like greedy matching, negative character classes, callbacks, etc. flow pretty nicely from these basic ideas.
- deleted 9y ago[deleted]
- rohitpaulk 9y agoA more mature version: https://github.com/VerbalExpressions/PythonVerbalExpressions https://github.com/VerbalExpressions/PythonVerbalExpressions
- iogf 9y agoThis one doesnt look to support lookahead/lookbehind nor it outputs valid inputs for the matches though. However, it is interesting the way of how you chain the VerEx to build the patterns.
- ddebernardy 9y agoI'm confused. Doesn't Python offer verbose regular expressions complete with comments? https://docs.python.org/3/library/re.html#re.VERBOSE https://docs.python.org/3/library/re.html#re.VERBOSE a = re.compile(r"""\d + # the integral part \. # the decimal point \d * # some fractional digits""", re.X) Or is there more at stake than making readable, self-documenting regular expressions?
- always_good 9y agoIt has the same advantages as combinators: the total computation can be broken down into pieces, tested individually, and then composed and reused, especially programmatically. Why use functions when you can have one big main(args: Array<String>) with inline comments? btw, a hello-world snippet like /\d+\.\d*/ isn't very scathing criticism of a library.
- rajaravivarma_r 9y agoI always wanted to do something like this. Because I always forgot how to do the 'look ahead' assertions. I even named the library 'hu_regex', where 'hu' stands for human, but never did anything further, as I couldn't think of any good builder patter for regex that seemed intuitive. Anyway, thanks for this. This comment section has introduced some good libraries as well.
- anentropic 9y agoHow often do you programmatically generate regexes though? I tend to find that regex patterns are effectively constants in the program And if it's just for readability it seems a bit of a sledgehammer
- iogf 9y agoI have been adding more docs to: https://github.com/iogf/crocs/wiki https://github.com/iogf/crocs/wiki It would be interesting to hear comments :)