5 ms·
OK, these kinds of regex tools get posted quite often. I get it, regex is very confusing at first. And some of these use-cases result in rather complex expressi
by robert_tweed 7y ago
OK, these kinds of regex tools get posted quite often. I get it, regex is very confusing at first. And some of these use-cases result in rather complex expressions nobody should be forced to write from scratch (you are still remembering to write unit tests for them though, right?)
But as someone who actually knows [some flavours of] regex fairly well, what I would really like, is a reference that covers all the subtle differences between the various regex engines, along with community-managed documentation (perhaps wiki pages) of which applications & API versions use which flavour of regex.
For example, the other day I wanted to run a find on my NAS. I needed to use a regex, but the Busybox version of find doesn't support the iregex option, so all expressions are case-sensitive. With some googling, I was able to find out that the default regex type is Emacs, but I wasn't able to find either a good reference for exactly what Emacs regex does and doesn't support, nor any information about how to set the "i" flag. In the end I had to manually convert every character into a class (like [aA] for "a") which was tedious, but quicker than trying to find a better solution or resorting to grep.
A related, annoyingly common pattern is that the documentation for `find` states that `--regex` specifies a regex, but it does not state which flavour of regex. The documentation for certain versions of `find`, which support alternative engines, note that the default is Emacs. From this I was able to infer (perhaps wrongly) that the Busybox `find` uses Emacs-flavoured regex, but ultimate I still had to resort to some trial-and-error. This problem is all too common in API documentation.
- geongeorgek 7y agoYou're totally right. Right now this tool only supports the javascript flavor of regex. That said, for all the simple expressions shown there it's more or less the same for most other engines. I guess that makes it okay.
- celeritascelery 7y agoThe O’Riley book “mastering regular expressions” has a whole section dedicated to it. As well as several tables. But it would be nice to have an online version.
- wyclif 7y agoAnd it's one one the best O'Reilly books. I went and checked because of your comment and just noticed there was a third edition that I missed, I have the second. Still a book worth studying.
- justaj 7y agoHonestly, as a noob, this is one of the biggest reasons I have such a hard time deciding to learn regex. Python flavor would probably be different than PCRE, which is probably different than JS flavor. Even worse is that it might be too late to standardize all the regex flavors because there is already so much written in different regex flavors that it just costs too much for them to become obsolete in the future. This is really demotivating.
- chirss 7y agoHonestly don't let this get you down, here's a learning plan (use regex101 to learn) 1) Learn PCRE regex. 2) Try regex golf or cross words to learn PCRE regex. 3) Take the quiz on regex101. Once you're done with all 3: Learn the minor/major differences in the other languages. There aren't many. For example this named capture group: (?<somename>someregex) Would look like this in a different language: (?P<somename>someregex) There's some differences about what language can and cannot do like recursion because someone thought it was a great idea to make javascript awful at regex, but that's besides the point. Regex is totally worth learning.
- new_guy 7y ago> Honestly, as a noob, this is one of the biggest reasons I have such a hard time deciding to learn regex. Clear your afternoon, and just learn it. Seriously, it takes a couple of hours at best and then - BOOM - you're done for the rest of your life.
- absorber 7y ago> you're done for the rest of your life. If that were so easy then I don't think much of these cheatsheets would exist.
- quickthrower2 7y agoThe basic regex is easy, infact an English word is a regex! A dot matches a single character. Star multiple of the previous character. Just that is useful for a lot of cases!
- chirss 7y agoregex101 does a good job at showing you what the selected variant can do.
- mklein994 7y agoI tend to go to https://www.regular-expressions.info https://www.regular-expressions.info when I need to find out which features are supported between dialects. Not always up-to-date, but has some good info.
- alexhutcheson 7y agoRE2 syntax[1] is a pretty good option to learn, because it's mostly a "lowest common denominator" - if it works in RE2, it should work in PCRE, Python, Javascript, etc. The reverse isn't true - there is a bunch of syntax that RE2 doesn't support by design, often to constrain performance bounds. Emacs regexps are unfortunately their own weird beast - they handle parentheses differently than other regexp engines, because Emacs assumes that you'll be running regexps on Lisp code a lot and want to easily match parentheses. The best documentation on that syntax is (confusingly) in the Elisp reference manual: https://www.gnu.org/software/emacs/manual/html_node/elisp/Syntax-of-Regexps.html#Syntax-of-Regexps https://www.gnu.org/software/emacs/manual/html_node/elisp/Sy.... [1] https://github.com/google/re2/wiki/Syntax https://github.com/google/re2/wiki/Syntax
- useragent86 7y agoIME Emacs provides a very pleasant way to write regexps using the rx library. ELPA also has the package xr, which converts Elisp regexps to rx format, and pcre2el converts PCRE to Elisp. So a regexp like \b[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,4}\b Can easily be converted like: (->> "\\b[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,4}\\b" pcre-to-elisp xr) To: (seq word-boundary (one-or-more (any "0-9A-Z" "%+._-")) "@" (one-or-more (any "0-9A-Z" ".-")) not-newline (repeat 2 4 (any "A-Z")) word-boundary)
- alexhutcheson 7y agoAgreed that rx is nice, but really only useful if you're writing elisp. 90%+ of people who need to interact with Emacs regexps aren't writing an elisp program - they're using Emacs interactively, or even using another program (busybox, GNU find, etc.) that uses Emacs regexps for historical reasons. For those people the differences in syntax between Emacs regexps and "normal" regex dialects are a pain.
- waz0wski 7y agoif you're on osx, the app Patterns is really good for testing regex, and also has quick references for a variety of regex 'engines' and also has decent matching explanations https://krillapps.com/patterns/ https://krillapps.com/patterns/
- 8bitsrule 7y agoBy coincidence, I found this link a bit earlier today. It tries to avoid flavors and exotic syntax. https://rexegg.com/regex-quickstart.html https://rexegg.com/regex-quickstart.html
- zaptheimpaler 7y agoIts like SQL - everyone has a dialect. For most things where a SQL/regex engine/parser isn't the core of what they do, it will never be a priority. The best approach IMO is something like this in priority order: 1. Stick to using the lowest common denominator like you did for case insensitivity. 2. If that becomes too cumbersome, then consider whether regex is the right tool for the job. Maybe you can use e.g Python/your favorite language with a known regex standard. 3. If there are no other tools and you're stuck with whatever flavor of regex one particular thing supports, only then invest time in learning the details. There is probably a book out there with the details even if there's no webpage. Then pray you never get to step 3 :)
- AlchemistCamp 7y agoTo me, the divide is pre and post-Perl. It's not so bad going between JS, Ruby and Elixir regex (possibly due to my use of a smaller set of features), but VIM regex disappoint me time after time.