14 ms·
It might just be a me problem, but I've always been wary of regexes. They're not too bad to write, but reading them back and understanding what's actually goin
by AussieWog93 2mo ago
It might just be a me problem, but I've always been wary of regexes. They're not too bad to write, but reading them back and understanding what's actually going on can get a bit hairy. Plus, all of the subtle differences between regex libraries seems like a bit of a footgun.
Obviously they have their place, but I know a lot of the older guys seemed to love them way more than the young.
- rhdunn 2mo agoVarious libraries (e.g. Python's `re` library) support comments and whitespace as an option allowing you to format the regex on multiple lines with commenting to document what each part does. I'm not sure if there are any regex libraries that support DSLs and easy composability (e.g. the email RFC regex would be easier to read/maintain if you could specify the individual parts like are defined in the RFCs).
- AussieWog93 2mo agoI honestly never knew that, should give it another go.
- klibertp 2mo agoI would recommend trying something like PyParsing[1] instead. Libraries like this allow you to compose the parser from language-level entities (object and functions, on top of regex and string literals). This means you can attach comments to those entities naturally within the syntax of the language. You also get much better error reporting out of the box, as well as a well-defined way of attaching transforming code to parts of the parser. There's a place for simple regexes, but complex regex DSLs (with comments and non-significant whitespace, etc.) are almost always less convenient than simply using your language directly. [1] https://pyparsing-docs.readthedocs.io/en/latest/HowToUsePyparsing.html#hello-world https://pyparsing-docs.readthedocs.io/en/latest/HowToUsePypa...
- frizlab 2mo agoSwift even has a `RegexBuilder` DSL which makes writing regular expressions pure code and type-safe. Pretty amazing tbh
- klibertp 2mo agoEmacs/Elisp has the rx library: https://www.gnu.org/software/emacs/manual/html_node/elisp/Rx-Notation.html https://www.gnu.org/software/emacs/manual/html_node/elisp/Rx... You get s-exp-based regex syntax (example for C-style block comments; there are shorter aliases too, e.g. `zero-or-more` can be written as `*`): (rx "/*" ; Initial /* (zero-or-more (or (not "*") ; Either non-*, (seq "*" ; or * followed by (not "/")))) ; non-/ (one-or-more "*") ; At least one star, "/") ; and the final / and you have rx-define and rx-let to defined named subforms: (rx-let ((comma-separated (item) (seq item (0+ "," item))) (number (1+ digit)) (numbers (comma-separated number))) (re-search-forward (rx "(" numbers ")"))) And this is just the regex builder - syntactic sugar - as it still just builds a single regex serialized to a normal string. I tend to use it everywhere, since it is guaranteed to always properly escape all backslashes (a major pain point in string regexes in Emacs), but it's also useful for building larger regexes from chunks and reusing chunks in multiple related regexes.*
- bryanrasmussen 2mo agothere is, I think, a divide between programmers that is pretty basic. Do they need a language that maps somewhat to written human language, or can they adapt to languages that do do not at all resemble the human languages they are familiar with. This divide is most probably cultural, programmers in Western societies often have pre-programming familiarity with English and thus they do not need to learn a language that does not match to how they understand languages to work (as might be the case with programmers from Asian countries or others where familiarity with English is not guaranteed) So if your primary gateway to programming languages are ones that slightly resemble a human language you are familiar with you may have lots of psychological blocks keeping you from making that final jump to reasoning in J, or APL, or even a DSL like regular expressions. Of course DSLs also have the problem that many programmers do not seem to fit well in things that do not have all the logical control operators they are used to, thus programmers who do not handle CSS, SQL or similar languages even though they are significantly simpler than a full featured programming language. In short, things that are very different from what you are used to will probably be difficult to learn, use, and remember, and the same goes for most of your coworkers.
- graemep 2mo ago> as might be the case with programmers from Asian countries or others where familiarity with English is not guaranteed Lots of Asian countries where familiarity with English is assumed in professional contexts. > So if your primary gateway to programming languages are ones that slightly resemble a human language you are familiar with you may have lots of psychological blocks keeping you from making that final jump to reasoning in J, or APL, or even a DSL like regular expressions. That raises the interesting possibility that J or APL might be more appealing to non-English speaking countries, or maybe where the dominant languages are not Indo-European (so not similar to English either). I wonder whether there is any evidence of this?
- bryanrasmussen 2mo agothe might at beginning of the clause was also meant to take into account that people might have familiarity with English, as I could not be certain, but probably should have been expressed better.
- tcp_handshaker 2mo ago>> I've always been wary of regexes. Be afraid: https://owasp.org/www-community/attacks/Regular_expression_Denial_of_Service_-_ReDoS https://owasp.org/www-community/attacks/Regular_expression_D... https://en.wikipedia.org/wiki/ReDoS https://en.wikipedia.org/wiki/ReDoS
- pygy_ 2mo agoThis is by design, the regexp syntax has been invented for write-only programming at the CLI, and graduated to ubiguitous programming language syntax because worse is better. The regular formalism is all about composability, and most languages don't offer a way to compose regexps, which is a real shame IMO.
- ddevnyc 2mo agoI know I am very very alone in this, but I've always found regex very readable. It's just not _quickly_ readable. You can't look at 20 characters of regex and read it and understand it as quickly as you would 20 characters of English. The biggest issue, I think, is people trying to do that. Regex is very information dense, it should be approached with the care of a mathematical formula or a sudoku rather than English prose.
- bazoom42 2mo agoThe readability should be compared to alternative ways to solve the same problem. Sure, regexes are not the most intuitive syntax, but it is compact and declarative. What is the alternative? Substring searches? Looping over characters? Hand-rolled recursive descent? Neither are obviously more readable, and intermingles the pattern with the mechanism.
- GuB-42 2mo agoRaku worked on this. As the Perl successor, they gave regex a lot of attention. The result still look like regex, but more powerful and with a more consistent syntax. They also made "/x" the default, which ignores unescaped whitespace and lets you put comments, so you can use spacing and indentation for readability. The general idea is that they are treated like actual programs rather than extended search strings. But Raku, despite some good ideas and what looks like a nice community is not mainstream to say the least. So I don't expect "RCRE" to become a thing anytime soon.
- m463 2mo ago> older guys seemed to love I actually don't love regular expressions. But honestly I think it is implementations that make me dislike using them. They are unclear in most programming languages. Using regular expressions in just about every language I've used has had this programming language vs regular expression language ambiguity that makes them hard to recommend in production past minimal complexity. I don't like to hand off code to my coworkers where it's unclear if a character is part of the quoting system, part of the programming language, part of the regular expression syntax, or a character to match literally. for example, what if the program variable foo contained "abc" and you wanted that to be matched by a regular expression. each language has a different way of doing this and reviewing the code has a high chance of an error unless the person is really pedantically accurate regarding regular expressions. for example a regular expression in bash vs python is different because of quoting and escapes. And what if you wanted to use it in a search and replace? What would help would be: - a very very syntax aware editor that could color the regular expression, showing language characters vs regular expression control characters vs literals - a tool for bidirectional conversion. Type in a pure regular expression and it will put out the expression in your programming language. or check an expression in the language and it will expand/annotate the regular expression. (maybe there are things like this?)