6 ms·
The true power of regular expressions (2012)
- cadamsdotcom 2mo agoSome people, when confronted with a problem, think "I know, I'll get my agent to solve it with regular expressions." Now they have three problems.
- jfyi 2mo agoIt's ok, someone else can maintain it.
- HelloUsername 2mo ago"Doom Using Regular Expressions" https://news.ycombinator.com/item?id=49094081 https://news.ycombinator.com/item?id=49094081
- PeterHolzwarth 2mo agoThis is wonderful! You should submit this to the HN feed.
- lbriner 2mo agoSomething that seems obvious but not always implied by people's comments is that people are rarely trying to match an entire document with a regular expression so it doesn't really matter that "HTML is not a regular language". If I am trying to e.g. count div tags with a regex like "<div" or whatever, then clearly this would work in 99.9% of cases and probably achieve what the poster is looking for. As soon as you also add character classes to ignore various parts of the document that you are not interested in like "<div[^>]*>" or whatever it is, then it is eminently useful even if the bit we are ignoring is not fully regular. One lovely thing about regex is how fast it is. I was asked to parse a massive CAN Bus log file for how many times some event had logged. This was the early 2000s and the file was 6GB, which was pretty big. I tried .Net's string.StartsWith or something and that took ages to run through the file. I did the same thing with a regex and it finished in like 5 seconds (HDD, not SSD!). I don't know how the magic works but it is very impressive.
- Gabriel54 2mo agoRegex can also be horribly slow - it depends on the particular regex you are using.
- flossly 2mo ago> Regex can also be horribly slow - it depends on the particular regex you are using. And the alternative approach we are comparing to.
- bob1029 2mo agoAre we also counting divs inside comments or literals within scripts? There is a reason this advice is default. The chances an edge case exist are probably a lot higher than anyone is prepared to accept. Even in the "simple" cases.
- reichstein 2mo agoRegExps is the way to tokenize, so it's not surpassing you can look for individual tokens using them. It's parsing that's hard , for example when it needs to match up braces, or start and end tags, even if either is easily matched by a RegExp. And you still need to be careful if the source you're looking in has any way to escape text or have different meanings for the same text. In source code, you should recognize comments and strings (and RegExp literals) so you don't match inside those. In HTML, you should recognize CDATA sections, including script elements. If they contain `<div`, it's not a tag. That's is, your 99.9% is probably too damn high.
- s_dev 2mo agoJust like jq there will be some lad along to tell us that "I don't like the syntax and find it confusing" not realising that's the exact superpower it presents is it's terseness is a key property to it's adoption. jq and regex really are sort of handy one liners that you invoke in other scripts and you explain what they do in your script with a comment.
- ourmandave 2mo agoEvery time I cut-n-paste a regex into code, I comment with the url of the spell book page I copied so future me can answer, "WTF does this do again?"
- bleuarff 2mo agoI'm wary of external urls in code. Some plaintext comment would come in handy for the day the link inevitably goes dead.
- brettermeier 2mo agoWhy could external URLs be a problem? And is it still one if you swap https to hxxps or something? What could go wrong with having a URL as a comment in code? I put URLs there sometimes and think it's very helpful.
- bleuarff 2mo agoYou don't control these external resources, and now the explanation of what your code does is tied to a site that could be taken down tomorrow, leaving you with a dead link and an unexplainable regex.
- RankingMember 2mo agoAgreed- I'm all for comments explaining a RegEx, but not hyperlinks in comments. The link inevitably goes dead and now you've left a helpful-looking present in a comment with dust inside.
- hn9rsvy2gx 2mo ago[dead]
- ape4 2mo agoThe last thing we need is huge regexs made by AI
- evilc00kie 2mo ago> The regular expression is very simple OT but this me-problem makes me angry every time I read it. Nothing is simple, otherwise it is trivial and not worth mentioning. I can't read over this without thinking that I'm not smart enough to wrap my head around something instantly.
- AussieWog93 2mo agoIt might just be a me problem, but I've always been wary of regexes. They're not too bad to write, but reading them back and understanding what's actually going on can get a bit hairy. Plus, all of the subtle differences between regex libraries seems like a bit of a footgun. Obviously they have their place, but I know a lot of the older guys seemed to love them way more than the young.
- rhdunn 2mo agoVarious libraries (e.g. Python's `re` library) support comments and whitespace as an option allowing you to format the regex on multiple lines with commenting to document what each part does. I'm not sure if there are any regex libraries that support DSLs and easy composability (e.g. the email RFC regex would be easier to read/maintain if you could specify the individual parts like are defined in the RFCs).
- AussieWog93 2mo agoI honestly never knew that, should give it another go.
- klibertp 2mo agoI would recommend trying something like PyParsing[1] instead. Libraries like this allow you to compose the parser from language-level entities (object and functions, on top of regex and string literals). This means you can attach comments to those entities naturally within the syntax of the language. You also get much better error reporting out of the box, as well as a well-defined way of attaching transforming code to parts of the parser. There's a place for simple regexes, but complex regex DSLs (with comments and non-significant whitespace, etc.) are almost always less convenient than simply using your language directly. [1] https://pyparsing-docs.readthedocs.io/en/latest/HowToUsePyparsing.html#hello-world https://pyparsing-docs.readthedocs.io/en/latest/HowToUsePypa...
- frizlab 2mo agoSwift even has a `RegexBuilder` DSL which makes writing regular expressions pure code and type-safe. Pretty amazing tbh
- isqueiros 2mo agoBefore the AI craze, I'd gotten quite good at writing regexes. Regexr was quite useful for decoding and composing them. I feel like they're going to become a lost art.
- jjice 2mo agoTotally agree. Selfishly, I was always the "regular expression" guy because they were a bit hobby space of mine (engine implementation and such), so seeing LLMs rip them is a bid of a bummer. Half the reason it's a bummer is because I've seen coworkers who don't know when a regular expression is very suboptimal performance wise, but the LLM has no problem spitting it out. Part of really understanding regular expressions is knowing when to not use them. The one that sticks in my head is when I was debugging some code that I was suspicious was causing our high memory consumption on a simple API service just to find out the regular expression was being used to strip a potential "data" front of a base64 encoded file (apparently someone thought we should do that instead of rejecting the payload). The regular expression scanned an entire base64 string that was up to 50 MB for the raw file, so about 66MB base64 encoded. I'll tell you what, replacing it with a loop over the first handful of characters solved all the problems. It should've never been a regular expression. If you see regular expressions as an archaic language that solve string problems, and now the magic box can make them for you, you're in for hell.
- isqueiros 2mo agoBeing the "regex guy" at work was also my thing. I was even in the process of making a regexr-like extension for vscode, but right about then everyone jumped ship to AI and making vscode extensions kinda felt like a last years thing. They really are a "tool for the job" type thing, and I've seen the abuses people put them through. The fact that we struggled to know when to reach for it before worries me that this will be exacerbated now that we don't even read our own code.
- ogogmad 2mo agoRegular expressions will always remain fundamental to computer science: They characterise all of those - and only those - conditions on bytestrings (or bitstrings, or Unicode strings, etc) which are checkable in constant memory.* In other words, they characterise the set of all "regular languages", which is a name for DSPACE(O(1)). Furthermore, regular expressions can be matched in O(n) time and O(1) memory, within a single left-to-right pass, which is the highest level of efficiency mathematically possible. Since they operate on bytestrings, they can be applied to computer memory and computer state itself, which are ultimately just bytestrings, and not just to text. To be fair, you might know all of that, but I wanted to highlight this. LLMs are a lot less efficient than regular expressions wherever both are applicable, simply because everything is less efficient than regular expressions. * By constant memory, I mean that the memory usage has a maximum value independent of the size or the contents of the input bytestring.
- trashb 2mo agoRegexes are great, they seem like magic when you use them right. They can solve your problems even if you don't use them right. Just make sure not to mix the flavors.
- klibertp 2mo agoWhile true in principle, writing grammars in regexes is problematic in practice: the syntax for the more advanced features (named submatches, lookahead, backreferences, etc.) is pretty complex, and refactoring the expression means you're working within a string literal, with no help whatsoever from your editor or IDE. My "go to" solution for parsing (and validating/matching) non-trivial grammars is a library that wraps regexes and allows you to structure the grammar with entities above substrings of a string literal (including arbitrary code for transformations). PyParsing for Python, scala-parser-combinators for Scala, Grammar in Raku, PetitParser in Smalltalk, PEGs in Janet, parser combinators in F#, and so on. These are mostly internal/embedded DSLs, which makes them much easier to use than the typical lexer/parser generators, while giving you all the power to structure and evolve the grammar easily. For simple grammars, a well-written library adds little overhead over plain regexes. However, grammars rarely stay simple - very often, during the course of development, you find edge cases or the need for extensions. If you started with a structured parser, you're fine: there are specific ways of evolving the grammar, and you can use normal refactoring tools to perform them. If you started with a regex, you quickly end up with a monster regex literal that becomes more brittle and harder to change with each modification. One important property I look for in parsing libraries is the support for left-recursion. Memoizing/packrat parser generators can handle it gracefully, which is important, because if I'm implementing a published grammar, I want to encode it as closely to the original as possible. For the same reason, I prefer having dedicated tools for associativity and precedence (so that I don't have to invent names for intermediate levels). TL;DR: yes, regexes are much more expressive than the "regular" in the name would imply, but they still have their limits. For parsing things, it's better to start with something that can work in the simple case fast (so no lex/yacc-style codegen from 2 separate external DSLs), but which also provides enough structure that adding good error handling, extending the grammar, attaching arbitrary code transformations, etc. won't be a big problem later.
- jakobnissen 2mo agoThis article misleads you by conflating regular expressions with specific implementations like PCRE, which also does non-regex string matches. Annoyingly, the article does a good job of explaining what a regex is and what the limitations of regex are relative to PCRE, so the author should understand that what they are talking about when they talk about NP-complete string matching is not regex, but PCRE-specific features. The distinction matters because regex absolutely can't match HTML, and because regex, unlike PCRE expressions, have guaranteed O(1) space and O(n) time complexity when matching a string of length n. When you use PCRE features for string matching, that may degrade to exponential time which makes it useless. For example, you can do denial of service PCRE attacks, but not denial of service regex attacks (unless you can query with some megabyte-large regex).
- jibal 2mo agoActually TFA is explicit about this: > Regular expressions in the formal grammar sense can (pretty much by definition) only parse regular grammars and nothing more. > But when programmers talk about “regular expressions” they aren’t talking about formal grammars. They are talking about the regular expression derivative which their language implements. And those regex implementations are only very slightly related to the original notion of regularity. > Any modern regex flavor can match a lot more than just regular languages. How much exactly, that’s what the rest of the article is about.
- skrebbel 2mo agoYou’re splitting hairs. The author is writing from the perspective of a PHP programmer (author is in fact a major PHP contributor), where the term “regex” has a single very clear definition, namely PHP’s PCRE-based implementation.
- jakobnissen 2mo agoNo, this is not splitting hairs. This is the author using the straight up wrong terminology. Regex can’t match HTML, and aren’t NP-complete. The fact that the author believes that “regex obviously means PCRE” is objectively wrong and misleading in the sense that all the things his article are about would have another conclusion if he actually talked about Regex. It’s like if there was a library called QuickSort which also included a SAT solver and I then wrote an article about how you can solve SAT-equivalent problems with quicksort (“in the programmer sense, which obviously means a SAT solver”)
- kwoff 2mo agoObligatory: https://stackoverflow.com/a/4234491 https://stackoverflow.com/a/4234491 (tchrist's "Oh Yes You Can Use Regexes to Parse HTML!", which refers to https://stackoverflow.com/a/1732454/459233 https://stackoverflow.com/a/1732454/459233 "TONY THE PONY, HE COMES" (I won't risk copy-pasting the "corrupted-looking" version, heh))
- Beestie 2mo agoomigosh - this is an internet classic that never gets old!
- jedisct1 2mo agoRelated: https://pcre-vera.github.io/pcre-vera/ https://pcre-vera.github.io/pcre-vera/ (WIP)
- zby 2mo agoIt's a pity that the Perl6 saga killed the new regex ideas - I wish the raku regexes were adopted elsewhere (like the Perl5 regexes were): https://docs.raku.org/language/regexes https://docs.raku.org/language/regexes
- wodenokoto 2mo agoThere was a time when everyone was like “you can’t parse html with regex, use beautiful soup!” And beautiful soup defaulted to a regex parser