9 ms·
The unreasonable effectiveness of f-strings and re.verbose
- throwaway81523 4y agoNow you have two problems. Or in this case, probably more than two.
- genericlemon24 4y agoAn infinity! https://xkcd.com/1313/ https://xkcd.com/1313/
- stavros 4y agoOkay wait a minute, what's that about the subtitles? Isn't that too small of a regex to accurately classify all the subtitles?
- wartijn_ 4y agoAccording to this site[0] it's about the name of the movies, not a about the subtitles of all the dialog in the movies. For example the title is "Star Wars", the subtitle is "The Empire Strikes Back" [0] https://www.explainxkcd.com/wiki/index.php/1313:_Regex_Golf https://www.explainxkcd.com/wiki/index.php/1313:_Regex_Golf
- stavros 4y agoAhh OK, that makes much more sense, thanks!
- Linda703 4y ago[dead]
- Waterluvian 4y agoOh ho ho ho… this is good. Anything that breaks apart regexes to make them easier to read and comment each logical unit is worth the extra lines and syntax. This is killer.
- bawolff 4y agoThere have been many attempts at doing this over the decades. It never catches on. I don't think its something developers really want.
- schoen 4y agoIt looks like this particular one has had some staying power, because it (or the PCRE version mentioned elsewhere in the comments) have been rather widely implemented. I didn't know about this and was a happy regular expression user without it, but this looks like a good feature for the specific use case of wanting other people to understand the structure of your regular expressions. And much more portable than I would have expected.
- semiquaver 4y agoStrongly disagree. I’ve seen this pattern used heavily inside many ruby codebases, it’s incredibly useful for making regular expressions readable(ish)
- NavinF 4y agoI’ve been doing this for years. Splitting up regexes and reusing sub-patterns is very common in new code.
- atoav 4y agoI am a developer and I really want it. For the next longr regex I am going to use it for sure, as I like to use f-strings extensively already.
- perceptronas 4y agoI suggest you take a look at parser combinators. They are quite readable compared to regexes
- recursive 4y agoThis looks like it's approaching the logical neighborhood where parser combinators live. I'm a fan of parser combination.
- dgl 4y agoIt's pretty much exactly what Perl 6 / Raku grammars are: https://docs.raku.org/language/grammars https://docs.raku.org/language/grammars
- smegsicle 4y agoi can't find it now, but i think there was a larry wall quip along the lines of, "from all the concepts to borrow from perl, why did python take regex?"
- arnsholt 4y agoIt looks similar, but the semantics are quite different. A Raku grammar is a recursive deacent parser, but this is still a regular expression in the end.
- tomatowurst 4y agothis is a game changer. i did not know f-string could be used like this, i was largely satisfied with being able to not have to use the awful 2.7 era percentage symbols. even more reason to love python, right now for me slots and dataclasses is my new obsession. there was a great article that was posted here that went into details about 3.8 and up that featured all these great python hacks
- KerrAvon 4y agoGood god. Your mind would be absolutely blown by what you can do in Ruby.
- tomatowurst 4y agowas that snark necessary? this article is about python. i have no interest in ruby
- DiggyJohnson 4y agoI read it as good natured humor - clearly dramatic. Just sharing my opinion, fwiw, I see how you might interpret it differently.
- agumonkey 4y agohaving interests in other languages is important
- anitil 4y agoWhat can you do in Ruby?
- tempest_ 4y agoAha I think you mean Rails, there can't be anyone left using it for anything outside of that.
- arthurcolle 4y agoI started writing a reply to your comment in an attempt to dissuade you from such heretical propaganda but then I decided to just hotlink this: https://imgs.xkcd.com/comics/duty_calls.png https://imgs.xkcd.com/comics/duty_calls.png
- girvo 4y agoThis reminds me quite a lot of the "pegs" module in Nim. https://nim-lang.org/docs/pegs.html https://nim-lang.org/docs/pegs.html identifier <- [A-Za-z][A-Za-z0-9_]* charsetchar <- "\\" . / [^\]] charset <- "[" "^"? (charsetchar ("-" charsetchar)?)+ "]" Thats a small snippet of how it is used. It's been one of my favourite parts of using Nim to be honest. The fact one can get similar ergonomics this way in straight Python is wonderful! I'm definitely going to leverage this. I've done similar in other languages, but it's never felt quite right. re.VERBOSE is also handy to know.
- WaxProlix 4y agoYou might already be aware of this, but 'pegs' likely refers to Parsing Expression Grammars [1], a super powerful and imho very chill concept which translates into great tooling in lots of languages. 1 https://en.wikipedia.org/wiki/Parsing_expression_grammar https://en.wikipedia.org/wiki/Parsing_expression_grammar
- girvo 4y agoYou surmised right, though I adore Nim's particular implementation of it compared to the times I've attempted it in, say, Javascript :) Parsing expression grammars are easily one of my favourite tools. Honestly, I find them superior to regexes.
- ogogmad 4y agoNever used them, but don't they have memory usage in proportion to the string being parsed? Regexes, LL and LR don't have this shortcoming. This should surely constrain their applications.
- abecedarius 4y agoAn f-string evaluates to a string and not to an object such as a compiled regex. For this there are tagged template literals in Javascript (which got them from E). Example: https://github.com/erights/quasiParserGenerator https://github.com/erights/quasiParserGenerator
- pdonis 4y agoSo you either call re.compile on it or, as in the example in the article, you call one of the re module's functions that takes a pattern string as an argument.
- abecedarius 4y agoParsing strings assembled out of strings is classically bug-prone. In the alternative I'm pointing out, you don't fill the holes with other strings, you fill them with already-parsed regex objects. I think it's a shame Python didn't follow this design, which predated f-strings.
- pdonis 4y agoCan you give an example of a regex engine that has the design you describe?
- abecedarius 4y agoI must have a serious bug in my writing about this (sorry), because this was never about regex engines -- it's about literals and domain-specific sublanguages in general. Composing DSL programs by string concatenation is such a famous source of security bugs you see it in top-10 lists. I linked to the very similar example of a PEG-parsing DSL. But any regex engine that can work with a parse tree shows the same principle, e.g. https://edicl.github.io/cl-ppcre/#create-scanner2 https://edicl.github.io/cl-ppcre/#create-scanner2
- pdonis 4y ago
- dahart 4y agoCan’t named pattern groups do the same thing the f-strings do here? (As a bonus named patterns work in old Python 2 code, which yes nobody should be using anymore, but just sayin’) I don’t have an opinion or even good mental model about the advantages of either choice, but I was pretty excited to learn about named patterns and promptly used it to make a hacky (but interesting to me) small parser for identifying tokens and keywords and operators with differing precedence.
- Izkata 4y ago> As a bonus named patterns work in old Python 2 code So does string formatting. You don't need f-strings for this pattern to work.
- webstrand 4y agoPCRE has builtin support for this kind of factoring, too: (?(DEFINE) (?<code> [A-Z]*H # prefix \d+ # digits [a-z]* # suffix ) (?<multicode> (?: \( \s* )? # maybe open paren and maybe space (?&code) # one code (?: \s* \+ \s* (?&code) )* # maybe followed by other codes, plus-separated (?: \s* [\):+] )? # maybe space and maybe close paren or colon or plus ) ) ( (?&multicode) ) # code (capture) ( .*? ) # message (capture): everything ... (?= # ... up to (but excluding) ... (?&multicode) # ... the next code (?! [^\w\s] ) # (but not when followed by punctuation) | $ # ... or the end )
- asicsp 4y ago`(?N)` where `N` is group number and `(?&name)` where `name` is named group are known as subexpression calls. The third-party `regex` module (https://pypi.org/project/regex/ https://pypi.org/project/regex/) supports this and more such PCRE features.
- ars 4y agoAnd because the PCRE library is integrated in a huge number of languages (it's almost hard to find a language that doesn't have it - I'm looking at you JavaScript), these types of REGEXs are actually widely available.
- riffraff 4y agoAfair, so does onigmo/oniguruma, with a mildly different syntax
- genericlemon24 4y agoI mention PCRE in passing at the end of the article, but I didn't know about (?(DEFINE)...); that's very, very cool!
- jhgb 4y agoIf the regular expression engine accepted tree structures instead of just strings, you could have first class definitions of fragments of regular expressions. Even better, you could define them as functions, so you could have parameterized fragments. So then you could just apply something like http://edicl.github.io/cl-ppcre/#create-scanner2 http://edicl.github.io/cl-ppcre/#create-scanner2 on the resulting expression tree without having to use the bizarre definition syntax above.
- amp108 4y agoThis seems like a weird, and possibly untrustworthy, hack. Does Python not have the equivalent of Ruby's `x` modifier?
- User23 4y agoThis looks like perlre /x.
- kwertyoowiyop 4y ago“Unreasonable effectiveness” in titles considered harmful.
- SeanLuke 4y agoI don't want to be that guy, but why in the world are f-strings (formatted string literals) called literals? They are clearly dynamically calculated expressions.
- digisign 4y agoThey're a weird combo of compile-time parsing and run-time expressions, with custom opcode. Early on they didn't have expressions but worked like .format(). Then they added expressions and the PEP title needed differentiation from it. Not entirely accurate title is now set in stone. I would have called them "interpolated strings" or even e-string but the f-string moniker had already caught on and there was no stopping it.
- zarzavat 4y agoThere’s two definitions of “literal” in widespread use. The first definition is as you say: an expression that has a constant value. The second definition is: an expression that is the primary syntactic form to construct a type. For example: “array literals” construct arrays, but may contain arbitrary expressions within. The first definition is more common in low-level languages where there is a place in the compiled executable to put constant data. These languages might call the second form an initializer rather than a literal. But in a dynamic language such as Python the distinction is less important.
- SeanLuke 4y agoLisp uses the first definition. And it's as dynamic as they come. Java also uses the first definition. I suspect a more proper description would be: #1 is the correct definition, formally used for 80 years now. #2 is incorrect, and is being abused by people in JS and Python who should know better.
- wodenokoto 4y agoThe real kicker, hidden between everything, is that you can combine f-strings and r-strings. fr”this is both an f- and an r-string” I had no idea. I wish Python allowed for custom string types. I would love a sql string type if for nothin else than to show my code editor how to highlight inside the string!
- digisign 4y agoThose quotes should be "ascii" quotes, not unicode.
- klodolph 4y agoThat’s not true, though! It’s written for humans to read, not for machines to parse, and any human reading this will realize what they’re supposed to be.
- _han 4y agoCode is very often written to be copied and pasted though! And you could be surprised by the amount of people not noticing the difference between the quotes.
- f1refly 4y agoThose people will notice when their compiler complains and will hopefully know better than to copy+paste something from an untrusted website next time.
- klodolph 4y agoThis code is very obviously just illustrating a point to the people reading it, it seems unlikely that anyone would want to copy and paste it. Lots of code snippets are incorrect code. For example, I often write C code like this: int x = ...; The line contains a syntax error, but I’m communicating to the people reading that x is initialized to some value.
- tokamak-teapot 4y ago
- digisign 4y agoIf you have a number of curly braces in the pattern, probably easier to use printf-style formatting to build the pattern, with %s etc.
- poleguy 4y ago99 times out of 100 when I think I might need a regular expression, I find it far better to code the search in python directly rather than using the regular expression engine at all. It's far easier to understand and you can run a regular debugger on it and use regular comments. In the 100'th case, I'll code most of the expression in straight python and a very small piece using the regular expression engine. By straight python I mean things like 'for', 'split', 'startswith', 'find', and regular character indexing. So for me this post is a solution to a problem that I just avoid.
- IshKebab 4y agoI mean it's definitely better, but if you end up with a regex that long you really shouldn't be using regex.
- kbob 4y agoDo sources still exist for Perl 4 and earlier? I decided to go looking for when Perl got verbose regex. The oldest thing I can find on perldoc.perl.org or cpan.org is 5.004 (1997), where they were an existing feature. EDIT: Found 4.036 sources (1993). A quick scan of the man page (troff source!) does not find verbose regular expressions. So it looks like they were introduced very early in the Perl 5 series. https://www.cpan.org/src/unsupported/4.036/ https://www.cpan.org/src/unsupported/4.036/
- crabbone 4y agoProgrammers adding features to languages the same way horror movie characters decide on entering an abandoned shack in the woods. No. f-strings are an awful idea. And combining them with r-strings is yet another awful idea. Please, never do that. I'm also not a big fan of extended / verbose regular expressions because it creates ambiguity in interpretation (a slightly different language to define regular expressions). It's a bad solution to the problem of building longer expressions, which should've been addressed in a different way: by making the language of regular expressions more modular, not through allowing more hard-to-interpret language details.
- ubq323 4y ago> No. f-strings are an awful idea. And combining them with r-strings is yet another awful idea. why?
- LynxInLA 4y ago“Here's the plan: When someone uses a feature you don't understand, simply shoot them. This is easier than learning something new, and before too long the only living coders will be writing in an easily understood, tiny subset of Python 0.9.6 <wink>.” ― Tim Peters
- sam2426679 4y agoI do a similar thing as suggested by the article, except by using Python's "concat between parenthesis" strings instead of Python's heredoc strings. The advantage of doing it this way is that there are no caveats (as mentioned in the article) with needing to unexpectedly escape certain characters. It looks like this: pattern = ( r'[A-Z]*H' # prefix r'\d+' # digits r"[a-z']*" # suffix ) No funky stuff with escapes, and you can indent to your heart's content.