12 ms·
Pomsky – A portable, modern regular expression language
- echelon 4y agoI don't know why I'd never previously considered regular expressions as being a compile/transpile target. It's pretty obvious from PL theory and makes a ton of sense. That said, after looking at this syntax, I'm not sure that this is much of an improvement. Maybe I've spent far too much time in Regex land [1], but I know I'd perform much slower in this. It's not particularly beautiful, either. The verbosity doesn't seem clearer. Variables and comments are great, though. We need to add them in future regexes. Overall, good idea. I'd like to see more takes on this. [1] https://jimbly.github.io/regex-crossword/ https://jimbly.github.io/regex-crossword/
- TuringTest 4y agoIt may be a good way to standardize regexp syntax for users of all levels of expertise. Every text editor, shell environment, programming language or desktop application seems to use regular expressions with a syntax which is slightly different to all the others, but not different enough to call it with a different name. This means that a newbie learning regular expressions will be thrown into an environment where it can learn the basic principles, but the rules it learns are not generalisable to all applications (e.g. do I match an arbitrary string with '*', or '.*' ? Can I reuse matched patterns with (1) or with {1}? Etc.) A new readable and easy-to-learn syntax that is nevertheless portable may work like markdown, as a simple-yet-universal new way to apply regexp that newcomers may learn with confidence and apply everywhere, replacing all the previous slightly incompatible versions.
- camtarn 4y agoMatching an arbitrary string with '*' is not a regular expression - it's a shell glob pattern. Different thing entirely. Maybe they should be the same thing, yes. And although I don't know for sure, reusing a matched pattern within the regex string using (1) seems like it would be very strange, because () are metacharacters for grouping.
- TuringTest 4y agoThey are valid patterns in MS Word and Notepad++ search and replace, though. (Word uses parentheses to create the group then \1 to reuse it in the replacement string, while Notepad uses $1). That's what I meant about every system using a different syntax for what's basically the same functionality.
- hnlmorg 4y agoI think that’s a fair summary. Interesting idea but the syntax adds nothing to the readability of regex. I’d be more impressed with this if it targeted multiple different regex targets, since not all implementations of regex are equal. But in its current state it has all the problems of regex plus now the problems of this new language on top. Looks a fun personal project though. Hope the developers enjoyed building it.
- jsnell 4y agoThose improvements are not really novel though. Perl regular expression have had variables, comments, non-significant whitespace basically forever.
- macintux 4y agoAfter learning Perl, nearly every language has been a disappointment when it comes to regexes and text handling in general. I couldn't believe it when Java was released with not only no regex capability in the language syntax, but no library either.
- dxbydt 4y agohttps://docs.oracle.com/javase/7/docs/api/java/util/regex/package-summary.html https://docs.oracle.com/javase/7/docs/api/java/util/regex/pa... pcre has always had java and scala support.
- zokier 4y agoJava Pattern != Pcre != Perl regex, they are all distinct implementations with own quirks
- macintux 4y agoDuring Sun’s initial Java 1.0 release roadshow I drove to Chicago to attend. At the time, java.util.regex didn’t exist. I don’t know how to identify when it appeared.
- Avshalom 4y agoSays at the bottom since 1.4 which would be 2002 according to wikipedia but the history of java version numbers seems weird and 1.4 might just be the last time the java.util.regex api changed.
- cutler 4y agoWorse, Java regexen still require metacharacters to be escaped despite the recent introduction of raw string literals. Would you include Ruby in your comparison given that it inherits its regex implementation largely from Perl despite Onigurama's slight differences?
- tmtvl 4y agoComments are a great idea, you're right! https://perldoc.perl.org/perlre#/x-and-/xx https://perldoc.perl.org/perlre#/x-and-/xx
- huqedato 4y agogreat achievement, but in real life what's its use?
- sfusato 4y agoIt's a language that compiles to regular expressions, aiming to make regexes easier to write and to maintain. It's compatible with many regex engines and polyfills Unicode support where necessary.
- AkshatJ27 4y ago> If you know RegExp's, the syntax will immediately make sense If I know RegEx, why would I use pomsky?
- DemocracyFTW2 4y agoIf you know RegEx, you have two problems.
- BtM909 4y agoIf you solve a problem with RegEx, you now have two problems!
- echelon 4y agoVariables and comments seem nice. Insignificant whitespace is needed to support it, and as an added bonus it would make it easier to break up patterns across multiple lines. The syntax changes and added verbosity do not seem great, though. They'd trip me up for sure. In general, I think I'd like to see a language more like "Regex 2.0", ie. an extension that doesn't depart too far from what we're used to.
- joshka 4y ago(?x) mode in regexes does comments / whitespace / multiple lines. +1 to the syntax
- mhitza 4y agoWhile a language like Raku (formerly known as Perl 6) is unlikely to catch up in the current landscape, it did bring a lot of improvements to regular expressions (link to section that starts to use interesting examples [1]). Way back somewhere in the 2010s when I was still keeping an eye on Perl 6, I was kind of hoping that all these improvements would make their way into some kind of PCRE v3 that all other tools that already use PCRE would switch to. Would have been nice. [1] https://docs.raku.org/language/regexes#Alternation:_|| https://docs.raku.org/language/regexes#Alternation:_||
- 4y ago
- zer00eyz 4y agoSome people, when confronted with a problem, think "I know, I'll use regular expressions." Now they have two problems. Im not so sure this solves the actual issue ;)!!!! Jokes aside, regex has its uses.
- rjh29 4y agoIt's such a bad take. I need to check if a string is a dash followed by two numbers. What's better, /^-\d{2}$/ or writing 10 lines of bespoke code to parse it?
- deleted 4y ago[deleted]
- xiphias2 4y agoNumber ranges look great, but it's much better to just extend + polyfill current regex syntax and keep compatibility. [:0-255:] could be an option for example to write a number range. Also regex / variable interpolation should be added to regexes in languages probably :)
- burntsushi 4y agoAuthor of Rust's regex crate here. > Also regex / variable interpolation should be added to regexes in languages probably I've often wondered and desired this myself. It feels like the single most useful change one might make, because it enables (albeit very basic) logical decomposition. It comes in handy when a regex works for your use case, but the regex is somewhat long but with many repeated parts. In my experience, that isn't too rare, and the regex readability would likely be helped quite a bit by some kind of decomposition.
- Someone 4y agoYou can do that by building up a string and converting that to a regex. Pseudo-code: let day = "[0-9]{2}" let month = "jan|feb|…|dec" let year = "[0-9]{4}" let regex = Regex(month + " " + day + ", " + year) Disadvantage is that, if your programming language has automatic syntax check of regular expressions, you probably lose that. The string concatenations also make things a bit slower, but that almost always needs to be done only once. Advantage is that this makes testing easier. For example, you can do let dayRegex = Regex(day) and then test that regex in isolation. That probably is overkill for this example, but if you change the regex to only accept numbers in [1,31], that can be useful. If you decide to have your regex to treat 30-day and 31 day months separately, things get more complicated. I wouldn’t go that far, though, and certainly not as far as handling leap years correctly. When regexes get complex, it’s time to go for a proper lexer/parser.
- burntsushi 4y agoOf course I'm aware of that technique and I've used it many times. :-) Doesn't change what I said.
- YesThatTom2 4y agoThis language is so… reasonable!
- tonnydourado 4y agoI have the impression that learning this would feel like learning Italian, after learning Spanish, and being a native French speaker. > If you know RegExp's, the syntax will immediately make sense Doesn't seem like a feature.
- mikepurvis 4y agoMaybe not a feature eventually, but absolutely during a transitional time, particularly if many of those who would consider such a transition are like me— they "know" regex, have written thousands of them, but even decades later continue to find them annoying, finicky, and hard to debug.
- deleted 4y ago[deleted]
- fegu 4y agoPomsky will look very similar to Pornsky in many fonts. quiet laugh
- yuvadam 4y agoYeah it definitely has a keming issue
- levesque 4y agoShould definitely find another name for it.
- geocrasher 4y agoCame here to say something similar. Definitely read it as "Pornsky" at first.
- thunderbong 4y agoI feel everyone should look at the examples before commenting! I don't know what the edge cases might be but this looks amazingly easy.
- deleted 4y ago[deleted]
- antman 4y agoFor spec review from business partners I have found "verbal expressions" the most useful flavor. Specifically because to review them you don't need to know regular expressions and the library has ports for most programming languages. Example: https://github.com/VerbalExpressions/JSVerbalExpressions#testing-if-we-have-a-valid-url https://github.com/VerbalExpressions/JSVerbalExpressions#tes...
- arthurofbabylon 4y agoCognition works in two ways: 1) highly deliberate, slow, patient and 2) spontaneous, quick, reactive. I believe software languages do a great job of leveraging both of these capacities of the minds of programmers. The first type of thinking is generally not composed in code. It happens staring out the window, or looking at data tables or lists and reading documentation, drawing out the architecture. The second type happens in code composition: programmers know what snippet will accomplish which outcome, and it just comes out when we know it's needed. Boom, we have the thing we imagined during the architecture phase. Maybe we look something up or slow down while writing code, but the composition is fairly reactive/spontaneous. It flows. So coding goes in this rhythm... thinking openly and patiently and deliberately, followed by a burst of code composition, and then more open/deliberate thought while staring out a window. Again and again. Maybe we reverse some of the quick/spontaneous thinking to rearchitect according to our slow/deliberate thought. Enter regular expressions. I have never, ever spontaneously composed a regex. I have to be very slow and deliberate. I have to recall the precise definition of every character in the regex as I read or write it. Sometimes I stare at a regex for 2 minutes, look up some operator, stare again for 2 minutes, refactor it. Finally I understand it and can now apply the regex spontaneously to whatever environment I'm working in. It always feels like first principles architecture, never like flow. This is odd, as regex is like a giant shorthand engine. Maybe I have not yet formed the neurological connections required for it to flow – but more likely this is a function of the language. There must be a syntax that better supports flow-state composition. (I don't think Pomsky accomplishes it.)
- pygy_ 4y agoThe regular formalism is, fundamentally, about composition, and the current syntax doesn't do it justice. That syntax was devised as a math notation, where sub-patterns were abstracted as one letter variables. You didn't have to deal with complex patterns at all, and it was very readable for that use case. It was then re-purposed as a write-only language for searching at the CLI (in `ed` and its descendants, then grep). They gradually graduated into what they are today in general programming languages because they were familiar and worse is better... However, composition got lost in the process. Swift takes a radically different approach, and gives RegExps a parser combinator syntax. So every sub-expression can be assigned to a variable and reused/tested independently. Here's an example (sorry I don't have it at hand in text form): https://pbs.twimg.com/media/FipgjUdUcAAl1xS?format=jpg&name=large https://pbs.twimg.com/media/FipgjUdUcAAl1xS?format=jpg&name=... I wrote a lib that does the same in JS (the lib predates the Swift syntax by years). Here's the same example ported to JS: https://flems.io/#0=N4IgtglgJlA2CmIBcBWAzAOgAwCYA0IAzgMYBOA9rLMgNoCMAnFhinlgLoEBmEChtoAHYBDMIiQgMACwAuYagWLlBM+CuQgAPFLQA+ADqCABEYAqU+EYBK8AOYBRAB4AHJIc0B6HbpAFC8BGIZCGV+CTokOjQQAF88IVFxSQArfkVlVXUJCDBnclIZI2BBchkAQQthKDwjETEoAGFhZxkAV1J4Gv8AR1a1Yk6jQlauHkcYoy4KMCMAciVc8n8AWg7beBdZw0MlQUJCgC94CgB5UgBZfMsAXiGRsYAKWYAqWYBKHdDCsGEATwAjG53UYQRxPAD870+e0KyngZ0uHSMt2GILBswA1FDBLt9kY4QirjZYK0gsIVMjgY9MZCPjivkYAHLwADuABkIIIgT0+jj4A8fgD+Vt9KR3jURYJsbjCsz2Zz4ABleC9fpAo6nC5XB5yjlcukyoxlQS-RnKXUKykeDAeVrbekwowAVVIsEpPLVT1kMmcsxqgsBT0I4rmSA8Hj9+K5hI6D2NpvNrL18DedOheIAbtB4OQmlR3SreQMHoYTB4AHoePClubLOvLB4ANWz5CMedgb3ryz9NYtXOVqr51eMRgD-JrJlmNB7I5MdXgjWabVjs3nkYJWo6xNJMnJMjjJrNgj7Kbew5Mk-YM4vTKTCoHRc6NbPNfni5a7WF7VgkZdHfPt7yv2hZqgBsz1tckH1teRgeAAJB4hhpg6eI-DIxBSJSWZQDm7YYBs8DECWggAAaGF2TYtm2whUJ29aGDQ7ZGOc8AAITsIY3rOIQYYeMQNGwPhjiiM4CAYAs5F1pB1z0aRGAyKQOQPHSyEQFwRgCsI6FSG8RQ1oaxSJDU34TLcaEYRgtgUK03H6aElDwBgsDkLYDzAN+NTzjEulGIYMT2oazhIrcUDkMQrRiCo4kdFp8D2AgkX7rMQXwNiKXyRsMgNBkaiFLc2G4QJhiheFiUYP85BQL8GDNM4ahQA8KVvL4RABIRwShBodA4EgOAAOyxPEIDzho4mEGkIC7JkMgaLE7AxEAA https://flems.io/#0=N4IgtglgJlA2CmIBcBWAzAOgAwCYA0IAzgMYBOA9... Here's the lib: https://github.com/compose-regexp/compose-regexp.js https://github.com/compose-regexp/compose-regexp.js
- asicsp 4y agoPrevious discussion: "Rulex – A new, portable, regular expression language" https://news.ycombinator.com/item?id=31690878 https://news.ycombinator.com/item?id=31690878 (242 points | 6 months ago | 189 comments)
- luuuzeta 4y agoNamed regexes (variables in Pomsky) remind me of Raku [1], which implements an improved flavor of PCRE regexes plus grammars in general as part of the language. [1]: https://docs.raku.org/language/grammars https://docs.raku.org/language/grammars
- jjice 4y agoThe grammars as part of the language is probably the most interesting thing about Raku to me. Never seen that in another language as a core concept like that.
- lizmat 4y agoFor those interested: https://docs.raku.org/language/grammar_tutorial https://docs.raku.org/language/grammar_tutorial
- jll29 4y agoRelated projects: 1. - Xerox xfst 2. - Xerox lexc 3. - Xerox twolc 4. - Mans Hulden's FOMA Repo: https://fomafst.github.io/ https://fomafst.github.io/ Paper: https://dingo.sbs.arizona.edu/~mhulden/hulden_foma_2009.pdf https://dingo.sbs.arizona.edu/~mhulden/hulden_foma_2009.pdf Demo: https://dsacl3-2018.github.io/xfst-demo/ https://dsacl3-2018.github.io/xfst-demo/ Tutorial: https://foma.sourceforge.net/lrec2010/ https://foma.sourceforge.net/lrec2010/ 5. Helsinki HfstXfst: Homepage: https://github.com/hfst/hfst/wiki/HfstXfst https://github.com/hfst/hfst/wiki/HfstXfst These tools go back to research by Lauri Karttunen and others at Xerox Research Center Europe in Grenoble, where an attempt was made to create highly efficient compilers and runtime libraries for finite-state transducers, i.e. to move beyond regular expressions to regular RELATIONS. This not only permits to formalize replacements (regular expressions with an "output tape"), but also creates reversible automata (input and output roles can be swapped) and leads to a domain specific language that describes transducers in very readable ways, including sub-automata naming, so that it can be useful for formal specification or linguistic rules (phonology, morphology, i.e. word or sound grammar). The latter two projects are open source clones of the former effort. Once you have used these for a week, you will never want to get back to ugly "ordinary" regexes again. Books: (a) https://www.amazon.co.uk/Finite-State-Processing-Synthesis-Lectures-Technologies/dp/163639115X https://www.amazon.co.uk/Finite-State-Processing-Synthesis-L... (b) https://www.amazon.co.uk/Recognition-Algorithms-Finite-State-Transducers-Processing/dp/1608454738 https://www.amazon.co.uk/Recognition-Algorithms-Finite-State... (c) https://www.amazon.co.uk/Finite-State-Techniques-Transducers-Bimachines-Theoretical/dp/1108485413 https://www.amazon.co.uk/Finite-State-Techniques-Transducers... (d) https://www.amazon.co.uk/Finite-state-Language-Processing-Speech-Communication/dp/0262181827/ https://www.amazon.co.uk/Finite-state-Language-Processing-Sp... (e) https://www.amazon.co.uk/Finite-State-Morphology-CSLI-Computational/dp/1575864347/ https://www.amazon.co.uk/Finite-State-Morphology-CSLI-Comput...
- jeltz 4y agoCan't say I see the point. It is uglier and more verbose than regexes without being easier to read.
- deleted 4y ago[deleted]
- g5095 4y agoShow me a pomsky for matching any valid email pls.
- Intermernet 4y agoMatching email addresses with any regular expression is fraught with errors. It can be done, depending on which RFC you are checking against, but in the real world it will eventually cause a problem. You're better off using an actual parser (which will probably implement a state machine), writing a state machine yourself, or being overly accepting of invalid email addresses and just relying on attempting to deliver an email to the address. See http://cubicspot.blogspot.com/2012/06/correct-way-to-validate-e-mail-address.html http://cubicspot.blogspot.com/2012/06/correct-way-to-validat... for more discussion.
- deleted 4y ago[deleted]
- moreati 4y agoSome regex engines/dialects also allow pre-defined named subroutines (variables) with `(?(DEFINE) (?<name>pattern)...)` http://www.rexegg.com/regex-disambiguation.html#define http://www.rexegg.com/regex-disambiguation.html#define, but it's a very niche feature. e.g. https://gist.github.com/moreati/9d974e5395829d737dc342715f15fc56 https://gist.github.com/moreati/9d974e5395829d737dc342715f15...
- vintermann 4y agoIt's probably better than regular expressions. However, is it enough better that it's worth learning yet another syntax? Well, maybe. What I REALLY like about this one, is that it fully reverses the quoting/escaping assumptions. The default assumption of old regexes that symbols should by default match themselves, I think is regexes' million dollar mistake. The literal matches are the least interesting part of regexes. If you reach for regexes, it's because you want something more complex than literal matches, and the syntax should be about taming that complexity. Even the Unix world sort of conceded that, in reversing the quoting assumption for ?, +, (), [] and | in egrep. The mistake was in stopping there. I will take a good look at this. I hope they provide good justifications for their choices.
- coffeeblack 4y agoOn first glance, I didn’t see anything better, tbh.
- SCLeo 4y agoI think the number ranges functionality is pretty handy. In addition, I also like how it supports comments. (However, I do think if the regex grows to needing comments, it is probably too complex.)
- cmovq 4y agoThe literal matches are useful in an editor search context where most of the time you want to search for the literal string.
- chongli 4y agoYes and editor search is one of the few reasons to use regex these days. If you're writing something more permanent then you should use a real parser library!
- nlitened 4y agoAdd a new library dependency when built-in one-liner regex does just fine? I don’t think that’s a good idea by any measure. Also, learning regex syntax has much better ROI and is a much more transferable skill, than learning a specific parser library’s syntax and bugs. Just invest 30 minutes once in your life into learning basic regex, for Christ’s sake, it’s not that hard!
- blindseer 4y agoI find something like this a lot more readable: https://github.com/jkrumbiegel/ReadableRegex.jl https://github.com/jkrumbiegel/ReadableRegex.jl It is in Julia, but if you have it installed locally it’s just a few taps away. You can even generate the regex, and use that in Python and just add the ReadableRegex in a comment nearby.
- dr_kiszonka 4y agoNot a Julia user, but this looks great!
- gaganyaan 4y agoThat reads better than regex, but at that point, why not just use parser combinators? At least for me, whenever I want something more complex than a basic regex, I go for a parser combinator library. Maybe that's what's happening under the hood anyways? I don't know Julia well enough to know.
- Alifatisk 4y agoWow this looks really cool, do I have to learn julia though?
- raydiatian 4y agoCan anybody explain how “any old” regex engine can just accept this new syntax? Are we transpiling to regex?
- camtarn 4y agoYes.
- ngomile 4y agoWould be nice if it was possible for the Pomsky playground to show informative modal boxes when you hover over some of the Pomsky style expressions. Kind of like RegExr which remains my favorite tool for quickly writing out a regular expression and seeing what its outcome will be. Very nice to be able to quickly see what some thing does and how it affects your query including documentation in an easily accessible location on the same page. Not having to navigate to a different page would be a boon for the playground. Very interesting though, I use Regular Expressions for the times when you need to extract information but find and replace functions/methods just aren't enough. Mostly for my scrapers.
- neilv 4y agoOlin Shivers defined the related "SRE" S-expression language for regular expressions: https://scsh.net/docu/html/man-Z-H-7.html https://scsh.net/docu/html/man-Z-H-7.html https://www.ccs.neu.edu/home/shivers/papers/sre.txt https://www.ccs.neu.edu/home/shivers/papers/sre.txt It works nicely with Scheme, including for programmatic generation. (I've done some regular expressions from heck, without the benefit of this, and needed extensive commenting just to keep a few/several levels of nested groupings straight. With S-expressions, that's trivial.)
- maweki 4y agoThat doesn't solve the most problematic part of regexing for me: having a different programming language embedded as a string in my other source code. The same happens with SQL in many codebases. The embedded string is often precluded from static analysis. Fieldnames or names of capture groups may drift apart from the outside code and maybe nobody notices. What we need are more embedded DSLs where we get static syntax and type checks and hopefully a good integration between the surrounding code and the embedded code.
- deleted 4y ago[deleted]
- Jorengarenar 4y ago>problematic part of regexing for me: having a different programming language embedded as a string While there is nothing stating programming languages need to be Turing-Complete (and there are in fact examples of non-TC programming languages, if only educational), I think calling regular expressions* a "programming language" may be a bit of a stretch. * Even when its capacities are as extended outside the mathematical definition as modern RegEx is. edit: specified I meant "non-TC programming languages"
- maweki 4y ago> and there are in fact examples of non-TC languages, if only educational Sure, SQL is only educational or not programming. There seems to be a lot of programming going on in the world generating a lot of value which you wouldn't recognize as such. Excel uses a terminating (and therefore not computationally complete) dataflow-based evaluation model but it's in no case programming. Are you sure? I'd think, telling a computer "if you recognize a pattern described as such, replace it with so and so", just regex with replacement, sounds very much like programming to me. Very weird where you draw the line.
- JadeNB 4y ago> Sure, SQL is only educational or not programming. > There seems to be a lot of programming going on in the world generating a lot of value which you wouldn't recognize as such. Excel uses a terminating (and therefore not computationally complete) dataflow-based evaluation model but it's in no case programming. Are you sure? I think this is not the most charitable reading of your parent. You're taking it as if they definitively rejected SQL and Excel as programming languages and arguing from that, rather than that maybe they didn't think of them. I think it would have been easy to make the case for SQL, Excel, and even (ir)regexes as programming languages without taking a sarcastic tone. For example: > I'd think, telling a computer "if you recognize a pattern described as such, replace it with so and so", just regex with replacement, sounds very much like programming to me. Very weird where you draw the line. could easily have asked instead, where do you draw the line?—which would surely be a more interesting discussion.
- sulami 4y agoI had a similar kind of idea for a long time, which I put into action a few weeks ago via a standalone transpiler of Emacs' rx macro to common regexp syntaxes.[0] I ended up getting interrupted and didn't completely finish it, but it generally works, though is probably riddled with edge cases. The basic idea of rx is to use S-expressions to describe regular expressions, and my elevator pitch would've been to embed rx invocations in shell scripts using $(syntax), the main use case being something like sed invocations. I still think it's a neat idea, and complex regular expressions tend to be hard to parse for humans. [0]: https://github.com/sulami/rx https://github.com/sulami/rx
- zokier 4y agoWhile I'm not sure about Pomsky specifically, I do think its nice that people explore the language space for regex more. General programming languages have huge variety of styles and syntaxes available, from APL to Haskell and Lisp, whereas regexes are pretty much the same everywhere. It feels like we stuck with the first thing that Kleene, Thompson et al thought of and for 50 years didn't even really try anything else.
- gamesbrainiac 4y agoThis was previously called Rulex. Glad to see it is getting traction :) I made a video on it a few months ago. https://www.youtube.com/watch?v=nPjCxwEdIIo https://www.youtube.com/watch?v=nPjCxwEdIIo
- deleted 4y ago[deleted]
- rekado 4y agoFor Scheme I enjoy using Irregex: http://synthcode.com/scheme/irregex/ http://synthcode.com/scheme/irregex/ The key benefit is to have a proper DSL instead of a DSL hidden away inside an untyped string.
- mekster 4y agoWhy don't languages have "grok" patterns in their standard libraries? It seems to only exist in log parsing ecosystems but this really helps with getting rid of little bugs and wrong parsing of specific regex patterns. Instead of doing "^\d+(\.\d+){3}$" for IP checking which is clearly wrong, you'd do "%{IPV4:ip}" which is so much better. List of known patterns : https://github.com/hpcugent/logstash-patterns/blob/master/files/grok-patterns https://github.com/hpcugent/logstash-patterns/blob/master/fi... Even for PHP a third party library only has 15 stars.
- pimlottc 4y agoPerhaps because many devs, like me, haven't heard of "grok patterns"[0] before. Because of that, it took me a while to understand what your post was saying, since I was reading "grok" as a normal word. It's not a bad idea, though you'd want to make it a more formal standard; since it's just used internally in one project, it may be subject to change based on that project's own needs. There could also be more documentation, showing examples of strings that are accepted and not accepted by each pattern, as well as advice for generating compliant strings. 0: https://logz.io/blog/logstash-grok/ https://logz.io/blog/logstash-grok/
- mekster 4y agoIt may have originated from a project but many other log parser projects such as Vector and fluentd have such support. https://vector.dev/docs/reference/vrl/examples/#parse_grok https://vector.dev/docs/reference/vrl/examples/#parse_grok https://github.com/fluent/fluent-plugin-grok-parser https://github.com/fluent/fluent-plugin-grok-parser
- andy81 4y ago.NET has something close in standard libraries, in that you can try converting a string to System.Net.IPAddress and other similar classes. It's for parsing rather than search though, you'd still need regex if e.g. searching unstructured logs for IPs. e.g. (powershell syntax for easy reproduction) returns parsed, or throw if invalid. [uri]"\\ServerName1\Folder\UserName" [version]'1.2.3.4' [ipaddress]'1.1.1.1'
- turnsout 4y agoTo me, the lead example on the homepage ("Basic") is a major red flag. This is not clearer than a traditional regular expression: 'Hello' ' ' Did you count the single quotes correctly? Even with syntax highlighting, people WILL mess this up.
- kstenerud 4y agoThe fundamental problem comes from assigning meaning to whitespace (in this case, concatenation). I had the same issues when developing KBNF ( https://github.com/kstenerud/kbnf/blob/master/kbnf.md https://github.com/kstenerud/kbnf/blob/master/kbnf.md ) which operates in a closely related space (grammars). In early development, I took a number of cues from existing work that turned out to be bad ideas, in particular using whitespace for concatenation (which all BNF dialects seem to do). Switching to '&' for concatenation (reading it as "x and then y") made things a lot clearer, as it would also do for Pomsky: 'Hello' & ' '+ & ('world' | 'pomsky')
- turnsout 4y agoYes, the ampersands are far, far more comprehensible. Also, as a binary format enthusiast (I built a binary "Codable" implementation for Swift), KBNF looks great! I'll be following your progress.
- cutler 4y agoI will never understand why a regex in Java still reqires metacharacters to be escaped despite the recent addition of raw string literals. Talk about stuck in the Stone Age.
- vanderZwan 4y agoThe idea of a compile-to-regex language is a neat one, that immediately makes it a lot easier to use in existing projects. If there's any interest in other takes on "better regexps", the Rosie Pattern Language has some neat ideas: [0] https://rosie-lang.org/index.html https://rosie-lang.org/index.html [1] https://www.youtube.com/watch?v=MkTiYDrb0zg&list=PLcGKfGEEONaBUdko326yL6ags8C_SYgqH&index=12 https://www.youtube.com/watch?v=MkTiYDrb0zg&list=PLcGKfGEEON...
- chubot 4y agoFWIW here is a list of other such projects: https://github.com/oilshell/oil/wiki/Alternative-Regex-Syntax https://github.com/oilshell/oil/wiki/Alternative-Regex-Synta... Feel free to add other projects, including the ones in this thread
- osigurdson 4y agoI would have done ‘Hello’ (‘World’ | ‘Pomsky’ | NULL) In order to avoid ?: stuff.
- osigurdson 4y agoI seem to need regular expressions about once every six months. Every time I do I wonder how this horrible little language became ubiquitous. However, I’m not going to pull in a dependency just to avoid it (at my current regex cadence at in any case).
- rhapsodic 4y ago[dead]
- biztos 4y agoI once worked on a product that was heavily invested in regular expressions, and fairly non-technical users generated more, often hundreds per day. Of course this led to a certain amount of UI whack-a-mole: people learn a mini-language really fast when they use it all day at their jobs, and people are creative; whereas the "computer people" really needed the system to not grind to a halt because of a massively inefficient regex. From this I learned to see regexes everywhere when I look at text; and that we should always consider the "threat model" of our own users being good at their jobs.
- account-5 4y agoI like regex, the syntax is succinct, powerful, and easy to learn. I learned regex before I learned any proper programming language, because you could use it in a text editor. I remember feeling like a wizard at the time. I do admit it can be a write once hopefully never need to read again thing, but I still love it.
- magic_hamster 4y agoThe only solution I've found for regex which is 100% compatible with the full capabilities, and is also portable to any language so you can learn just one syntax, is regex.
- Alifatisk 4y agoI couldn't find how you turn off/on case-sensetive?