3 ms·
Ah! I was only referring to ERE as specced by POSIX, yes. > It's not difficult because getting it correct is difficult, it's difficult because making it corre
by dan-dev 4y ago
Ah! I was only referring to ERE as specced by POSIX, yes.
> It's not difficult because getting it correct is difficult, it's difficult because making it correct and fast is not straight-forward.
It seems to me that you take my project as a proxy for whether regex derivations are a feasible way to deal with unicode-ready regexes that also support complements and so on. That was not my intention, I was merely attempting to show an easy implementation of regex derivations that can deal with unicode and can be extended to support complements. With this project and the paper I linked, it seems to me that answering whether this particular kind of constructing regex DFAs is a possible way to achieve what you are looking for or not should be rather straightforward.
(To my last knowledge, a regex complement is not easy to add in the presence of extra features like backtracing and lookahead.)
- burntsushi 4y ago> With this project and the paper I linked, it seems to me that answering whether this particular kind of constructing regex DFAs is a possible way or not should be rather straightforward. Yes, you're correct. It just isn't that interesting because it will fall over in any kind of real practical usage. :-) Full disclosure, I'm the author of ripgrep, so I have some particularly relevant experience in the specific domain of making regexes work well for the masses. (It also forms my bias with what I care about.) Like I said, I wasn't necessarily trying to pick on you. There are others in this thread that are seemingly ignoring the engineering challenges, and only looking at the theory. Then using that as a basis to proclaim that regex engines are "stuck" in the 80s/90s. Consider my perspective here: if someone pipes up and says, "yeah hey it is actually easy to add those kinds of features to regex engines." Well, then, it's important to include the caveats with that. That's what I see as my role in this thread. Otherwise, it looks like general purpose regex engines are crippled for no apparent reason.
- dan-dev 4y agoOh well, first of, thank you for writing ripgrep! Great work, love that tool and use it daily :) All good, thank you for sharing your POV! I find both theory and the practical engineering feats quite interesting, and actually understood the original question as a theoretical question, not a question about general-purpose regex engine. Quite possibly that this is even how the asking person was intending it, in hindsight. Last thing on topic (talking now about having complement and unicode in a usable regex engine): My personal, gut-feeling-based hunch is that by adding all these features not strictly related to regular languages, we kind of evolved our regex enginges and expectations towards them into a corner of the design space that works well, but might not work well when one adds complement. It feels like a constructivist world without negation and stuff on top that works, and reverse-engineering what works with negation and what not will likely be hard work. As a theory-ish kind of person (without much insight into regexes or formal grammars), I wish there would be some kind of insight into what are the atoms of a regex and which atoms can be combined witch which other atoms - I would bet that fancyFeatures and unicode+complement would have a high algorithmic complexity already in theory. Any final tipps on the repo I linked to not disappoint users expectations? :)
- burntsushi 4y agoI think you got it right by putting "university project" in the description. :-) I wrote my fair share of projects like that too. The incentive structure of academia that leads to software like that, though, is also a big reason why I left. The theory is super important (that's why I work on a regex engine that only provides support for regular languages), but I also think the engineering aspect is just as important. And academia, or at least the corner I was in, just did not care about the engineering side of things at all. This in turn leads to publishing results that are hard to reproduce, because if you don't share code that is at least somewhat robust, it's going to be very costly for someone else to build on your work. Anyway, sorry about the side rant haha. But all good.
- dan-dev 4y agoAll good! And sorry for not making the effort to understand your position first, I want to work on that. > The incentive structure of academia that leads to software like that, though, is also a big reason why I left. The theory is super important (that's why I work on a regex engine that only provides support for regular languages), but I also think the engineering aspect is just as important. And academia, or at least the corner I was in, just did not care about the engineering side of things at all. This in turn leads to publishing results that are hard to reproduce, because if you don't share code that is at least somewhat robust, it's going to be very costly for someone else to build on your work. Totally get it. This also baffles me - papers in computer science that are hard to reproduce? Yes I know, publish or perish and all that BS. But I don't want to accept this: I mean I get it for psychology obviously, and then also material sciences and so on, but for compsci?... Next to proofs, programs are probably one of the most reproducible things that nature has given given us. Good for me and many others that you cho(o)se to spend your time on building tools :)
- avgcorrection 4y agoBest not to get uppity on the subject of regexp when confronted with the great burntsushi… he can get irate.
- avgcorrection 4y ago> Full disclosure, I'm the author of ripgrep, so I have some particularly relevant experience in the specific domain of making regexes work well for the masses. The unwashed ones?