3 ms·
My take is that most regexes are easier to read than the alternative code. You do have to step through them and think about what they match, but the same is tru
by thethirdone 4y ago
My take is that most regexes are easier to read than the alternative code. You do have to step through them and think about what they match, but the same is true for any other code.
> the very existence of RegEx101.com ought to bring shame on our industry.
> ...
> You have a desire to build something hard to debug.
Regex101 is there to make it easy to debug regexes. It would be MUCH harder to debug a complex parsing function than a regex with Regex101.
> You don't trust compilers.
I don't even understand the idea here. You need a compiler for the regex. Why would someone that doesn't trust compilers use regexes?
- Jtsummers 4y ago> It would be MUCH harder to debug a complex parsing function than a regex with Regex101. I encountered, this week, a 42-line function that determines if a number is a valid string representation of a double. The function is wrong and has been wrong for a decade or more. The code is obtuse using unclear state variables (multiple boolean flags) to accept or reject the string at various points. It could have been a regex, and we could have seen at (almost) a glance what was intended. There are several other functions in the same file doing similar things so this ended up being something like 300+ lines of both wrong and difficult to understand code that could have been maybe 30 total lines of relatively easy to read and debug regexes.
- throwanem 4y agoYeah, I'd definitely replace that whole function - with a try/catch-wrapped (or NaN-checked, etc.) attempt to parse the string as a double. Not a regex, which then still would have to be tested to make sure it's exactly as strict as whatever parser it's going to be passed to, anyway.
- Jtsummers 4y agoFair, but I'm not even sure where this value ends up yet so I don't want to parse it (vice validate it), just getting it to reject "1-2" as a double would be a start (currently accepted, note that it's not accepting expressions just values). Stage one of fixing this thing is just tackling obviously wrong code like this, their creative shared pointer implementation that's also wrong (the memory will be freed while other pointers still have access to it, and the reference counter is not in the shared pointer but the object it contains so checking the reference counter at that point will result in use-after-free errors), and their fantastic use of iterators that will crash the program if that code is ever called (they didn't understand iterator invalidation). And yesterday I found that they read a file (in its entirety!) one character at a time, tacking the results onto the end of a string. There is no parsing logic at that point, it's literally just building a giant string. And all of that is in what should be straightforward logic any trained novice can handle, the real critical code which solved a somewhat novel problem (not novel novel, but not an everyday problem) gets even more creative.
- throwanem 4y agoBy the sound of it, this may well indeed be the rare case in which jwz's dictum fails to hold, and using a regex will yield one problem fewer - or, failing that, maybe kick a hole in the side and go looking for a river or two. In any case, good luck...
- fiddlerwoaroof 4y agoI agree, I find regular expressions to be the clearest way to describe families of strings that share a specified structure and no longer find them hard to read. This sort of argument has always struck me a bit like the way people will say “Greek is hard to read because you have to learn the alphabet”: learning the alphabet for Greek isn’t the hard part of learning Greek and, similarly, learning the meaning of the characters used in regexes isn’t the hard part of using regexes. The hard part is learning to think in the language that uses those characters.
- marginalia_nu 4y agoRegular expressions are virtually always easier to read because they're so much terser. I unroll a fair bit regular expressions because they are typically an order of magnitude slower, and it's almost never an upgrade in readability. Compare for example something trivial like [a-f0-9]{32} with its unrolled form public boolean hashTest(String path) { int runLength = 0; int minLength = 32; if (path.length() <= minLength + 2) return false; for (int i = 0; i < path.length(); i++) { int c = path.charAt(i); if ((c >= '0' && c <= '9') || (c >= 'a' && c <= 'f')) { runLength++; } else if (runLength >= minLength) { return true; } else { runLength = 0; } } return runLength >= minLength; }
- lmm 4y ago> It would be MUCH harder to debug a complex parsing function than a regex with Regex101. Use a parser combinator library and your parser is much easier to debug, because everything's compositional and you can unit-test it.
- bazoom42 4y agoYeah, when complaining about multi-line regexes, show us the alternative imperative code. Is it really easier to understand and maintain?