4 ms·
* RegEx opens your application up for DoS attacks * RegEx is not very readable * RegEx can be (very) slow * It's not trivial to write RegEx code that achieve
by sharpercoder 9y ago
* RegEx opens your application up for DoS attacks
* RegEx is not very readable
* RegEx can be (very) slow
* It's not trivial to write RegEx code that achieves your goal in a high-quality way. Often quircks and edge-cases are missed.
I'm not saying that you should never use them, but oftentimes a (much) better alternative of achieving your goals is present.
See https://blog.codinghorror.com/regular-expressions-now-you-have-two-problems/ https://blog.codinghorror.com/regular-expressions-now-you-ha...
- geofft 9y agoAs the article you link argues, use of regular expressions when they are inappropriate is bad. This particular case - finding and replacing certain characters with other characters - is pretty well-suited to the problem, and is probably more readable than a bunch of open code to do the same thing. (I'm not sure what you mean by DoS attacks - are you referring to the exponential case of backtracking? If so, don't use a regex engine with that problem, and don't use lookbehind/lookahead assertions, which aren't needed to solve this problem.)
- sharpercoder 9y ago> This particular case - finding and replacing certain characters with other characters - is pretty well-suited to the problem, and is probably more readable than a bunch of open code to do the same thing. No, it is not a good solution to the problem; you ignore my earlier comments. English or latin text is not comprised of the sole ASCII characterset; it contains characters outside this set (quoting other languages, names, imported words for example).
- laumars 9y agoGood thing most regex engines handle unicode ;) Honestly, I do get your point about inappropriate use of regex, but this kind of simple text manipulation is well suited for regex. The biggest argument against using regex for this kind of problem is performance verses writing the same code programmatically in the host language (assuming you're using a fast AOT compiled language). However even that is a non-issue given the small quantities of text you're decoding. Also I'd bet the regex in this instance would actually work out more readable because the transformations are basic so you're localising the text manipulation to simple rules rather than multiple lines of byte array reading and thus also potentially having to manually build in your own rudimentary unicode support too.