4 ms·
After 20 years of software development I‘ve come to adopt a best practise: Whenever I start writing a regular expression, I stop and write a „manual“ domain sp
by cel1ne 9y ago
After 20 years of software development I‘ve come to adopt a best practise:
Whenever I start writing a regular expression, I stop and write a „manual“ domain specific parse function instead.
Saved me a LOT of debugging time.
Since I can now use kotlin pretty much anywhere (jvm, browser, shellscripts) this is easy because of the superb stdlib („startsWith“, „lastIndexOf“, „substringBeforeLast(...)“)
The time saved I invest in Unittests for the parser.
- darethas 9y agoYep. I primarily use Go, so you are often forced down this way because it uses a simpler regex engine. I used to complain but in hindsight I realized it was a blessing in disguise. Particularly in Go's case, it has excellent character set library support, especially unicode, so those really tricky corner cases with unicode characters are non-existent now as well. I will be happy if I never see a regex with a unicode range again.
- hinkley 9y agoI can't shake the feeling that Regexp could be written just as efficiently as a fluent interface with a more human friendly syntax. I've been telling Jr devs bucking for promotion for years to explain what they're doing in plain english, then write code that looks like that. Basically telling them to skip right over the "gee look what a clever fuck I am" stage and write good code instead of creating riddles. The Regexp problem just screams this at me. What am I doing? I'm looking for a line that starts with a capital T, then has some quantity of alphanumeric characters greater than n (if n is not 0, 1, or infinity, this requires extra work in Regex), followed by an equals sign with or without whitespace characters around it. Give me an API that does exactly that, instead of Regex. Something gets lost in translation every time. I think the fact that the origin of Regex is the command line interface is pretty telling. We didn't and we don't have a convenient way to type in imperative code on a command line. So an arcane syntax was created so you could do the whole thing in a quarter line of text. Speaking as someone who has had a Unix shell for 25 years, and routinely works on mini tools for their fellow developers, I don't think we actually type stuff into a shell that often anymore. The difference between documenting a one-liner in a README and just building a shell script that does the same thing is not that big. There's a difference in development effort but building a script can allow you access to a debugger. Personally, I'd be willing to pay that tax any day.
- cel1ne 9y agoI agree. That’s why I pointed out that a necessary addition to get more productivity out of writing custom parsers is a) easy usability in shellscripts B) easy tooling for unittests I started writing shellscript in kotlin, so i get the unittests for free
- lalaithion 9y agoParsing combinators are the API you want. regexReplaced n = do char 'T' spaces x <- concat (replicateM n alphaNum) y <- concat (many alphaNum) spaces char '=' return (x ++ y) The above is a Haskell function that does the parsing required above, returning the alphanumeric characters if the parse succeeds and returning an error if it does not. You may not speak Haskell, but this is probably still more readable than (n) => {new RegExp(`T\s(\w{${n}}\w)\s*=`)}, which is the Javascript function that does a similar thing.
- xfer 9y agoBut are they as efficient as regexp? I personally prefer regexp combinators.
- lalaithion 9y agoOptimizing them is a bit more work, but they can outperform hand-rolled C code: http://www.serpentine.com/blog/2014/05/31/attoparsec/ http://www.serpentine.com/blog/2014/05/31/attoparsec/
- ZenoArrow 9y ago> "I can't shake the feeling that Regexp could be written just as efficiently as a fluent interface with a more human friendly syntax." You may be interested in the Parse dialect of Red: http://www.red-lang.org/2013/11/041-introducing-parse.html http://www.red-lang.org/2013/11/041-introducing-parse.html Also worth noting that Red can be embedded in any program that supports a C function interface, through using LibRed. http://www.red-lang.org/2017/03/062-libred-and-macros.html?m=1 http://www.red-lang.org/2017/03/062-libred-and-macros.html?m...