7 ms·
For some reason regular expressions have the lowest expiry date in my mind's cache. I had to relearn them at least 10 times. Watching that famous Udemy programm
by macando 7y ago
For some reason regular expressions have the lowest expiry date in my mind's cache. I had to relearn them at least 10 times. Watching that famous Udemy programming course where regex is explained in great detail with examples and state machine diagrams didn't help.
- 52-6F-62 7y agoSame here save for some of the basics. Regardless, I love this tool: https://regex101.com/ https://regex101.com/ I usually hammer out a prototype there before testing in any method if I have to write anything non-trivial or that may have edge cases I want to test out.
- kbenson 7y agoRegular expressions are a powerful tool, but unless you use them often after learning them (and are in a language where it makes sense to do so), it can be hard to make it stick. I've had amazing success in using a regex in situations where you wouldn't think it would work as well as other solutions. For example, I've gotten more than an order of magnitude speedup in parsing a well known simple (but fairly large) XML data set using regular expressions instead of the fastest XML parsing libraries I could find at the time. Sometimes the less efficient tool that does only what you need is much faster than the highly optimized tool that handles all the special cases that don't matter for the particular job.
- macando 7y agoYou use whatever works for your case. When I hear parsing with regex I always think of this legendary Stack Overflow answer :) https://stackoverflow.com/a/1732454 https://stackoverflow.com/a/1732454
- kbenson 7y agoYeah, that's a classic. But it's also mostly about parsing arbitrary HTML, and that's where "well known" comes in. The data in question looked somewhat like: <doc> <bool name"foo">1</bool> <int name="bar">12345</int> <str name="baz">test string</str> <float name="quux">1.234</float> </doc><doc> <bool name"foo">0</bool> ... </doc> but with a lot more fields per-doc. To pull out each <doc> as a string (using a regex) in a while loop and to then parse the doc into a hash of key-value pairs (another regex) that are stored into an array is less than twenty lines of pretty standard Perl, include exception handling and error reporting: my @items; my $count = 0; while ($item_xml =~ m{<doc>(.*?)</doc>}gsmi) { try { # process <doc> my $item = $1; my $i = {}; $i->{$1} = $2 while $item =~ m{<(?:arr|date|str|bool|int|float) name="([^"]+)">([^<]+)</[^>]+>}gsmi; push(@items, $i); $count++; } catch { warn "Error :: $_"; } } print "$count items found\n"; And if you think the regex to pull out the field values is hard to read, it could always be written in the extended format like so (this is overly verbose, but you should get the idea) my $field_parse_re = q{ # Parse opening tag < # Start of opening tag # Any of the allowed tag types (?: arr | date | str | bool | int | float ) \s # space between tag type and name attribute name="([^"]+)" # Save tag name as $1, or first returned item > # End of opening tag # Parse tag contents ([^<]+) # One or more characters that are not <, in $2, or second item returned # Parse enough of the closing tag to make sure we got all the contens </ }smix; # Now use it $t->{$1} = $2 while $ticket =~ m{$field_parse_re}gsmi; In any case, that's a trivial amount of work to beat the fasted XML parsing I could find (and I surveyed a few libraries) by like 14x, IIRC.
- macando 7y agoBeing an engineer means finding/building the best tool for the task at hand. Sometimes it feels great to write a short and effective snippet of code instead of examining how some bizarre API works.