5 ms·
This regex that took about 10 minutes to generate was amazingly effective and helped me earn my highest ever daily rate https://gist.github.com/dr-kd/d43c884fb
by singingfish 3y ago
This regex that took about 10 minutes to generate was amazingly effective and helped me earn my highest ever daily rate https://gist.github.com/dr-kd/d43c884fbac0089d8523 https://gist.github.com/dr-kd/d43c884fbac0089d8523
- caesil 3y ago>that took about 10 minutes to generate How?
- Matheus28 3y agoI'm assuming: 1. Generate a list of suburbs names separated by | 2. Simplify the regex: 2.1. Convert to a NFA 2.2. Turn it into a DFA 2.3. Merge redundant states 2.4. Turn it back into a regex
- KomoD 3y agonfa? dfa?
- zdimension 3y agoNondeterministic finite automaton and deterministic finite automaton. The theory behind state machines.
- KomoD 3y agoThank you
- singingfish 3y agohttps://gist.github.com/dr-kd/36df691fa395567cb73d1368e4345c6a https://gist.github.com/dr-kd/36df691fa395567cb73d1368e4345c...
- veeti 3y agoDo they ever coin new suburb names? What happens then?
- singingfish 3y agoYes they do. This was good for the early/mid 2000s - new suburbs are going to be a problem as australia post no longer give the list away for free.
- wodenokoto 3y agoHow is that better than just checking against a list of suburb names?
- _0ffh 3y agoThe regex is a compressed representation which saves memory, and it's also likely to be quite a bit faster which saves cycles. I consider it a clever bit of optimisation.
- singingfish 3y agoIt's also handling what state the suburb is in, and can deal with various kinds of bad data (e.g. lack of uniformity of spacing and capitalisation).
- phanimahesh 3y agoHow did you generate this?
- singingfish 3y agoperl module Regexp::Assemble