4 ms·
Hey HN! One of the authors here. This was a fun project to work on. Happy to answer any questions :) Code + data here: https://github.com/nicholaslocascio/dee
by nicklo 10y ago
Hey HN! One of the authors here. This was a fun project to work on. Happy to answer any questions :)
Code + data here:
https://github.com/nicholaslocascio/deep-regex https://github.com/nicholaslocascio/deep-regex
- vessenes 10y agoHey, awesome work. Do you have any generated regexs? It would be nice to see some examples, especially if surprising in some way.
- nicklo 10y agoSure! Some positive ones: 1) Spot-on prediction: PROMPT: lines with 3 or more characters or lower-case letters PRED: ((.)|([a-z])){3,} GOLD: ((.)|([a-z])){3,} 2) Learned to generalize and produced a simpler regex: PROMPT: lines with a character and the string 'dog' PRED: .*(.)&(dog).* GOLD: .*((.)+)&(dog).* 3) Also learned to generalize and produced simpler regex without duplicate logic: PROMPT: lines not containing a letter PRED: .*~(([A-z])+).* GOLD: (.*)(.*~([A-z]).*) 4) Handling multiple references correctly: PROMPT: lines using 'su' after 'sun' or 'soon'. PRED: .*(sun|soon).*su.* GOLD: .*(sun|soon).*su.* Though I find the mistakes interesting as well! 1) Issues counting properly: PROMPT: lines containing a 5 letter word beginning with 't' PRED: .*\bt[A-z]{5}\b.* GOLD: .*\bt[A-z]{4}\b.* 2) Misallocation of parenthesis (to be fair, the prompt is slightly ambiguous): PROMPT: lines with 'dog' follwed by 'truck' and a lower-case PRED: (dog).*((truck)&([a-z])).* GOLD: (dog.*truck.*)&(.*[a-z].*)
- throwwit 10y agoMistake 1 - Looks like the classic off-by-one! Definitely the boundary point for a Chomsky Grammar type. Modifications to the code for processing problems like the Sorites paradox would be interesting.
- utopkara 10y agoCongrats! The regex translation dataset generation is a great idea! 10000 lines could be the MNIST for regex :-)
- nicklo 10y agoThanks! Starting this work, we realized that large regex datasets (large enough to apply deep-learning to) were difficult to come by. So we came up with a methodology that allowed us to make a pretty decent-sized dataset for cheap. We are glad to share it :)
- joe_the_user 10y agoIt's an interesting project but to be honest, Regular expressions are kind of terrible as a piece of software if one is looking to embed them in larger software, especially if the regex gets at all large. Have you can considered generating something like a formal grammar, a recursive-descent parser or a software library?
- nickpsecurity 10y agoDespite this work being neat, that's still the best way to do it that I'm aware of. First reason is that English is imprecise enough that CompSci invented formal specifications to solve the problems that created. This is an immediate step back from precise requirements. Second, there's a ton tooling to automatically generate parsers from precise grammars. The languages, even BNF, are pretty easy to teach people. There's also text languages & spec methods that can make comprehensible stuff that regex's or BNF might muddy up. All of these take almost no CPU effort to deterministically produce a result from. So, the old ways are still better for this domain if it's a production system whose cost or results matter. These methods might be useful for search/query by casual users, though. Or people that come from a foreign language likely to express queries in a weird way.
- jgalt212 10y agoAre you familiar with the work of the Machine Learning Lab at the University of Trieste? And, if so, could you quickly comment on how the two approaches differ? http://machinelearning.inginf.units.it/ http://machinelearning.inginf.units.it/ http://regex.inginf.units.it/ http://regex.inginf.units.it/
- nicklo 10y agoYeah their work is very cool! The two tasks are a bit different however. Their system is for the task of generating regular expressions given a set of positive and negative examples of what the regex should match. It uses genetic algorithms and other techniques to optimize and search for a regex that fits all the given examples. In our case, we have no examples to test against, only a natural language (English) description of what the user wants the regex to do. This is an inference problem more than a search problem as we've got one shot to give our best guess without any tests to check against and modify our answer.
- jgalt212 10y agocool, thanks for the explanation.