5 ms·
DeepRegex: Neural Generation of Regular Expressions from Natural Language
- RankingMember 10y agoAs soon as I saw the title I started thinking of how great it would be to have a solid NaturalLanguage -> Regex translator, something I've always wanted. Regex is so powerful, but sometimes it takes so long to get it to do what you want it to.
- jasonjmcghee 10y agoI couldn't see myself ever using a nl to regex translator. English especially is incredibly ambiguous and regex are simple and, as you noted, very powerful. Implementing a regex interpreter personally enabled to create even the trickiest regex. I'd highly recommend it! Knowing how to write complex regex also makes sed/grep your best friend. :)
- deleted 10y ago[deleted]
- flanbiscuit 10y ago> how great it would be to have a solid NaturalLanguage -> Regex translator This idea sounds good, but as soon as you start getting slightly more complicated you'll be writing paragraphs: Try writing this in a natural language format: <a\s+(?:[^>]*?\s+)?href="([^"]*)" That's a regex to get the value of an href from anchor links. "match "<a " then do not match a ">" if it exists, followed by a space, if it exists, then match a "href=", then begin a capturing group, then match anything but a '"' 0 or more times, then close a capturing group, then match a '"' I'm sure that's not even correct but you can see what I mean. I can see this idea being a good tool for learning though, especially for smaller regex
- hamandcheese 10y agoIt's not wise to attempt to parse html with a regex :) http://stackoverflow.com/questions/1732348/regex-match-open-tags-except-xhtml-self-contained-tags/1732454#1732454 http://stackoverflow.com/questions/1732348/regex-match-open-...
- modeless 10y ago> Try writing this in a natural language format: > [...] That's a regex to get the value of an href from anchor links. A true natural language to regex system would take your short natural language description as input, not paragraphs describing the task in more detail. Of course this would require a lot of domain knowledge about HTML, but that knowledge is readily available out there on the internet. I think it's no longer crazy to imagine a system which could read the internet, learn about HTML, and apply that knowledge to answer your natural language regex query. This is clearly far beyond where we are today, but I think a few orders of magnitude larger neural nets would be able to handle this task, and the hardware guys are hard at work getting us there. The pace of improvement will be much faster than Moore's law over the next couple of years as the first optimized neural net hardware becomes available.
- mynewtb 10y agoThere is one that I always meant to try but I cannot find it again. One chained methods together or something.
- nicklo 10y agoHey HN! One of the authors here. This was a fun project to work on. Happy to answer any questions :) Code + data here: https://github.com/nicholaslocascio/deep-regex https://github.com/nicholaslocascio/deep-regex
- vessenes 10y agoHey, awesome work. Do you have any generated regexs? It would be nice to see some examples, especially if surprising in some way.
- nicklo 10y agoSure! Some positive ones: 1) Spot-on prediction: PROMPT: lines with 3 or more characters or lower-case letters PRED: ((.)|([a-z])){3,} GOLD: ((.)|([a-z])){3,} 2) Learned to generalize and produced a simpler regex: PROMPT: lines with a character and the string 'dog' PRED: .*(.)&(dog).* GOLD: .*((.)+)&(dog).* 3) Also learned to generalize and produced simpler regex without duplicate logic: PROMPT: lines not containing a letter PRED: .*~(([A-z])+).* GOLD: (.*)(.*~([A-z]).*) 4) Handling multiple references correctly: PROMPT: lines using 'su' after 'sun' or 'soon'. PRED: .*(sun|soon).*su.* GOLD: .*(sun|soon).*su.* Though I find the mistakes interesting as well! 1) Issues counting properly: PROMPT: lines containing a 5 letter word beginning with 't' PRED: .*\bt[A-z]{5}\b.* GOLD: .*\bt[A-z]{4}\b.* 2) Misallocation of parenthesis (to be fair, the prompt is slightly ambiguous): PROMPT: lines with 'dog' follwed by 'truck' and a lower-case PRED: (dog).*((truck)&([a-z])).* GOLD: (dog.*truck.*)&(.*[a-z].*)
- throwwit 10y agoMistake 1 - Looks like the classic off-by-one! Definitely the boundary point for a Chomsky Grammar type. Modifications to the code for processing problems like the Sorites paradox would be interesting.
- utopkara 10y agoCongrats! The regex translation dataset generation is a great idea! 10000 lines could be the MNIST for regex :-)
- deleted 10y ago[deleted]
- duaneb 10y agoThe applications of this for scraping alone are amazing
- staticautomatic 10y agoImagine translating natural language into XPath queries. Now that would be something.
- deleted 10y ago[deleted]