13 ms·
Regex Golf
- Trufa 13y agoI like the concept but the word choice doesn't seem too "regular", it is more about catching all the particulars rather than finding a pattern as far as I can tell.
- ZirconCode 13y agoThere are good patterns, the first regex for me was /foo/, the second /ick$/, for example. They're there.
- davidbrent 13y agoI have the most trouble with regex, and something about seeing my matches as I type made this incredibly useful for me!
- tokenizerrr 13y agohttp://regexpal.com/ http://regexpal.com/ or for a paid desktop program http://www.regexbuddy.com/ http://www.regexbuddy.com/, which is most excellent
- rschmitty 13y agoRegexbuddy is great. Well worth the coin I paid way back when
- Inversechi 13y agoThe best online tool I've seen is http://www.debuggex.com/ http://www.debuggex.com/
- mcescalante 13y agoI've used the previously mentioned regexpal quite often. http://www.rubular.com http://www.rubular.com is another one for Ruby only.
- Spittie 13y agoI'll add http://www.regexper.com/ http://www.regexper.com/ to the various sites everyone posted, I find that this one output the nicer graph.
- ToastyMallows 13y agoWhile people are recommending online regex tools.. http://regex101.com/ http://regex101.com/
- lsv1 13y agoFun learning tool for regex.
- chrismorgan 13y agoPractise, maybe. Learning, probably not so much. Certainly fun, though, if you can cope with them!
- zedadex 13y agoLearning, kinda. I've often been motivated to learn something after being given a problem that can only be solved (or be solved much more easily) using it. Every other Excel trick I know is a result of that.
- daGrevis 13y agoI like the idea, but words there seems to be pretty random. I can't figure out the pattern, not even talking about writing regex... :(
- maxerickson 13y agoI got a 3 character pattern for the first puzzle and a 4 character pattern for the second puzzle (and reading a comment here I see there is a 2 char pattern for the second puzzle). Only used 1 special char between them (anchoring something). So the first two sets have repeated 3 character sequences, and in one of them they follow an additional pattern. For the third sequence, 8 character solution, lots of specials.
- rplnt 13y agoRead a) name of the puzzle b) the little help below the score. It gives out the pattern you should (not) look for. edit: Not in Glob. What is Glob about?
- schoen 13y agoIt's about implementing * as a wildcard character: https://en.wikipedia.org/wiki/Globbing https://en.wikipedia.org/wiki/Globbing I haven't figured out how to solve this with the parts of ERE and PCRE that I know. (I definitely don't know the entirety of PCRE.) It's straightforward for me to write a substitution using regular expressions to create a pattern-matcher for a given glob (just anchor the ends and replace literal ? with . and literal * with .*) but here we have to do it inside a single regular expression. I don't think there's a BRE solution if the number of stars is unbounded because I don't think this is a regular language.
- bencoder 13y agowhat's the pattern on "Abba"? I thought it was just to exclude doubled letters but I have doubles on two words on the left hand side as well (noisefully and effusive, in case the word lists are the same)
- shdon 13y agoExcluding doubled consonants that have the same vowel on both sides of the pair: ^((?!([aeiou])([^aeiou])\3\2).)*$ The following also works for the testcases and is shorter: ^((?!(.)(.)\3\2).)*$
- surreal 13y agoCan get 2 more points with: ef|^((?!(.)\2).)*$ But I reckon yours wins for having a pair of breasts in the middle
- moron4hire 13y agowell, if you are strictly interested in matching with the fewest characters, regardless of the apparent pattern: ^(.)(.).*\2\1$
- danielweber 13y agoYour anchors at the end make it match nothing.
- moron4hire 13y agosorry, misreplied. This was for palindromes.
- abus 13y ago5. Abba (^unv|.u|z|ph|mi|st|vi|tan) haha
- deleted 13y ago[deleted]
- sciolistse 13y agothat's twice as long as it has to be.. not that it matters.
- maxerickson 13y agoWhy publish them? You only need 1 range.
- Sir_Cmpwn 13y agoI got ranges at 202 with [a-f]{4} 198 on plain strings with .\*foo[lt]?.* Then I stopped trying.
- gorhill 13y agoPlain string: just 'foo' works.
- easy_rider 13y agofml
- shdon 13y agoIndent your regexes by two spaces and precede them with a blank line, that will preserve asterisks ;)
- deleted 13y ago[deleted]
- deleted 13y ago[deleted]
- shdon 13y agoAnchors (208): k$ Ranges (202): ^[a-f]+$ Backrefs (199): (\w{3}).*\1
- chaz 13y agoShouldn't the objective be to get the lowest score if it's called "golf?"
- danceonfire 13y agoThe goal in code golf is lower character count, not lower score - which is also the case here.
- talmand 13y agoTherefore, the score it provides should be a running tally of how many characters you've used to make it match the scoring system of golf; which is the lowest number of strokes wins. The scoring system for this is incremental which is the opposite of golf. A proper scoring system with this would provide a character limit (par) for each section and the goal would be to write a shorter regex formula to complete the task. Final score would be how many characters under or over the total character limits (course par) you scored. Seems this is more like Regex Darts or something like that. But it's fun nonetheless.
- danceonfire 13y agoI don't get what this is all about. The title was (most probably) derived from Code Golf, which is a competition in coding something with as few characers as possible. Code Golf was derived from Golf, where you want to use as few turns as possible. The score going up and not down, which is done because you get more points the more objectives you fulfill, does not change the objective of this game or what it is based on.
- rschmitty 13y ago> The score going up and not down, which is done because you get more points the more objectives you fulfill, does not change the objective of this game or what it is based on. Golf scoring penalizes you with more "points" by how many strokes you take. If you are playing a Par 4 and it takes you 6 swings to get in the cup you just got _penalized_ +2 If your partner gets in the cup in 2 swings he is awarded -2 Therefor "Regex Golf" has its scoring reversed
- ZirconCode 13y agoFor "6. A man, a plan", I thought it was impossible to match palindromes with regex, am I wrong?
- quarterto 13y agoIt gives you hints below the score, for 6 it's: You're allowed to cheat a little, since this one is technically impossible. No idea how you cheat... EDIT, SPOILERS: I get 170 with ^(.?)(.)(.).?\3\2\1$
- mryingster 13y ago176 with ^(.)(.).*\2\1$
- deleted 13y ago[deleted]
- hyp0 13y ago^(.)[^p].*\1$ # 177, "cheat a little"
- deleted 13y ago[deleted]
- Yen 13y agoYou're pretty much correct. True regular expressions don't have the expressive power to decide whether arbitrary-length strings are or aren't palindromes. That said, 'regular expressions', as used in most programming languages, have extensions that extend the expressive power. One such extension is the matched group & backreference, used in other commentor's answers. From a theoretic stance, these aren't really 'regular expressions', but that's what we call them in practice.
- joelanman 13y agoam I being a bit slow? Why doesn't [^g-z] work on 'ranges'?
- martinml 13y agoBecause for example "beam" matches [^g-z]. That is, it has a letter somewhere that is not between g and z (namely e and a). I came up with ^[a-f]+$ but I'm guessing it could be shorter :) Edit: ah, every word in left column has 4 a-f letters. So [a-f]{4} is a shorter match.
- deleted 13y ago[deleted]
- joelanman 13y agoah you're right :) I was being slow, what I thought I wrote was 'words consisting only of letters that arent g-z'
- bluedino 13y agoAnother dumb question - why does 'beam' not match [bdf][ae] ?
- maxerickson 13y agoIt's complaining because it does match.
- bluedino 13y agoOh. I told you it was a dumb question.
- deleted 13y ago[deleted]
- deleted 13y ago[deleted]
- tareqak 13y agoI got Four for 196 with (.).*\1.\1.*\1 and Order for 156 with ^a*b*c*d*e*f*g*h*i*j*k*l*m*n*o*p*q*r*s*t*u*v*w*x*y*z*$
- chrismorgan 13y agoFour: 199 (.)(.\1){3} Order: 168 ^a*b*c*d*e*f*g*h*i*l*m*n*o*p*r*s*t*w*y*z*$ (You just didn't remove unused letters.)
- spystath 13y agoOr just (.).\1.\1.\1 for Four (198)
- deleted 13y ago[deleted]
- MereInterest 13y agoSo long as we are going down this route, you can shave off another character with (.)(.\1){3} (199).
- maxerickson 13y agoCan chop out one check to save a character: (.)...\1.\1
- gorhill 13y agoOrder for 198: ^[^o].....?$ Probably not what was wanted, but it works (or maybe it was to trick people onto a false path)
- The_Double 13y agoDoes anybody know how to do math with regex? (triples) And conditionals don't seem to work?
- danielweber 13y agoMy guess is that 147 must appear the same number of times as 258, but I'm not sure if that's even expressible in regexp.
- recursive 13y agoNo. 111 is a multiple of 3.
- johnlbevan2 13y agoSadly that wouldn't work for: 140091876 147 = 4 times 258 = 1 time
- johnlbevan2 13y agoRefined: [0369] can appear any number of times [147] and [258] must appear an equal number of times, or for any remaining: [147] must appear a multiple of 3 times [258] must appear a multiple of 3 times
- dudus 13y agoFrom: http://quaxio.com/triple/ http://quaxio.com/triple/ ^([0369]|[258][0369]*[147]|[147]([0369]|[147][0369]*[258])*[258]|[258][0369]*[258]([0369]|[147][0369]*[258])*[258]|[147]([0369]|[147][0369]*[258])*[147][0369]*[147]|[258][0369]*[258]([0369]|[147][0369]*[258])*[147][0369]*[147])*$
- moron4hire 13y ago569pts: ([^31]0|31|[017]2|[03]03|[^1]4|(900|01|7)5|6|[48]7|[57]8|09)$
- chch 13y ago
- easy_rider 13y ago(.+|)foo(.+|) Lol so awesome this. I oblige. Much love!
- ryanthejuggler 13y agoPowers: ^((((((((((x)\10?)\9?)\8?)\7?)\6?)\5?)\4?)\3?)\2?)\1?$ I feel like there's gotta be a sneakier way of doing this.
- rplnt 13y agoMine is a bit shorter, though a bit more "meh" as well. ^x{32}$|^(x{2}){1,8}$|^(x{64})+$|^x$
- galen_tyrol 13y agoimproved, gives 80 ^(x|(xx){1,9}|x{32}|(x{64})+)$
- rplnt 13y agoI tried to get rid of those redundant ^ and $ but it somehow didn't work. I probably forgot to put it all inside one group.
- danielweber 13y ago
- Pxtl 13y agoI hit enter and nothing happens.
- danielweber 13y agoUse the yellow box, not the name box that gets your focus when you land on the page. This confused me for a few minutes.
- deleted 13y ago[deleted]
- deleted 13y ago[deleted]
- hadem 13y agoReminds me of Vim Golf. http://vimgolf.com/ http://vimgolf.com/
- deleted 13y ago[deleted]
- danielweber 13y agoMany times I would match the exact opposite opposite of what I wanted. Is there a general rule for inverting regexps? ^ and ?! don't seem general purpose.
- josephlord 13y agoYou need to pin the match to the start and end of the string with ^ and $ respectively otherwise the negation just matches an empty string or other irrelevant string.
- danielweber 13y agoJust to follow my thought pattern. I start with (.)(.)\2\1 Now to invert I change it to (?!(.)(.)\3\2) As you say, the negation matches an empty string. So I put on anchors: ^(?!(.)(.)\3\2)$ But now it matches nothing. I'm not even sure what that regexp says. The entire string is a negation? Would anything match that regexp? I see elsewhere on this page that the answer involves putting in an extra dummy character, putting that new negation-and-dummy-character in parens, and then requiring that, between the anchors, there be 0-or-more of negation-and-dummy-character. ^((?!(.)(.)\3\2).)∗$ Two questions: 1. Why are my backrefs still \3 and \2? I added another pair of parens. (I thought ?! might not count, but it counted in my second example above. 2. Why does abba no longer match? It has no matches to the negation-and-dummy character construct, which ∗ should match, right? NB: I used ∗ as my asterisk to avoid bb-code.
- josephlord 13y ago^(?!.*(.)(.)\2\1) You may also need to fill in the places where it could be anything. The above worked on abba for me. BTW a double space indent then formats as code on HN I think.
- Procrastes 13y agoYour approach scores higher, but it only matches the (imaginary) space before the "good" words. I went with: ^(?!.(.)(.)\2\1).$ with the thought that if I really wanted those matches I would want the whole strings. Fun game!
- deleted 13y ago[deleted]
- josephlord 13y ago^(?!(..+)(\1)+$) Why does that work on primes? I got it by mistake when fiddling with the parenthesis locations but I was expecting to have to deal with xx separately.
- surreal 13y agoNice find. It works because it rejects "2 or more x's" repeated "2 or more times". So xx doesn't get rejected, but any multiple of that (xxxx, xxxxxx, ...) will be. The same way xxx doesn't get rejected, but any multiple of that (xxxxxx, xxxxxxxxx, ...) will be. You've solved it using the actual definition of prime numbers, no trickery needed. Well played. FYI, you don't need brackets around the \1, so can score 286.
- josephlord 13y agoThanks, I was going for the definition of primes but was just going a bit loopy about the negation. I didn't want to trim surplus characters until I understood it but yes I can see those brackets are unnecessary.
- Aissen 13y agoMore interesting than the definition of primes, it's almost the definition of multiplication that is embedded in this regex. We have two numbers(of occurrences) being multiplied: - the first one is represented by the group (..+) it represents the number of occurrences n between 2 and +∞ - the second one is represented by (\1)+. We will repeat the first number m times, between 1 and +∞ times. So the result of the multiplication is n*(m+1), which cannot be a prime. We just have to take the opposite with negative lookahead. It's very beautiful indeed. See http://regex101.com/r/qN2fQ8 http://regex101.com/r/qN2fQ8 or http://www.regexper.com/#^%28%3F!%28..%2B%29\1%2B%24%29 http://www.regexper.com/#^%28%3F!%28..%2B%29\1%2B%24%29 to follow the above explanations.
- vijucat 13y agoThanks for the web site links! Both are pretty interesting and I actually learned something from the detailed description(s) that regex101 provides. (I learned that for (\1)+, "Note: A repeated capturing group will only capture the last iteration. Put a capturing group around the repeated group to capture all iterations or use a non-capturing group instead if you're not interested in the data")
- jhight 13y agoSpoiler alert (201 points): f[ao][no]
- onaclov2000 13y agoGlob: 277, it didn't match all and not match the others, but it's reasonably high. ^([bcdlmpwr]|\*[efptv])
- moron4hire 13y ago378: ^(\*(er|[fiptv])|b|c(?!a)|do|le|mi|p|re|w)
- chingjun 13y ago379: ^\*(er|[fiptv])|^([blpw]|c[hor]|do|re|mi) and it has "do re mi" in it!
- moron4hire 13y ago380: ^.[^bds].*[^e-kjotz] .* [^eiz].+[^lx]..$
- hadrel 13y agoGlob (333) without cheating (replace ⁕s with asterisks, they get turned into italics): ^(\⁕?)(\w⁕)(\⁕?)(\w⁕)(\⁕?)(\w⁕) .⁕ ((.(?!\1))+|\1)\2((.(?!\3))+|\3)\4((.(?!\5))+|\5)\6$ Edit: ((.(?!\1))+|\1) is used to conditionally match .+ iff a * has been found. .(?!\1) Matches any character if it is followed by \1. When * has been found then it matches no character, when * is not found it matches every character. Edit 2: Formatting to avoid the *s becoming italics :/
- josephlord 13y agole[^*]|co|ito|dr|^p|su|gi|nr|hw|fa|[eo]b|ide Glob 376 although it isn't pretty and better may be possible. Indent two spaces with a blank line above to avoid code mangling.
- deleted 13y ago[deleted]
- falsedan 13y agoGlob 380 ^([wlpb]|c[hor]|do|re|mi|\*[pifvt]|\*er)
- GregP91 13y agoen(?!c)|rr|il|eat|^(do|p|b|c(?!a)) Glob 386 although not the cutest way to do it.
- jpsim 13y agoGist with my answers: https://gist.github.com/jpsim/8057500 https://gist.github.com/jpsim/8057500 If you look at the revisions, you'll see my 1st iteration was mostly identifying patterns, then with more and more cheating (and looking at this thread) to squeeze every point possible.
- chingjun 13y ago5. 193 ^(?!.*(.)(.)\2\1) 11. 379 ^\*(er|[fiptv])|^([blpw]|c[hor]|do|mi|re)
- galen_tyrol 13y agohttps://gist.github.com/jonathanmorley/8058871 https://gist.github.com/jonathanmorley/8058871 3121 points
- johnlbevan2 13y agoTernstroemiaceae contains four es; anyone else hit that issue / know what that's not in the valid results for Four? I'm guessing there's a pun in there that I didn't get :/
- elwell 13y agoI figured out Ranges! abac|accede|adead|babe|bead|bebed|bedad|bedded|bedead|bedeaf|caba|caffa|dace|dade|daff|dead|deed|deface|faded|faff|feed Edit: /s
- abus 13y agoPrime: ^x{2,3}$|^x{5}$|^x{7}$|^x{11}$|^x{13}$|^x{17}$|^x{19}$|^x{23}$|^x{29}$|^x{31}$|x{33}
- criswell 13y ago[a-f]{4,} works as well... but I do like your solution!
- danielsamuels 13y ago[a-f]{4} is shorter.
- josephlord 13y ago^[a-f]*$ Is the same length too.
- dsschnau 13y agoMe and a coworker totaled 3079 points. Anyone beat it?
- ekke 13y agoGot hooked and to 3202, but ca 10-20 more points could be gained according to answers here and there: https://gist.github.com/jonathanmorley/8058871 https://gist.github.com/jonathanmorley/8058871 Kudos to the author of the game, good job. PS. http://regexcrossword.com/ http://regexcrossword.com/
- endophage 13y agoSeems very very very broken. The regex "[a]" apparently matches "crenel"
- pomfpomfpomf3 13y agoIt doesn't. The green ✔ indicates that you've completed the task successfully — that is, your regex does not match "crenel".
- shawabawa3 13y agoIt's a bit confusing. Ticks on the left means it did match, ticks on the right means it didn't
- scott_karana 13y agoFor Abba... Why doesn't (.)(.)\2[^\1] work? I thought backreferences matched the captured literal, so negating it would match? But this looks the same as (.)(.)\2\1...
- annasaru 13y ago(.)(.)\2\1 ?
- goldenkey 13y agoYou cannot negate a capture, only a character literal. A capture might only be a character but it is NOT a literal,
- scott_karana 13y agoGotcha. That makes perfect sense. Thank you! :)
- goldenkey 13y agoYou got it :-)
- xarien 13y agoWhy is Kesha an optimal answer? (#2 k$) ;)
- amix 13y agoJavaScript solver: https://gist.github.com/amix/8063003 https://gist.github.com/amix/8063003 ;-)
- hyp0 13y agochallenge: use machine learning to find the best solutions. They might improve on those intended by exploiting accidental regularity in the corpus - though charmingly, the golf-cost of regex length helps combat this overfitting. They might also find genuinely cleverer solutions.
- cpeterso 13y agoWhat purpose do the "plain strings" serve?
- ColinDabritz 13y agoHrm, on number 8 "Four" using: (.)(.*\1){3,} I got all but the "do not match" for "Ternstroemiaceae" The challenge appeared to be to match words with four instances of the same letter. "Ternstroemiaceae" contains four 'e's, and thus should be in the "match" column, instead of the "don't match" column, no? Did I miss something?
- denkfaul 13y agoLook closer, there's something different about the matches and this word.
- sbirch 13y agoAn interesting bit on the computational complexity of solving this problem (with a slightly different scoring function): http://cstheory.stackexchange.com/questions/1854/is-finding-the-minimum-regular-expression-an-np-complete-problem http://cstheory.stackexchange.com/questions/1854/is-finding-...
- sriharis 13y ago^ answers all questions.
- ddebernardy 13y agoDoesn't seem to do anything on an iPad...
- JadeNB 13y agoWhat are the rules? That is, are these Perl regexes, POSIX regexes, …? (Come to that, what is this site? Going up one level to alf.nu gives me a lot of suggestions for what I can do by modifying the address, but no clue of who's doing it on my behalf.)