4 ms·
Solving Wordle using information theory
https://orb.binghamton.edu/nejcs/vol8/iss1/6/ https://orb.binghamton.edu/nejcs/vol8/iss1/6/
- kens 4mo agoWhat I'm interested in is the best starting word. Using Shannon entropy, the paper finds that it is "tares".
- bee_rider 4mo agoI think it depends on whether you use a conventional dictionary as the population of possible words, the valid guess list, or the (known, at the time, for the originally version of the game at least) possible solution list. I guess you could also use the possible solution list minus the words that had already been guessed in previous iterations. I mean, you had to give yourself an artificial constraint because IIRC the next solution was actually built into the page anyway, not obfuscated in any way.
- devNoise 4mo agoA couple years ago, I wrote some js to sort a 5 letter word file. The value for a word was based on letter frequency in each letter location. Then I took the words that didn't have repeating letters. The I choose a word at the top that I liked and grep'd for some other words to use for the 2nd and 3rd try. In 3 tries, you have used over half the alphabet and all vowels. This usually lets me get the word by the 4th try. Here are some word sets you can start with. - pares, moity, bundh|bunch - pares, monty, build|guild - pares, doily, bunch|munch - pores, banty, guild|mucid - pores, canty, build|guild - pores, manty, build|guild - bares, ponty, guild|mucid - bares, moity, punch|dunch - bares, monty, guild|child - cares, ponty, build|guild - cares, moity, bundh|gulph - cares, monty, build|guild
- futune 4mo agoI did this immediately when wordle came out, too. Here's what I got for top 5 starting words along with the entropy: "tears" 6.00570508021708 "teras" 5.95925128791865 "teals" 5.90122552646963 "tares" 5.88711668336419 "lears" 5.86158270234831 Using another metric which I called consistency (I think it was just minimizing the biggest bucket possible afterwards, trying to beat absurdle fast), the best starting word was tied between "arise", "raise", "aesir", "reais" and "serai". All of these have a worst-case scenario of 168 valid words remaining. I find it interesting that different implementations seem to come up with slightly different rankings. Maybe the dictionaries are different?
- ano-ther 4mo agoPaper: https://orb.binghamton.edu/nejcs/vol8/iss1/6/ https://orb.binghamton.edu/nejcs/vol8/iss1/6/
- jezzamon 4mo agoI thought this was old news? I remember people making videos about using information theory to solve Wordle back when it was particularly hyped. (After writing this I checked, there's even a 3 blue 1 brown video on this) My favourite along those lines was solving wordle in 1 guess using the distribution of coloured squares on social media https://www.kaggle.com/code/benhamner/wordle-1-6 https://www.kaggle.com/code/benhamner/wordle-1-6
- GL26 4mo agoYes 3blue1brown made a series of videos which explains it : https://www.youtube.com/watch?v=v68zYyaEmEA https://www.youtube.com/watch?v=v68zYyaEmEA ; https://www.youtube.com/watch?v=fRed0Xmc2Wg https://www.youtube.com/watch?v=fRed0Xmc2Wg .
- seirl 3mo agoAlso same exact name as the paper?
- csande17 4mo agoIt's not clear how the strategy in the article differs from the one used by Wordle Bot, the analysis/feedback system that Wordle links to on the victory screen at the end of the game: https://www.nytimes.com/interactive/2022/upshot/wordle-bot.html https://www.nytimes.com/interactive/2022/upshot/wordle-bot.h... The first page of the published paper ( https://orb.binghamton.edu/nejcs/vol8/iss1/6/ https://orb.binghamton.edu/nejcs/vol8/iss1/6/ ) also claims that the game was developed by "Josh Wordle", so maybe it just isn't the highest-quality scholarship in the world.
- wasabi991011 4mo ago> also claims that the game was developed by "Josh Wordle", To be fair, that is a 1 letter typo; the developer is in fact "Josh Wardle".
- madcaptenor 4mo agoI had wondered if getting wordle in 1 based on social media data was possible! Now I know.
- kps 4mo agoTo paraphrase von Neumann, isn't that how everyone does it?
- _ache_ 4mo agoYeah, exactly. My girlfriend did it too, using information theory. In my solution, I did chose a word that minimized the size of the larger set of possible guesses. It's not strictly information theory, but it's a good approximation of finding 'the word that brings the most information [this turn]'. I suspect that this researchers work involves formalizing and proving the optimality of their solution.
- gkoberger 4mo agoThis isn't groundbreaking science, sure, but it is a great way to get people interested in a topic. After all, it wouldn't be on Hacker News if the word Wordle wasn't in it. I'm a huge fan of using things like this to teach science, math and engineering. We started using "Solve 100 Wordles programmatically" as our technical interview, and people _love_ it. They get really into it and have fun. It's pretty easy to do inefficiently, and it's great to watch people build on it and try to improve their scores. It has two benefits: 1/ everyone clearly understands the problem 2/ people see it as fun rather than a drag.
- tialaramex 4mo agoAnother small benefit: Everybody understands this isn't what the job is. Nobody is hiring you to beat Wordle, and you solving this is clearly not somehow on the path to their actual task, they are asking you if you can write software which solves a problem and you're demonstrating that you can do that, which makes sense. I think "Solve 100 Wordles programmatically" sounds like a lot of work, so that'd probably be a "No" from me unless it was last hurdle for a job I was enthusiastic about but unlike "Write a program to solve this class of graph problem" I at least wouldn't be worried that you're trying to get me to do work for free. Actually Wordle solver as Code Review task sounds like maybe a more interesting live interview than the one we do today. "Here's this mediocre Wordle solver, what is your feedback in review?" has the advantage that they've probably seen a Wordle puzzle before but it's not an example problem they've seen in fifty textbooks.
- gkoberger 4mo agoWell, the 100 Wordles is just "Solve one Wordle" in a for-loop. If you're an even somewhat decent engineer, it takes under 10 minutes to get it working (inefficiently). Then we encourage people to do whatever they want next: improve their average score, build a frontend UI for it, solve on Hard Mode, etc. In the past, we never did technical interview questions like this. We always asked people to bring their own project, and work the way they want to. However, with the addition of AI, we hit a wall: we want people to feel they can use AI in a way that mimics how they'd actually work day-to-day, BUT we also need a simple check to make sure they understood engineering basics.
- plants 4mo agoplug to my blog post where I did something similar a few years back! https://willbeckman.com/wordle.html https://willbeckman.com/wordle.html :). Not sure if this is an identical solution, but it was a fun little project.
- dredmorbius 4mo agoI'd quickly realised that a set of words which covered most of the alphabet (20 words, leaving b, g, j, q, v, and z excluded) allowed solving virtually all Wordle puzzles. The game quickly lost any challenge. wimpy crowd thank fuels Altering order might give faster results. The order presented leaves the most common letters (e, t) for last. Z is quite uncommon, q is virtually always followed by u, similarly common pairs such as ch, sh, and th, as well as three- and four-letter combinations ing and tion, though those won't show frequently in five-letter words of default Wordle. It would be possible to vary word choice based on revealed matches and hits, but if your goal is simply to solve (rather than minimise attempts), the above list works quite well.
- magneticnorth 4mo ago> The game quickly lost any challenge. I only play on hard mode for this reason. My next guess must always be a possible answer based on my current information, and that varies the puzzle enough from day to day that I still find it enjoyable to play occasionally.
- drivers99 4mo agoDoes that not lead you to situations where you have to guess the last letter on each round? Or do you choose words to avoid getting trapped in that situation in the first place?
- copypasterepeat 4mo agoI try to avoid the common traps, but it can be tricky. I also impose this additional constraint on myself, which the game doesn't enforce, that I can't reuse letters that have been marked gray. Sometimes you just can't think of the next word, or might be tempted to use a gray letter because that way you could get more information from other letters, but I avoid using them.
- magneticnorth 4mo ago
- starky 4mo ago>In simulations, their approach solved 99% of Wordle puzzles, while the traditional method solved just 90%. This seems wrong to me, getting a 98%+ solve rate for Wordle is pretty common.
- boothby 4mo agoI crushed wordle within a few days of its popularity entering my sphere. It was pretty easy to brute-force a decision tree minimizing the average number of guesses using a lowly python script and a few days of qpu time. Don't Wordle[1] is significantly more interesting; I've got a solver but the maximum score takes my lowly python script upwards of a day (per day) to solve using brute force. For now, I solve it with a heuristic that terminates in about 20 minutes. My old wordle solver was useful to find a good but suboptimal tree for identifying the answer in 5 undos or less. Today: Don't Wordle 1491 - SURVIVED Hooray! I didn't Wordle today! ..... 8089 ..... 4647 ..... 2492 ..... 1026 .Y... 231 ..G.. 100 Undos used: 3 100 words remaining x 10 unused letters = 1000 total score My puzzle ethics are: you can and should download the dictionaries of valid answers and valid guesses, you're allowed to keep them separate, but you must not keep the list of answers in its original order. [1] https://dontwordle.com/ https://dontwordle.com/
- cvoss 4mo agoAnother great variant is Unfair Wordle [1]. The opponent does not fix the answer upfront but instead evades the player's guesses as long as possible, providing you with the least information it legally can give (according to the usual rules) while still preserving a valid game completion path. The result is that your guesses end up looking extremely unlucky in retrospect. [1] https://tweakimp.github.io/unfairwordle/ https://tweakimp.github.io/unfairwordle/
- arcastroe 4mo agoWoah. I was surprised to win on the first try: cramp, ghost, blind, bulky, bevel, bezel
- cvoss 4mo agoNice! I'm not actually sure what its mechanism is for providing the "least information". It could be smart and reply in a way that maximizes the number of remaining consistent answers. Or it could be greedy and try to report as many "grays" as possible, then as many "yellows" as possible, then resorting to "greens". The latter seems more likely to me, since its easier to implement.
- trollbridge 4mo agoHacker News commenter uses grep -i ^.u...$ /usr/share/dict/words | grep -i c | grep -i -v '^..c..$' | grep -i -v '^...c.$' to crack today's Wordle
- macintux 4mo agoI admit after a half hour of flailing a couple of days ago I got desperate enough to use /usr/share/dict/words. Turns out "emoji" isn't in my copy.
- trollbridge 4mo agoThose are my favourite Wordles. It becomes more of a challenge.
- adamgordonbell 4mo agoIt's trivial to determine the best guess at any point based on what options it cuts out. But I ended up building an alphaWordle, using MCTS and a reinforcement loop just to get a feel for how AlphaGo approach to solving games works. It's not a 'smart' way to solve it, but its pretty instructive and I could compare its moves to the theoretical best move to see it progress. https://github.com/adamgordonbell/bitter-lesson-demos https://github.com/adamgordonbell/bitter-lesson-demos
- senderista 4mo agoIsn’t “maximize information gain at each step” already the standard approach to such problems?
- _whiteCaps_ 4mo agoIf you want a challenge, try https://lirdle.com/ https://lirdle.com/ One letter per line is a lie.
- anigbrowl 4mo agoI think most people understand such problems are analytically tractable, but not everyone understands that the challenge and pleasure of a game is being able to do so unaided.
- waffletower 4mo agoI made a Wordle solver in Clojure a while back which is backed by an English word frequency database. My family believes it is cheating to use it. Much more fun to write that code than to play the game though :D
- Bratmon 4mo agoDidn't like 20 different YouTube channels do this exact thing back when Wordle was relevant?
- aidenn0 4mo agoI wrote a wordle solver that was fairly straightforward; it just brute-forced all of the possible outputs, minimizing the maximum number of remaining possibilities. The very first step was too slow to be interactive (IIRC it took about a minute), but fortunately can be precalculated. With a good first-guess, the number of remaining words is small enough that you can just brute-force. Note that this is not theoretically optimal; you would want more than one level of lookahead for that, but it's good enough to solve almost the entire YAWL dictionary in 6 guesses or less.
- whycome 4mo agoThe NYT world bot that "reviews" your game also has entropy as one of the things it addresses I think. I prefer to not have or use predetermined words. It's most fun to actually feel like you have a chance at solving it in one, and have the challenge of more or less information from first play. It's interesting to compare the solve distribution that the game lists when you're done. I wonder what information can be gleaned from that?
- silentmafia 4mo agoNobody mentions that this was already done by 3Blue1Brown? https://youtu.be/v68zYyaEmEA https://youtu.be/v68zYyaEmEA
- PrimalPower 4mo ago[flagged]
- egberts1 3mo agoI have a tree. I always started with 3 words POINT LASED CRUMB This eliminates not only the most frequent letters but least frequent letters (Bloom filter), leaving rest of infrequent letter set to guess by narrow deduction. Never failed me.