Wordle, two ways

Self-directed final project · Academic

Wordle gives you six guesses to find a five-letter word, with feedback for each letter after each guess: green (right letter, right spot), yellow (right letter, wrong spot), gray (not in the word). There are 2,315 possible answers.

Formally, this is a partially observable problem. The hidden state is the answer, each color pattern is an observation, and the set of answers still consistent with everything observed so far is the belief state. A Bayes filter maintains that belief; after each guess, a word survives only if it would have produced exactly the feedback seen. The filter starts at 2,315 words and collapses from there.

This project builds two agents on that shared filter. They differ only in how they choose the next word to guess.

Two ways to pick a candidate

Both agents evaluate a guess the same way: bucket the remaining answers by the feedback pattern each would produce. A guess that splits the answers into many small buckets is a good one, because few candidates will remain for whatever comes back. A guess that leaves one big bucket is a bad one, for the opposite reason. The agents differ only in how they judge the buckets:

The entropy agent is risk-neutral. It treats the bucket sizes as a probability distribution and picks the guess with the highest Shannon entropy — the most informative question on average, assuming the answer was drawn at random.

The second agent is risk-averse. It looks only at each guess’s largest bucket — the worst case — and picks the guess whose largest bucket is smallest. It is minimax in nature: it gives up some performance on average to ensure no feedback pattern ever results in a large group of candidates.

The opening guess

Every game starts from the same belief state, so the first guess is the same every time. Each agent computes its own:

AgentOpenerExpected bitsWorst-case bucketCould it be the answer?
Entropysoare5.886183no
Risk-averseraise5.878168yes

soare (a young hawk) is not in the answer list at all. The entropy agent spends the first turn purely on information. The risk-averse agent has a real tie to settle: five words reach the best worst case of 168 candidates, and two of them (arise and raise) are themselves possible answers. It takes raise, which carries the most information of the five and could also win outright.

Because the opener is deterministic, it is computed offline once and committed as a constant — scoring 12,972 candidate guesses against all 2,315 answers live would cost seconds and produce the same word every time.

Head to head

Both agents solve every one of the 2,315 possible answers, guessing from the full 12,972-word allowed list:

44 1,217 990 63 1 1 67 1,045 1,129 73 1 2 3 4 5 6
Guesses per game across all 2,315 possible answers — the full benchmark, each agent using its own opener and the 12,972-word guess pool.
View as table
Guesses EntropyMinimax
1 01
2 4467
3 1,2171,045
4 9901,129
5 6373
6 10
Average 3.463.52
Worst 65

Neither agent ever loses. The entropy agent is better on average — 3.46 guesses to 3.52 — but the risk-averse agent never needs more than 5, while entropy requires 6 exactly once. Each agent wins on the metric it was built for and gives up ground on the other.

The single word that costs the entropy agent a sixth guess is waver. It sits in a family of answers separated by one letter — wafer, wager, water, waver — and on the fifth turn the entropy agent is left choosing between two of them with nothing to separate them. Walking into that is what the risk-averse agent spends its average performance avoiding.

The ordering survives a change of opener, which suggests it comes from the two objectives rather than from one lucky word. Running both agents on a shared salet opener instead gives 3.43 and worst case 6 for entropy, against 3.48 and worst case 5 for the risk-averse agent. That comparison also shows the cost of choosing an opener greedily: salet beats both agents’ own openers, despite scoring worse on both of their criteria at move one. One move ahead is not the same as best overall.

Against a human baseline — five random games, played blind and then re-run through both agents — the agents matched or beat the human on every word, and the human lost one game outright (humph). A typical human average sits around 4–5 guesses; the agents’ is under 3.6.

Try it

Think of a five-letter word or let the agents race on a random word. The candidate count next to each agent is the belief state; watch it collapse from 2,315 to 1.

Implementation notes

  • The in-browser agents are a TypeScript port of the original Python project, running in a Web Worker. It has the same Bayes filter, the same bucket scoring and the same tie-break order, and it returns the same guess count as the Python project for every one of the 2,315 answers. Feedback patterns are base-3 integers (3⁵ = 243 per guess).
  • The original project precomputes the full 12,972 × 2,315 guess-by-answer pattern table. The port skips the table: with the opening guess precomputed, every later turn faces a candidate set small enough to score in milliseconds.