3 ms·
I did something similar recently for a game I'm working on[1]. I didn't know it was called a Markov chain, but one thing I found is that if you take the two pre
by codeka 14y ago
I did something similar recently for a game I'm working on[1]. I didn't know it was called a Markov chain, but one thing I found is that if you take the two previous letters to generate the next one, the results are a little less random and seem a little more natural.
The more letters you take to generate the next one, the closer to the original source data you get, but with a big enough corpus of source data, you can still make random names using three or four letters.
[1] http://www.war-worlds.com/blog/2012/07/generating-names http://www.war-worlds.com/blog/2012/07/generating-names
- AlexeyMK 14y agoYep, Programming Pearls seems to agree with you - 'order-2' Markov chains are usually more natural than 'order-1' chains: http://www.cs.bell-labs.com/cm/cs/pearls/sec153.html http://www.cs.bell-labs.com/cm/cs/pearls/sec153.html.
- pserwylo 14y agoIt was probably still a Markov chain. But when Markov chains are used in this way, they are usually referred to as n-gram models (https://en.wikipedia.org/wiki/N-gram https://en.wikipedia.org/wiki/N-gram) where "n" is the number of previous items you investigate (in your case, a 2-gram or bigram model). Ideally, you would want to go back infinitely many times. My understanding of why Markov models are great is because they reduce the computation required, by not looking back too far, yet they still achieve good results. The more you look back, the more possible combinations (in this case, of letters) there are to consider. As the number of possible combinations of letters increases, the chance of each combination appearing in your training corpus decreases. N-gram models are often used at the whole word level. That is, instead of the "next letter", they are interested in the "next word". This leads to interesting ways to perform spell checking, based on the context of the surrounding words. For example take the following sentences: "I did that to" "I did that too" The to is a spelling mistake, even though the word is an English word with correct spelling. Imagine how many times the phrase "I did that to" occurs in a large enough corpus, compared to "I did that too". My understanding is that Google has the largest corpus and largest "N" (they have a 5-gram model). The cool thing is that they have released it under a CC license (http://books.google.com/ngrams/datasets http://books.google.com/ngrams/datasets).
- user24 14y agoTwo years ago I heard (from a lecturer who knew people at google) that Google were using a 6-gram model internally.
- waterhouse 14y agoI once made a word-based Markov chain and fed it a corpus of my instant messages. I did that to amuse myself. I did that too crudely, though; I didn't give it any notion of beginning or ending a sentence (although I kept capital letters and periods as part of each token, simply because that was easy to do). I don't think there were any other corpuses that I did that to.