6 ms·
Generating Text with Markov Chains
- kleer001 6y agoYup, those are Markov Chains. Not to sound snotty or mean, but... so what? Why should we be interested? Looks kinda like you followed a tutorial and did a write up.
- anaerobicover 6y agoIt's a well-written practical demonstration of an interesting thing that [some people might not know about yet][0], with links for further reading. What's the problem? [0]:https://xkcd.com/346/ https://xkcd.com/346/
- not_knuth 6y agoAlso https://xkcd.com/1053/ https://xkcd.com/1053/
- anaerobicover 6y agoDangit, that's the one I meant to link to! I copied the link without checking. Thanks!
- kleer001 6y agoNo problem. Didn't say I had one. Just curious, wanted more information about why OP posted, what made their post special. Still, didn't seem very novel or deep. Might as well have been someone posting their first timing circuit test, water/air tunnel simulation, or computer vision project. So, good for them. But I thought HN wasn't a glad handing cheerleader chorus. Searching for Introduction to Markov Chains gets tons of hits at various levels of explanation that are nearly a decade old.
- healeycodes 6y agoI tried to code up Markov chains when I was first learning to program but I found that many resources didn't have clear/terse enough code snippets. Most articles also didn't walk you from zero to understanding. So I'm trying to fill that gap and help out anyone who is trying to learn the basics.
- gsich 6y agoThe same result generated with "AI" doesn't get this question.
- kleer001 6y agoExamples? IMHO they probably should if they're this basic.
- mkaic 6y agothank you for this clear and instructive write up! i appreciate the high-level, concise breakdown.
- not2b 6y agoI published a program to do basically this on Usenet in 1987. Someone created a fixed version on github at https://github.com/cheako/markov3 https://github.com/cheako/markov3
- fit2rule 6y agoI probably learned Markov from you then.
- spiderxxxx 6y agoI've done something similar without first learning about Markov Chains. One of my more interesting experiments was creating messed-up laws. I fed it the constitution and alice in wonderland, and it made the most surreal laws. The great thing about them is they don't need to know about language. You could make one to create images, another to create names for cities. I made one to create 'pronounceable passwords'. It took the top 1000 words of a dictionary, and then it would spit out things which could potentially be words, of any length. Of course, the pronounceability of a word like Shestatond is debatable.
- monokai_nl 6y ago8 years ago I implemented something like this for your tweets at https://thatcan.be/my/next/tweet/ https://thatcan.be/my/next/tweet/ It still causes a spike of traffic every now and then from Twitter.
- zhengyi13 6y agoThis is a neat introduction to the subject. If you want to see more of what they can do, and you have had any exposure to Chomsky, you might also appreciate https://rubberducky.org/cgi-bin/chomsky.pl https://rubberducky.org/cgi-bin/chomsky.pl.
- inbx0 6y agoFor the Finns reading this, there's a Twitter bot "ylekov" [1] that combines headlines from different Finnish news outlets using Markov chains. Sometimes they come out pretty funny > Suora lähetys virkaanastujaisista – ainakin kaksi tonnia kokaiinia [1]: https://twitter.com/ylekov_uutiset https://twitter.com/ylekov_uutiset
- hermitcrab 6y agoI wrote 'Bloviate' to mess around with Markov chains. From Goldilocks and the 3 bears it prodoces gems such as: “Someone’s been sitting my porridge,” said the bedroom. You can download it here: https://successfulsoftware.net/2019/04/02/bloviate/ https://successfulsoftware.net/2019/04/02/bloviate/
- phaedrus 6y agoI used to write Markov-based chat bots. Something I thought I observed, but tried and failed to show mathematically, is the possibility of well-connected neighborhoods of the graph (cliques?) leading to some outputs being more likely than their a priori likelihood in the input. For example, once a simple Markov bot learns the input phrase "had had" it will also start to generate output phrases a human would assign 0% probability to like "had had had" and "had had had had". This in itself isn't a violation of the principles behind the model (it would have to look at more words to distinguish these). The question is whether more complicated loops among related words can create "thickets" where the output generation can get "tangled up" and generate improbable output at a slightly higher rate than the formal analysis says a Markov model of that order should do for given input frequencies. An example of such a thicket would be something like, "have to have had had to have had". Essentially, I'm hypothesizing that the percentage values of the weighted probabilities of transitions does not tell the whole story, because the high-level structure of the graph has an add-on effect. A weaker hypothesis is that the state-space of Markov models contains such pathological examples, but that these states are not reachable by normal learning. Unfortunately I lacked the mathematical chops / expertise to formalize these ideas myself, nor do I personally know anyone who could help me explore these ideas.
- ben_w 6y agoI have observed the same phenomena with both simple Markov chains and whatever Apple used for autosuggest on the iPhone. That said, your specific example reminded me of a specific joke sentence: https://en.m.wikipedia.org/wiki/James_while_John_had_had_had_had_had_had_had_had_had_had_had_a_better_effect_on_the_teacher https://en.m.wikipedia.org/wiki/James_while_John_had_had_had...
- phaedrus 6y agoI actually hadn't seen that one (I did know of the "Buffalo buffalo" family of sentences). Regarding observation of the phenomenon, in a way I first observed the inverse of it. My first proofs of concept didn't learn probabilities, only structure. (One might say all edges had the same probability.) Yet aesthetically they performed just as well as versions which included probability in the algorithm.