4 ms·
The part about users coming up with their own formats for tagging including using "word-dash-word" as a format reminded me of a story. When I was in high schoo
by phaedrus 5y ago
The part about users coming up with their own formats for tagging including using "word-dash-word" as a format reminded me of a story.
When I was in high school I wrote chatbots. One of my earliest started out as a "mad libs" program that accepted three-word sentences and spit out a reply by randomly picking one remembered word from each "category" (of the three words).
My intent was these three would usually be "Subject Verb Object", however in one of my first lessons in how users will use your software differently than you intend, I was surprised to find my classmates "tricking" the chatbot into learning sentences with more than three words.
My simplistic parser merely looked for there to be exactly two spaces in the input sentence (along which to split it), and my friends blithely ignored this limitation by just putting dashes in the input like, "this-sentence has-more-than three-words." My first thought was you're using it wrong, the generated sentences won't make sense anymore, etc. But of course the result was more funny and more flexible.
So I looked at what they were doing and thought, ok, how can turn this "abuse" into an intended feature? They're using dashes ('-') to join words together, and the program is using spaces (' ') as places to cut sentences apart. What if the program considered every space between words as both a place to join words and a place to cut a sentence apart?
The second problem is that with three words and three columns in S-V-O order you know at all times, positionally, which bucket(s) you're throwing learned words into and which bucket(s) you're pulling from. How can I take this concept and divorce it from the dependence on fixed positions and fixed number of words?
(Mind you I was 13-14 years old then, self-taught and family and school didn't have internet yet, so this was some deep abstract thinking to be doing without any guidance, experience or examples to go on.)
So I thought ok in the three-word case, let's look at the middle word as it's the most symmetrical case. Assuming you've picked the middle word, now you want to pick the word before it from a list of words that can come before, and the word after it from a list of words that can come after. We can't rely on the absolute position to look up these lists, as it can vary. But we can key the "before" and "after" lists on what the middle word is. So every word gets its own unique list of words-before and list of words-after.
So now we've got two key ideas: learn links between pairs of words (interpretation of what the dash was being used for) and choose from before and after lists (interpretation of what the space was doing). It could be trivially scaled to 2nd-order and higher as well: instead of just learning links between pairs of words, a program could learn links between links, and so on.
Anyway that's how you invent a Markov chat bot from first principles (though I wouldn't learn the name until later). All because some users started putting-dashes to-make-compound-words.