3 ms·
Well, this is stupid and darkly hilarious. I'm only slightly ashamed to admit I created something similar, albeit much less sophisticated, several years ago. B
by nyx 4y ago
Well, this is stupid and darkly hilarious. I'm only slightly ashamed to admit I created something similar, albeit much less sophisticated, several years ago.
Before Reddit started worrying about advertiser friendliness and cleaned up its act, there was a thriving network of hate subreddits. Probably a lot of overlap with the /pol/ population, based on the amount of overt, disgusting racism to be found there.
Anyway, I wrote some crappy Python that would go visit my carefully curated list of racist cesspool subreddits, hoover up all the post and comment text, and add it to the corpus, then some more crappy Python that would ingest the corpus and do Markov chain stuff to spit out some fairly convincing internet hate speech. I think the key to my success was that frothing racists in comment threads typically aren't putting forward the most cogent arguments anyway, so it's a pretty low bar.
I didn't post this little project or write it up anywhere, because I felt bad enough having brought it into the world, but it was good for a chuckle, at least for a little while.
- geocar 4y agoDid you know you can run the chain in-reverse and turn it into a filter? Finding a good threshold one-sided can be hard, but this is basically how a lot of spam-detectors work: They record the chains seen in /g/ and the chains on /pol/ and now you can make statements about which board the comment probably belongs on with, simply by doing some analysis on the frequency of chains seen in one corpus versus another.