4 ms·
I'm trying to make social media moderation more democratic and using that to decide fuzzy questions like "should this post be censored", or "is this misleading"
by evandwight 5y ago
I'm trying to make social media moderation more democratic and using that to decide fuzzy questions like "should this post be censored", or "is this misleading" [0]. While the crowd's answer won't be perfect it will help sort through a lot of the noise and feel better than the decision of whatever mod happened to create the subreddit.
The problem: how can I make decisions based on a sample with a binary question. I think the central limit theorem applies, and I need to account for various priors and missing votes. Is there an existing solution to this problem? The server is written in nodejs, if that matters.
[0] - https://efficientdemocracy.com/about/what-is-this https://efficientdemocracy.com/about/what-is-this
- chriswarbo 5y agoThis might be a good use-case for the "bayesian truth serum" http://economics.mit.edu/files/1966 http://economics.mit.edu/files/1966 This applies when our questions are not just trying to learn about the world (e.g. 'Our survey discovered that 10% of posts are considered misleading'); we are going to use their answers to decide on actions, e.g. removing posts, attaching warning labels, etc. Those answering the questions know this, and (if they have a preference over which action will be taken) are incentivised to give more extreme answers. A classic example is an ice cream company surveying shoppers about the flavours they like: if I truthfully answer that I like chocolate slightly more than strawberry, this will have a small effect on the survey result, and hence the company's new product flavours. However, if I falsely say that chocolate is the best flavour I've ever encountered, and that strawberry makes me vomit, that will have a much stronger effect on the survey result, and make it more likely that the company will make the chocolate ice cream that I prefer. The "bayesian truth serum" counteracts this by asking each question in two parts: there's the initial question we want answered, as well as an additional question: "how do you think others will answer?". For example: - "I find this misleading" and "I think 80% of respondents will find this misleading" - "I rate strawberry as 4/5" and "I think 10% of respondents will give strawberry 1/5; 20% 2/5; 50% 3/5; 15% 4/5; and 5% 5/5" The first answers (the ones we care about) are weighted based on two conditions: how closely the estimated distribution matched the real answers, and how 'surprisingly popular' the first answer is. To see why this cancels-out the incentives to lie: our best chance of affecting the result is to choose a 'surprisingly popular' answer, since this will contribute more weight to the result. However, these two constraints exactly cancel out: - The answers we predict are popular, will also be those we predict are unsurprising (after all, we could predict them!) - The answers we predict will be surprising, will also be those we predict are unpopular (that's why it would be surprising if they were popular!) It turns out that the rational strategy, for swaying decisions as much as possible towards the outcomes we want, is to answer the first part truthfully. A similar analysis applies to answering the second question (the estimates) truthfully. In that case there are two things to consider: - We want our estimates to be as close as possible to the true distribution, in order to maximise our response's weight. - We want to engineer our estimates such that the answers we disagree with get a high estimate, and hence appear 'unsurprising' (reducing the weight of those responses). Our estimates must sum to 100%, so decreasing the 'surprisingness' of one answer must increase the 'surprisingness' of the others. The effect we have on each answer's weight will be small, but it will affect every response which chooses that answer. Hence to have the largest impact, we need to decrease the 'surprisingness' of those answers we think will get the most responses. Yet that exactly what we've been asked for (an estimate of how popular we think each answer will be!)
- evandwight 5y agoThat's a very interesting system! I've been reading about it but I'm not sure how it applies. > (if they have a preference over which action will be taken) are incentivised to give more extreme answers. With yes and no answers, how can answers become more extreme? If you are asked "Is this misleading? Yes/No" and it's only marked as misleading when "Yes" is the majority, then you are incentivized to answer with your true opinion. If you want the post to be marked misleading, then answering yes increases the chance that it is marked as such. The bayesian truth sereum information score makes sense when you are trying to reward people for truthfully answering your questions; for example by paying them [0]. When asking "Is this misleading?", how do you use the information score to compute who won? [0] - http://www.eecs.harvard.edu/cs286r/courses/fall10/papers/DW08.pdf http://www.eecs.harvard.edu/cs286r/courses/fall10/papers/DW0...
- chriswarbo 5y agoFor yes/no questions I think (but haven't checked the math) that the incentive is to shift the distribution closer to my opinion. If my opinion is, say, 75% that it's misleading, then the truthful response would be a coin toss with bias 75%. However, if I know my answer will affect censorship, etc. then I may try to predict the resulting distribution, and vote yes if I predict it's less than 75%, and no if I predict it's more than 75%. For example, I may be more "trigger happy" if I think people are more likely to believe something uncritically; I may be more of a "devil's advocate" if I think something is under-represented, or less likely to be taken seriously.
- a1369209993 5y ago> how closely the estimated distribution matched the real answers > A similar analysis applies to answering the second question (the estimates) truthfully. How does this avoid (or compensate for) downweighting the preferences of people who are legitimately ignorant about what everyone else thinks (and consequently give estimated distributions that hardly match the real answers at all)?
- chriswarbo 5y ago
- cdaringe 5y agoI worked on a similar idea last year. What I did was take urls to content, scrape the content, and pipe it through a machine learning a evaluator to apply various labels and warnings to content. Lastly, add some nice embeddable UI to surface the report. I got it to a decent state, but didn’t know how to propagate it or inject it into social communities. I wanted people to be able to tag it on Facebook, and it would reply with an informational card with the analysis and summary. https://github.com/dino-dna/informed-citizen https://github.com/dino-dna/informed-citizen
- evandwight 5y agoCool! Was it supervised learning? I feel like machine learning isn't at the level where it can tell if something is misleading, unless it's from a known sketchy source.