3 ms·
I'm curious why FB chose to manually label the data vs collecting feedback from users. Eg via a dialog on a large sample of users and articles asking whether t
by erdevs 10y ago
I'm curious why FB chose to manually label the data vs collecting feedback from users. Eg via a dialog on a large sample of users and articles asking whether the article was clickbait, or asking users to rate it, or some such.
The benefits of this approach, besides a lesser degree of tedious manual data review and entry up-front, include: ease of retraining the system as clickbaiters inevitably adapt in this new arms race (can regularly gather fresh user feedback and feed the updated corpus and labels into the system), and perhaps also closer affinity with what users actually think qualifies as clickbait (as opposed to FB's internal definition). This soft of approach may also lead to more differentiated filtering on a personalized basis or an affinity-group basis... ie they'd have the opportunity to model user-behavior features and create differentiated filters based on user behavior/preferences.
I'm sure there were very good reasons for going this way, so I'm not second-guessing. Just curious what the tradeoffs were in the decision, if any knows or can make educated speculation.
- heydenberk 10y agoIt would be gamed. The people who have the most incentive to take part in a "crowdsourced" labelling exercise to reduce clickbait are, of course, the people generating the clickbait.
- harigov 10y agoNot if they have no control over who is asked to label. You can evaluate how trustable labels are by manually reviewing few random ratings and extrapolating it on the entire population.
- chejazi 10y agoMost people don't consciously think "News Feed is full of clickbait"; instead they value the overall experience less, e.g. "I find News Feed less entertaining/worthwhile than BuzzFeed". They aren't aware of the curation going on under the hood, and soliciting feedback is considered more disruptive than anything.
- ben_jones 10y agoI met someone who was somewhat internet-addicted to buzzfeed. They would spend hours a day going through buzzfeed pages and get a tremendous source of enjoyment from it. So it's understandable that Facebook finds (no matter how hard they have to look for it) the value that users see in such articles and implement a filter that retains as much of it as possible.
- erdevs 10y agoI can see that as a reasonable concern. Two points though: 1. Speaking personally as a user, I wouldn't mind this a bit. And would actually love the ability to provide feedback more readily. I often hate the content of my feed and I would feel much better if it seemed like I had greater input on its filtering. 2. Whether gathering feedback has a negative effect or not seems like a testable hypothesis and this could be tried and measured, rather than simply speculated about. Given the above, I'm not sure if this would be a rigorous rationale for avoiding the active feedback experiment.
- Domenic_S 10y agore: #1, you can. Each post on your feed has a "show less like this" option on it.
- mandeepj 10y ago>I'm curious why FB chose to manually label the data vs collecting feedback from users. FB is collecting data from users in a way. Yes, it is not a feedback. Since it is a learning process so getting feedback for each post is going to be an obtrusive behavior. For more details, please refer to following line in the article - To address clickbait headlines, we previously made an update to News Feed that reduces the distribution of posts that lead people to click and then quickly come back to News Feed.
- erdevs 10y agoI did read that. They note that this passive feedback / preference inference was insufficient in terms of achieving their aims. Which is why I'm asking about potentially gathering active feedback / explicit preference.
- gnicholas 10y agoCrowdsourcing to the users is an interesting alternative. My startup actually considered doing just this in our browser plugin [1]—adding a feature that lets people assign a clickbait rating to a link. We would then get a crowdsourced rating for links, and low-rated links would be grayed out for our plugin users. We ultimately hoped to release the data on rankings publicly, so that a predictive algorithm could be created and used by others as well. This feature aligns moderately with our mission/product, which is about reading efficiently on the web. But we've had other priorities so far and haven't built it yet. I think the reason that crowdsourcing wouldn't make as much sense for FB is that their audience is less early-adopter than ours. Some people wouldn't know what clickbait is. It's not just about whether you like the article or not—it's specifically about whether the headline mischaracterizes or inappropriately teases the content. FB probably decided that they wanted to train the algorithm carefully, so they used an internal team instead of a crowdsourced solution. 1: https://chrome.google.com/webstore/detail/beeline-reader/ifjafammaookpiajfbedmacfldaiamgg https://chrome.google.com/webstore/detail/beeline-reader/ifj...
- erdevs 10y ago> Some people wouldn't know what clickbait is The term "clickbait" need never be used. One could describe the characteristics of clickbait and ask about them. "Do you feel this article was meaningful?" "Did you feel this article required to click too many times?" "Did you feel this article had a misleading headline?" Or whatever more refined version of these questions might make sense.
- gnicholas 10y agoI suppose if you were to ask several questions you could avoid this term. But none of the above questions, taken alone, captures the essence of clickbait. Not sure how many people would be willing to answer 2-4 questions many times over. And people who are familiar with the term clickbait might wonder "why are you beating around the bush—just ask me if it's clickbait already!". And as others have pointed out, clickbait creators might try to find ways to game the survey system to make it less effective.
- thoughtPolize 10y agoyou know how facebook has been in the news for manipulating news? Now they can legally do that. They're telling you upfront that some human has influence over the flow of information to users on their platform. This is marketing it as a feature, giving them full legal immunity.