4 ms·
> A team at Facebook reviewed thousands of headlines using these criteria, validating each other’s work to identify a large set of clickbait headlines. What a
by jackfrodo 10y ago
> A team at Facebook reviewed thousands of headlines using these criteria, validating each other’s work to identify a large set of clickbait headlines.
What a soul-crushing job that would be.
- pigscantfly 10y agoI'm extremely surprised they didn't outsource to Mechanical Turk or leverage their user base in some way for crowdsourcing the labels.
- darkstar999 10y agoThey may have. Did they release any technical details on this?
- riebschlager 10y agoIt was actually only number six on our top ten soul crushing jobs list. Number one will surprise you! Click here!
- lostgame 10y agoOwch. That was painfully good.
- gohrt 10y agoLooks boring, won't click Fun fact: https://medium.com/i-data/29-reasons-youre-reading-this-article-fbf4671327e3 https://medium.com/i-data/29-reasons-youre-reading-this-arti... "29 reasons you’re reading this article or why odd-length BuzzFeed listicles perform better than even ones"
- wldcordeiro 10y agoSeems like a similar thing to what goes behind pricing something as 1.99 instead of 2.00
- goodJobWalrus 10y ago> Looking at ten thousand published BuzzFeed listicles over a period of three months I found a statistically significant difference in the performance of odd-length listicles compared to even ones. They published at least 10,000 listicles in 3 months! now, that is what I call a soul-crushing job (writing those articles)
- goshx 10y ago- You wont believe how much money they make - Engineers are shocked with these results - Here is the top 10 headlines they found. #6 will make you cringe - The CEO wrote the sweetest message to them Perhaps they should simply ban buzzfeed or use all of their headlines as examples.
- pc86 10y ago> Perhaps they should simply ban buzzfeed What a kind, just world that would be.
- deleted 10y ago[deleted]
- potatoyogurt 10y agoYou'd miss out on the handful of really good real journalism articles per month that Buzzfeed publishes in that case, which would be kind of a shame. Maybe that's how they justify the expense of doing real journalism -- it makes it harder for other sites to ban their domain.
- paulgb 10y agoHN bans buzzfeed submissions outright, but I would have loved to see the community's comments on this article, for example: https://www.buzzfeed.com/andrewrice/the-fall-of-intrade-and-the-business-of-betting-on-real-life https://www.buzzfeed.com/andrewrice/the-fall-of-intrade-and-...
- quotemstr 10y agoThanks. That was an interesting article. "Roman bookies ran numbers on the election of Renaissance popes until Gregory XIV banned the practice on penalty of excommunication. Around the turn of the 20th century, Wall Street brokers openly traded election futures and newspapers quoted their prices like modern opinion polls. Strumpf estimates that at the peak of this practice, in the election of 1916, around $10 million was bet on these markets — more than $200 million in today’s dollars. By the end of the New Deal era, though, the electoral markets had all but disappeared, due to both competition — modern polling pushed the betting lines out of the newspapers — and legal crackdowns."
- mcpherrinm 10y agoYou only really need a few seconds to classify a headline. Spending an afternoon classifying data with a few colleagues isn't the worst thing ever. You're making a big change to a system used by many millions of people. There's probably no easier way to validate your system is working as expected.
- Kiro 10y agoIf you call that a soul-crushing job you can't have had many bad jobs. Most people would only dream of a "fun" job like that.
- alanh 10y agoIt's also for a great cause :) that helps.
- themihai 10y agoIt may be fun for a few hours but reading clickbait headlines excessively is like eating vomit. I'm sure these people have quotas to meet so I doubt even more the "fun" element.
- lifeformed 10y agoIt would be fun designing a system that can automatically detect those headlines, using machine learning and whatever other tricks you can think of.
- rspeer 10y agoI have a heuristic that gets evaluated over time on randomly-sampled tweets. (It's the core of the Python "ftfy" package, which fixes Unicode mistakes based on a heuristic for whether text "looks right".) I frequently have to read the randomly-sampled tweets for debugging purposes. And, yes, random tweets are often so dumb that my brain slightly regrets the time it spent reading them. But that is far outweighed by the benefit that the Internet is delivering me fresh test data all the time. On the whole, I enjoy the notion that I am converting stupid babble into something somewhat useful.
- themihai 10y agoThat's different than reading randomly-sampled tweets or clickbait headlines all day long, is it? At FB(AFAIU) they have people doing just that which makes sense(to reduce costs) but it's a miserable job regardless the scope(i.e. to make more money for Facebook or help people who `friend` rough entities or both.)
- minimaxir 10y agoOver a year ago, I did a statistical analysis of BuzzFeed's clickbait (http://minimaxir.com/2015/01/linkbait/ http://minimaxir.com/2015/01/linkbait/) and found that it is highly formulaic. In fairness, the linkbait game has changed since then, with Medium posts using calls-to-actions and "just" needlessly in their headlines. Then again, with FB's machine learning expertise, I'm surprised they need a team to manually classify linkbait posts at all.
- room271 10y agoEven with expertise, you need labeling of some kind for supervised learning.
- AndrewKemendo 10y agoAnyone who has built a data set for machine learning systems (labeling images etc...) has done this. It's just part of the way you teach ANNs.
- erdevs 10y agoThis is often required work in supervised learning / any machine learning with labels. Sometimes you get better result with human judgment on the labels, and that means you have to grind through a meaningful corpus of data. If you believe what you're doing is important and beneficial, it doesn't feel like too much of a grind. Ultimately this was probably about a week of work, if every labeling participant independently reviewed ~3-5K articles. Some people might have chosen to mechanical turk this, or to gather feedback from a subset of users and use those labels. Doing that might be a better approach as it might not only be less labor-intensive up-front but also allow the system to be regularly retrained easily (ie by gathering new label information as clickbaiters inevitably try to adapt their headlines). I imagine there were reasons for starting off this way, though.
- room271 10y agoIndeed, I've seen at least one paper where they used Mecanical Turk and aggregated results (by showing the same examples to multiple people) to check for quality. It sounds like they did almost exactly the same thing here. Paper: 'Antisocial Behaviour In Online Discussion Communities' (Cheng et al., 2015) Link: http://arxiv.org/pdf/1504.00680v1.pdf http://arxiv.org/pdf/1504.00680v1.pdf (Search for 'Mechanical Turk' / section Data Preparation -> Measuring text quality.)