18 ms·
What I am wondering, and this is probably a dumb question, is why this has not been automatized? Can't content moderation be done with modern and strong machine
by msapaydin 6y ago
What I am wondering, and this is probably a dumb question, is why this has not been automatized? Can't content moderation be done with modern and strong machine learning based systems? There must be plenty of training data on this, and just like a spam filter which does not require humans in most cases, this should also be automatable. Why is it not?
- amelius 6y agoPerhaps because the training of these systems requires human supervisors?
- msapaydin 6y agowell, just as in spam filters, some supervisors would be necessary to automate things in the beginning, but eventually this should be less needed?
- amelius 6y agoYes, one would hope so. My guess is that we're still in the data-collection phase. It's probably a tough problem because you want the number of false-negatives to be extremely small.
- pixelbath 6y ago>this is probably a dumb question Not dumb, but I do think it's not thought-through. You're proposing a simple solution to what is a hugely complex problem, and throwing ML at it just isn't going to work. To get "plenty of training data," a human still has to classify all of that, leading to the problem of viewing that much unpalatable content by a human. You also have to train your network, which requires humans to verify the accuracy of training, hence viewing the content again. If it were as easy as text-based spam filtering, this wouldn't even be a discussion.
- msapaydin 6y agoI am guessing that by now plenty of data has already been generated which could be used to train future systems..
- pixelbath 6y agoI'm guessing you don't actually understand the complexity involved. How does it differentiate between normal ranting and hate speech? How is hate speech classified in the US vs Saudi Arabia? How does it tell the difference between someone asking a child innocent questions and asking them sexually-related questions on camera? Does the algorithm get trained to flag videos about depression that might lead to suicide, or does it say they're supporting getting help FOR depression and leave the video? What about subjects that aren't already in the corpus of "flag these naughty things"? You still have to get a human to look at those; most likely a data scientist who knows what the algorithm is doing and what needs to be done to correct the training. Machine learning as-is will not get us there, so in the meantime the only option is moderation by people described in the article. It can be outsourced elsewhere, but it's just shifting the responsibility to a different subset of people.
- msapaydin 6y agoCertainly there is more complexity than spam filters, and cultural differences, subtle nuances which are hard to train a system to auto-classify; things that need to be generalized to new cases that arises- that may be wrongly classified in the beginning, etc. Nevertheless such systems can be built from simple cases to more difficult ones step by step and gradually reduce the demand on humans to perform these tasks. Based on some other comments on this thread, I realize FB is already using these methods mostly in languages with plenty of data, and someone also in this thread has made an interesting comment about using machine translation to apply these systems to other languages. I think this will be less of a problem in the near future. (edited)
- ElliotH 6y agoWhat everyone else said, but also the presence of false positives means there'll be appeals. So you have a choice or ignoring the appeals and being accused of censoring people, or processing them - and that tends to need a person.
- fsociety 6y agoThe issue is that this is adversarial in nature. Certainly there are ML systems catching content before moderators.. however since it is adversarial people are pivoting to circumvent the automated systems.
- nitwit005 6y agoIt has been. Most of these sites catch a ton of images and video automatically when they are similar enough to prior known content. That doesn't matter from a staff point of view though. You have a queue to work through. You'll be putting in an 8 hour day dealing with the stuff the system doesn't catch. The automation just means they don't need as much staff.
- three_seagrass 6y agoYep. One way to view automation with machine vision / perception is that it can cover ~85-95% of the true positives. You're still going to get false positives and false negatives that need human review, and at a scale of Facebook, that's a lot of humans.
- msapaydin 6y agoI am just hoping that those currently uncovered cases will be "milder or more nuanced" cases that will be less damaging to the psyche of human moderators and will, once labeled correctly, improve the coverage rate of automated moderators.
- agentdrtran 6y agoIt's not automatable at all, for example Facebook's guidelines change by the day and sometimes by the hour, and many cases are not able to be analyzed by machine as the error rate is too high.
- racl101 6y agoNot to mention the fact, that we'd be subjecting AIs to the worse that humanity has to offer. This is the kind of stuff that creates Ultrons or Skynets who decide that humanity is trash and that it ought to be wiped out. Lol.
- SamuelAdams 6y agoIt's a really hard problem to automate. You can scan photos against a list of hashes to sort out known bad / illegal content, but new content is generated every single day. Also YouTube has increased the amount of auto-moderation but that leads to many legitimate videos being taken down, with little recourse for those legitimate users. Some examples: https://news.ycombinator.com/item?id=22487403 https://news.ycombinator.com/item?id=22487403 https://news.ycombinator.com/item?id=5068626 https://news.ycombinator.com/item?id=5068626 https://news.ycombinator.com/item?id=18742640 https://news.ycombinator.com/item?id=18742640 https://news.ycombinator.com/item?id=19065859 https://news.ycombinator.com/item?id=19065859