3 ms·
>this is probably a dumb question Not dumb, but I do think it's not thought-through. You're proposing a simple solution to what is a hugely complex problem, an
by pixelbath 6y ago
>this is probably a dumb question
Not dumb, but I do think it's not thought-through. You're proposing a simple solution to what is a hugely complex problem, and throwing ML at it just isn't going to work. To get "plenty of training data," a human still has to classify all of that, leading to the problem of viewing that much unpalatable content by a human. You also have to train your network, which requires humans to verify the accuracy of training, hence viewing the content again.
If it were as easy as text-based spam filtering, this wouldn't even be a discussion.
- msapaydin 6y agoI am guessing that by now plenty of data has already been generated which could be used to train future systems..
- pixelbath 6y agoI'm guessing you don't actually understand the complexity involved. How does it differentiate between normal ranting and hate speech? How is hate speech classified in the US vs Saudi Arabia? How does it tell the difference between someone asking a child innocent questions and asking them sexually-related questions on camera? Does the algorithm get trained to flag videos about depression that might lead to suicide, or does it say they're supporting getting help FOR depression and leave the video? What about subjects that aren't already in the corpus of "flag these naughty things"? You still have to get a human to look at those; most likely a data scientist who knows what the algorithm is doing and what needs to be done to correct the training. Machine learning as-is will not get us there, so in the meantime the only option is moderation by people described in the article. It can be outsourced elsewhere, but it's just shifting the responsibility to a different subset of people.
- msapaydin 6y agoCertainly there is more complexity than spam filters, and cultural differences, subtle nuances which are hard to train a system to auto-classify; things that need to be generalized to new cases that arises- that may be wrongly classified in the beginning, etc. Nevertheless such systems can be built from simple cases to more difficult ones step by step and gradually reduce the demand on humans to perform these tasks. Based on some other comments on this thread, I realize FB is already using these methods mostly in languages with plenty of data, and someone also in this thread has made an interesting comment about using machine translation to apply these systems to other languages. I think this will be less of a problem in the near future. (edited)