24 ms·
There are ~32.000 tags. No surveillance system is using ImageNet tags to classify people into Buddhists or Not-Buddhist. Most researchers ignore these tags and
by ipsa 7y ago
There are ~32.000 tags. No surveillance system is using ImageNet tags to classify people into Buddhists or Not-Buddhist. Most researchers ignore these tags and focus on a 1000 classes, and know that 32k performance is not good (and these artists have no intention of making it work at all). What they are uniquely trying with this Art Project is as much research as it is activism. Note that "mantrap" is defined in synset as "A trap for catching trespassers", and that you are bound to find weird stuff among over 30k categories (imagine what you can say with the 32% most popular words in French...).
This is a photo in question: https://memepedia.ru/wp-content/uploads/2019/09/imagenet-1.png https://memepedia.ru/wp-content/uploads/2019/09/imagenet-1.p...
This is the route the network took:
person, individual, someone, somebody, mortal, soul (6978) > female, female person (150) > woman, adult female (129) > smasher, stunner, knockout, beauty, ravisher, sweetheart, peach, lulu, looker, mantrap, dish (0)
So it was (politically) correct on the first three categories, and the last one was either a crapshoot (and she could also have gotten to the subcategory of "prostitute" > "streetwalker, street girl, hooker, hustler, floozy, floozie, slattern") or she really is posing in a common "beautiful woman"-way. (The global description for this route is "A very attractive or seductive looking woman" and often triggers for females with tilted heads and lip curls).
You can turn any faces dataset into a labeled face color dataset, so if a black person being subclassified as "negro" is problematic bias or encoded racism, then all such datasets are suspect. Noisy labeled data is the norm, not some horrible exception to be avoided at all costs.
- dannyw 7y agoThe fact that we have a machine learning algorithm that labels _humans_ as smashers, prostitutes, or convicts is a problem, irrespective of whatever technical justification can be made for it.
- ipsa 7y agoBut who artificially created that problem? The artists. There is no marketing company that has intelligent billboards scanning the public for prostitutes. There are no researchers seriously using CV to classify convicts (at least not in the West, and not with ImageNet). That could be a malicious usage problem. This is either a non-problem or Armaggedon for all ML CV datasets, because you certainly can use most CV datasets to train a crappy classifier to output offensive labels. If I train a people photo tagger using a dataset used for combating poaching monkeys in Africa, then who is at fault? Certainly not the researchers who published that dataset with the idea that the data would be used with common sense and scientific rigor, not adversarially -- to make a political point attacking the very existence of that data. The "exposed" bias is trivial. It is the technical justification that should be all that matters for a canonical academic dataset. Science does its best to be apolitical, but then politics ("red bull drinking white men train racist and sexist classifiers") is forced upon it, and we can't really have a productive conversation about bias and ethics anymore. AI needs common sense knowledge of the world to improve. Censorship so science does not offend our sensibilities, would only make it so Google Image Search (a machine learning algorithm) does not return any images of people when you search for "prostitute". Heck, the AI would never learn the difference between a male and a female prostitute. Destruction of accessible knowledge so we (aka: people on Twitter who think AI is the terminator, or the director of the internet) don't get offended by some primitive ML model-as-art-project forced to make errors or awkward classifications. That's a sentence the academic ML community could do entirely without. No benchmarks or duckface selfie would be hurt. No unfortunate third-world souls hired to scan 20 million + internet crawled images for wrongthink, only for the machine to do the unsupervised learning in a hug box, not a black box. Oi mate, you got a loicense fer that label? Just wait until the activists find out who wrote the first 100 8's added to MNIST. Nobody but MIT would be associated with her, if they found out what she did.
- deleted 7y ago[deleted]