7 ms·
Over the years my heuristic has turned into: "Did the team formulate their problem as a supervised learning problem?" - If not it's probably BS. In longform if
by formalsystem 7y ago
Over the years my heuristic has turned into: "Did the team formulate their problem as a supervised learning problem?" - If not it's probably BS.
In longform if anyone is interested https://medium.com/@marksaroufim/can-deep-learning-solve-my-problem-a-type-theoretic-heuristic-e57f4d1658f https://medium.com/@marksaroufim/can-deep-learning-solve-my-...
EDIT: I would consider autoencoders, word2vec, Reinforcement Learning examples of turning a different problem into a supervised learning problem
EDIT 2: Social functions like happiness, emotion and fairness are difficult to state - you can't have a supervised learning problem without a loss function
- sanchezdev 7y agoI basically agree with this rule. I find that my colleagues who overly hype unsupervised approaches typically don't have much experience working on ML problems without labeled data. My suspicion of this comes from the fact that whenever I give a talk on ML I always have a wealth of personal experience to draw on for examples. My colleagues almost always reuse slides from projects they never worked on.
- sillysaurusx 7y agoI'm a little surprised to see this sentiment. Some of the most important advances in the field have been unsupervised tasks: - OpenAI: Dota 2 (PPO), GPT-2... - NVidia: StyleGAN, BigGAN, ProGAN...
- sanchezdev 7y agoThose are certainly important advances, but they don't really apply to most business needs for AI or ML.
- MiroF 7y agoI work in the industry on NLP tasks. Unsupervised learning has been behind the largest developments in the last decade in the field.
- sanchezdev 7y agoI don't disagree with your point, but the unsupervised aspect of NLP typically isn't useful on its own. Usually it's a form of pre-training to help supervised models perform better with less data. From Google in 2018: "One of the biggest challenges in natural language processing (NLP) is the shortage of training data. Because NLP is a diversified field with many distinct tasks, most task-specific datasets contain only a few thousand or a few hundred thousand human-labeled training examples. However, modern deep learning-based NLP models see benefits from much larger amounts of data, improving when trained on millions, or billions, of annotated training examples. To help close this gap in data, researchers have developed a variety of techniques for training general purpose language representation models using the enormous amount of unannotated text on the web (known as pre-training). The pre-trained model can then be fine-tuned on small-data NLP tasks like question answering and sentiment analysis, resulting in substantial accuracy improvements compared to training on these datasets from scratch."
- MiroF 7y agoAs I said, I'm an NLP researcher and practitioner, so you don't need to quote this at me. The unsupervised aspect is the engine driving all modern NLP advancements. Your comment suggests that it is incidental, which is far from the case. Yes, it is often ultimately then used for a downstream supervised task, but it wouldn't work at all without unsupervised training. Indeed, one of the biggest applications of deep NLP in recent times, machine translation, is (somewhat arguably) entirely unsupervised.
- sanchezdev 7y agoI didn't mean to make it sound incidental although I do see your point. Just wanted to chime in with how important having a labeled dataset is for a successful ML project.
- fspeech 7y agoI think the point is labeling itself is very difficult except for special and limited domains. Manually constructed labels, like feature engineering, are not robust and do not advance the field in general.
- Der_Einzige 7y agoEcho the other commentator. Unsupervised techniques are the only reason NLP works as well as it does.
- zamubafoo 7y agoI would argue that GAN's by definition aren't unsupervised, they just aren't supervised by humans. Additionally, OpenAI's game stuff also has similar arguments against it.
- MiroF 7y ago> I would argue that GAN's by definition aren't unsupervised You can define the terms how you want - but in terms of how they're understood in both industry and academia, you are incorrect.
- hervature 7y agoThe discriminator definitely is supervised but the generator is unsupervised. I.e., it has no labels on its targets.
- sillysaurusx 7y agoI'm not sure that's correct. The discriminator and the generator both learn to match a training set. You don't need to label the training set at all. You can just throw 70,000 aligned photos at it. I think I see what you're saying, but that might be a different definition of "supervised". It seems impossible for one half of the same algorithm to be supervised and the other to be unsupervised. But I like your definition (if it was renamed to something else) because you're right that the discriminator is the only thing that pays attention to the training data, whereas the generator does not.
- drongoking 7y agoYou miss the point of the slides. His point isn't about supervised vs unsupervised, it's about the general areas where AI seems to excel and fail. It's being used to predict social outcomes where it does very poorly, may be inscrutable, and is unaccountable to the public. Your examples (deep learning applied to perception) are what he argues AI is generally good for.
- gok 7y agoI don't follow that. The recidivism predictor was supervised. Conversely, AlphaZero is unsupervised and certainly not BS.
- hervature 7y agoAlphaZero is not unsupervised. It is a reinforcement learning algorithm, it knows exactly what the outcome of the game is.
- nokcha 7y agoThe terms "supervised machine learning" and "unsupervised machine learning", by their ordinary English meaning, make it sound like all machine learning is partitioned into one or the other. But a lot of the literature in machine learning considers reinforcement learning to be neither 'supervised learning' nor 'unsupervised learning'. See, e.g., section 1.1 of [1]. [1] Richard Sutton and Andrew Barto, Reinforcement Learning: An Introduction, second edition. MIT press, 2018.
- fyp 7y agoThere's a lot of gray area between unsupervised and supervised learning. For example self-supervised learning: https://www.facebook.com/722677142/posts/10155934004262143/ https://www.facebook.com/722677142/posts/10155934004262143/
- confrnz 7y agoIronically, the algorithm you pose in that comment itself, is a BS algorithm in it of itself. "Formulate the problem as X" - what is your input for how a problem is formulated? That you personally like how it was formulated? "Probably," - OK, so you assign probability scores? Or do you mean, "likelihood based upon my guess?" Finally, how do you measure performance? Your own assessment of how good you were at it?
- xeRTRex 7y agoAuto-encoders have been more successful in fraud and anomaly detection then supervised methods. For the uninitiated: the basic concept is to reduce the feature space (i.e. the things you know) to a lower dimensional space, then decode back into the original space. When enough differences arise between the original and reconstructed variables, the event may be flagged for a human to review (or some triage process).
- amelius 7y agoI wonder if a similar approach can be used for a classification task where one or more classes have only few training examples (those would be similar to "anomalies", I suppose).
- chimi 7y agoIt's hard to verbalize this, most of it is "intuition" but I think it boils down to "supervised learning is BS." Humans are smarter than computers. How can a human teach a computer how to do something when the human itself can't teach another human that something? We haven't solved that problem. The snake is eating its tail. You can't teach a human how to do something when the methodology to do that is the student trying something and the teacher saying "Yes" or "No". Well.... why? Why is it yes or why is it no? What is the difference between what the human or the computer, or in general, the student, did and what is good or correct? And then you still have to define "good" and many times that means waiting, in the case of the PDF linked to above, perhaps many years to determine if the employee the AI picked, turned out to be a good employee or not. And how do you determine that? How do you know if an employee is good or not? We haven't even figured that out yet. How can we create an AI to pick good employees if human beings don't know how to do that? Supervised learning isn't going to solve any problem, if that problem isn't solved or perhaps even solvable at all. In other words, over the years, my heuristic has turned into, "Has a human being solved this problem?" If not, then AI software that claims to is BS.
- wilg 7y agoWell... why is it necessary that we can teach a human to do something in order to teach a machine to do it?
- mark-r 7y agoTeaching a human is a heuristic for understanding the problem well enough to teach a machine.
- chimi 7y agoI agree and rather than post a sibling response, I'll add that I think it's necessary today, simply because we don't have AGI, yet. And also point out that we are talking about determining if AI is snake oil or not. There may be some scenarios where we can teach a computer to do something we can't teach a human to do, I can't think of any off the top of my head, but if we can't, then I'm going to be super doubtful that an AI software can do it better than a human, if at all. AGI, in the singularity sense, will be solving problems before we even identify them as problems. Experts in a field can do this for the layman already and I think it's possible. Some don't. I do. It'll be super interesting when it flips! When the student becomes the master and we, as a species, start learning from the computer. You can kind of get a sense of this from the Deep Mind founder's presentation on their AI learning how to play the old Atari game Breakout. He says when their engineers watched the computer play the game, it had developed techniques the engineers who wrote the program hadn't even thought of. Even still, the engineers could teach another human how to play Breakout, so yes, I do believe they did in fact create a software to play Breakout better than they could.