6 ms·
I know very little about machine vision, so forgive the naïvete of this question: > an ability humans have that AI lacks: the ability to understand when a sce
by hydrox24 8y ago
I know very little about machine vision, so forgive the naïvete of this question:
> an ability humans have that AI lacks: the ability to understand when a scene is confusing and thus go back for a second glance.
Wouldn't the machine return a low confidence score when a scene is confusing? If not, why is this difficult to get around?
If so, why can't we just call this a computation difficulty problem, simply requiring a way to go back and spend more effort later when required? A problem like this would simply require better computers, and patience.
In other words, Is this problem as deeply rooted as the article suggests? Or is it simply a problem with the popular approach to machine vision?
- useful 8y agoEach is harder: * return a result * return a list of results * return a list of results with a confidence weightins * return a list of results with multiple weightings (confidence and complexity) how many if statements would you need to write for each if this was applied to a 3x3 grid like tic-tac-toe? how would you know you are correct?
- ht85 8y agoI think the article is referring to a scene that would be unlikely, not make sense given the context. Like identify a grey mass in the middle of the street as en elephant. "Confusing" would be a poor choice of word here though.
- skierscott 8y ago> a low confidence score Neural nets should return a low confidence score. But, the popular approach (described below) ignores that. Neural nets ignore confidence because of a technique called softmax [1]. This happens as the final operation of a neural net, and is required for training. Softmax is a tool to make an array of positive numbers look like a probability distribution: out = x / x.sum() x[i] is a class prediction, but x.sum() != 1. Say if the network was uncertain, x[cat, dog] = [0.03, 0.01]. These are small values that do not imply great confidence (the network was trained on vectors with out.sum() = 1. The network would predict “dog” using softmax because out[dog] = 0.75 > 0.25 = out[cat]. But then in inference/prediction, the confidence is ignored. What if x.sum() is small? That would imply that the network is uncertain. [1]: https://en.m.wikipedia.org/wiki/Softmax_function https://en.m.wikipedia.org/wiki/Softmax_function
- gok 8y agoNote that you can get a form of confidence by just not applying softmax to the output during inference. Softmax is primarily to aid in training.
- deleted 8y ago[deleted]
- PavlikPaja 8y agoHow well do neural networks train with no normalization at all, compared with softmax?
- CodesInChaos 8y agoYou need to perform some kind of normalization, since probability must be between 0 and 1 (and being wrong on a confident prediction gives huge penalties using the popular maximum likelyhood loss functions). But you can use component wise normalization (sigmoid) instead of combined normalization (softmax). These correspond to the assumption that the classes are independent (component wise sigmoid) or mutually exclusive (softmax).
- PavlikPaja 8y ago"probability must be between 0 and 1" - why? (I get it's used in mathematics, but I see no reason why a NN would have to output probability that way.) "and being wrong on a confident prediction gives huge penalties using the popular maximum likelyhood loss functions" - It should.
- p1esk 8y agoI see no reason why a NN would have to output probability that way For classification tasks, the labels are usually encoded as a one hot vector (one in the position of the correct class output, zeros everywhere else). If you don't normalize outputs to be between zero and one, it becomes a regression task - you are essentially asking the model to fit your one hot encoded label. That's not desirable, because we don't care about the actual value of the output for the correct class. Whether it is 0.1, 1.1 or 1001 it is the correct output as long as it's larger than outputs for other classes. That's why we want to take the largest output, and scale it in a way that it's always less than one. Its distance from one depends on how much larger it is than other outputs (the confidence of the model in this prediction). Without normalization, the model that outputs 1000 for the correct class and tiny values for all other classes would get severely penalized because the labels says it should be 1 in that position (so the error is 1000-1=999), even though the model made the correct prediction. There's some confusion about this (e.g. https://news.ycombinator.com/item?id=18054447 https://news.ycombinator.com/item?id=18054447 ), so hopefully my explanation makes sense.
- 21 8y ago> simply requiring a way to go back and spend more effort later when required Present neural nets have no "more effort" or "less effort" knob. For the same input they always produce the same output. The article says a different thing: humans rerun the "algorithm" again, but this time with the knowledge that something was wrong in the previous run (lets say color), and this knowledge will tweak the way it runs again, maybe even running a completely different algorithm the second time (a network more specialized in color but which is worse at shapes)
- taeric 8y agoNaively, doing an ensemble of different networks is effectively this. I don't know if folks doing that, though. Pretty sure these edge cases don't matter nearly as much as lack of data for most of us. Consider, can you recognize that distant relative you have never seen? Why not?
- p1esk 8y agoYou can easily add a "knob" like this to any model. For example, you can first run an input through a fast, but less accurate object detector (e.g. YOLO), and if not satisfied with results (or automatically triggered by low confidence scores), run the same input through a slower, but more accurate model (RCNN).
- p1esk 8y agoYou can also have a "generalist" model, trained to differentiate major object classes (e.g. furniture from animals), and multiple "specialist" models, each trained to differentiate between variants of the same object category (e.g. breeds of dogs). That approach was described in Hinton's knowledge distillation paper.
- sgt101 8y agoWell - you want to believe the best method you have, if the best method produces a low confidence approach then that's that really - you can't "believe" the next best classifier instead - because by definition it's likely to be wrong. Humans re-appraise the scene in light of their surprise or confusion. We need a cognitive model for vision... This is all leading back to David Marr, who wuda thunk.