3 ms·
The first part of this is fine, but the last sentence is not. Confidence can be extremely useful in AI. Is the “accuracy confidence” metric you’re using right e
by joxel 4y ago
The first part of this is fine, but the last sentence is not. Confidence can be extremely useful in AI. Is the “accuracy confidence” metric you’re using right every time? No. But using the correct techniques you can create a confidence metric that takes a model with high accuracy (~99%) and couple with a confidence score that is oftentimes right and tuned to specifically flare up when the 1% of bad results come in, and you can be looking at a model that increases its overall prediction capability by an order of magnitude or two. Confidence is a huge part of usable AI.
- godelski 4y ago> The first part of this is fine, but the last sentence is not Your following explanation suggests you have misunderstood me. > Confidence can be extremely useful in AI. Is the “accuracy confidence” metric you’re using right every time? No. >> They indicate how confident the model is of the result, not how likely the prediction is to be accurate. What I said here does not mean that model confidence is not useful. It simply means that model confidence is different than the real world probability of an answer being correct. Allow me to use an example: Suppose we train a model on coin flipping. If we ask the model what the probability of flipping a heads is the model might say "51% with a confidence of 99%". The model has learned through observation and is going to be biased by good/bad "luck" in observation sampling. Of course this example is something we can solve analytically but the disagreement in the analytical solution and statistically learned solution are precisely what I'm trying to explain. I agree, confidence is a useful metric in ML. Very useful. As someone who does generative modeling I can attest to its usefulness first hand. But models are statistical and we need to have a deep understanding of the limitations of our metrics and be careful in understanding what our metrics precisely mean. That is all I mean.
- joxel 4y agoThis makes sense where you're coming from. I still think you're doing a bit of disservice to some of the uncertainty estimation methods for DL. Batch normalization, ensembling, dropout etc. have shown to be very close approximations for Bayesian methods which, as far as I know, are probably the closest we get to mathematically modeling "real-world probability". If you have information that goes against this I'd love to hear it because I'm just relating my understanding of the situation at this time.
- godelski 4y ago> If you have information that goes against this I'd love to hear it because I'm just relating my understanding of the situation at this time. I do and I wrote it. My example is not one that is specific to DL or any type of architecture. My example works for any arbitrary statistical model. What you're missing is a nuanced difference between word definitions. Likelihood and probability are not the same thing. (You're also clearly misreading those papers and have not dug much into DL because there's tons of models that work on modeling probabilities directly, and this is specifically the area I work in)
- joxel 4y agoDo you have some papers of DL that models probabilities directly? You're not just talking about Bayesian Neural Networks right? I've had ideas for having a neural network predict both its target and its estimation of error (by using the actual residual in the loss function to compare to predicted error estimation). Is that similar to what you're talking about?
- godelski 4y agoA few thousand in fact. Look for Ian Goodfellow's taxonomy. I'd have more patience but you're demonstrating a clear lack of domain knowledge but spoke with extremely high confidence earlier.
- joxel 4y agoHas anyone ever told you that you’re an asshole? Imagine telling someone you don’t want to talk to them because of hubris while talking to them like you currently are. I also work in DL, I’m an early career employee and have some experience and am interested to learn. I also have publications involving combinations of uncertainty methods and PINNs and show that certain applications of uncertainty modeling do work in those environments for practical solutions. The way you are describing things is just a little different than what I’m used to, and you decided to use that to insult and belittle me.