4 ms·
Unless you are using some very exotic method i don't know, you are wrong about discrete classifier not having a score. Everything I know extracts an integer fro
by igorkraw 4y ago
Unless you are using some very exotic method i don't know, you are wrong about discrete classifier not having a score. Everything I know extracts an integer from a vector of scores over options, with the most common being argmax over likelihood trained with CrossEntropy
- fjkdlsjflkds 4y agoNope. Decision trees typically output binary (or multinomial) values directly. Sure, you could tune the node thresholds (or optimize them during tree construction), to vary the trade-off between specificity and selectivity, but this may not be straightforward for deep trees or those with many leaf nodes, and it will never be as simple as simply tuning a single threshold (unless you are using decision stumps, rather than decision trees).
- igorkraw 4y agoHow do you construct your decision tree though? All methods I'm familiar with use a criterion which comes back down to using a score, which also ties down to the set theoretical foundations of probability theory: assign a score and then normalise to to obtain a probability.
- radarsat1 4y ago> Decision trees typically output binary (or multinomial) values directly. Is there some Bayesian version of decision trees where you model each decision in the tree as a beta distribution or something? Then you could simply choose the leaf or path with the highest confidence. I just tried a cursory search about this but came up with a lot of theory justifying the use of thresholds as an approximation to a Bayesian interpretation, but I couldn't find any actual implementations that propagate a confidence measure or probability distribution through the tree, so I don't know if this is standard methodology. I can imagine having confidence in the final decision to also be useful for model ensembling. Curious what the right way to do this might be, or to know if there are no advantages vs. simple threshold trees.
- fjkdlsjflkds 4y ago> Is there some Bayesian version of decision trees where you model each decision in the tree as a beta distribution or something? Then you could simply choose the leaf or path with the highest confidence. Sure... there's many things you could do "force" a decision tree to output a real scalar/vector (that you can then threshold): for example, you could use a regression tree where leafs are logistic/multinomial GLM. My point is that most types of decision trees (including CART decision trees) do not work like that by default (and they are considered pretty standard classifiers in ML literature), unlike what was claimed by igorkraw. > I can imagine having confidence in the final decision to also be useful for model ensembling. Curious what the right way to do this might be, or to know if there are no advantages vs. simple threshold trees. Sure, you could use ensembles of decision trees (e.g. random forests) to estimate a continuous value that perhaps (but not surely) correlates with "confidence" or "probability", which you can then threshold to obtain a binary value. But, again, this is beyond what a "vanilla" decision tree is attempting to do (note: a decision tree is usually not explicitly trying to model the probability of each sample belonging to class X or Y).
- radarsat1 4y ago> (note: a decision tree is usually not explicitly trying to model the probability of each sample belonging to class X or Y). I realize that. Maybe my question wasn't clear. I was specifically trying to ask if there exist interesting or useful models beyond standard decision trees that might do so.
- fjkdlsjflkds 4y agoI think I got your question right (but perhaps I was not totally clear in my answer). In a nutshell: sure, there are lots of such possibly useful models that use trees (you can even come up with new ones yourself). Random forests or GLM regression trees are two examples. My point is that, because standard decision trees do not attempt to explicitly model probabilities (they just try to partition the data in "homogeneous groups", in some sense), you have no pre-baked assurances that your "score" will actually be well calibrated (i.e., no assurances that it will actually be a good predictor of the raw probability/confidence). Sure, you could come up with some "decision tree"-like (i.e., partitioning algorithm) Bayesian method that would (maybe) ensure reasonable calibration of "confidence"... but, at that point, are you even using "decision trees" anymore? Probably not.