4 ms·
The issue with NB for explainability is that the model scores can be very badly calibrated due to the "naive" assumption. You find especially on longer document
by sweezyjeezy 3y ago
The issue with NB for explainability is that the model scores can be very badly calibrated due to the "naive" assumption. You find especially on longer documents, the NB scores basically clump around 0 and 1 due to multiplying a bunch of dependent scores together as if they were independent. This means you can't really use them to assess how 'confident' the model is on its decision.
IMO logistic regression, or better fasttext (rank-limited logistic regression) should be the "default" baseline models you should use for NLP classification. They are trained with cross-entropy loss, so the scores are generally fairly well calibrated to an actual confidence score. Moreover since everything is linear (or at least logit-linear), you can do all the same explainability tricks as with NB (a little more effort for fasttext, but it is doable).