3 ms·
> but in practice it should always be combined with a calibration algorithm (which add a trivial O(n) cost to the process). Why not just use logistic regressio
by otabdeveloper 11y ago
> but in practice it should always be combined with a calibration algorithm (which add a trivial O(n) cost to the process).
Why not just use logistic regression at this point? The only benefit of Naive Bayes over logistic regression is that Naive Bayes is simpler to code.
- _dps 11y agoThe calibration cost is trivial compared to the coefficient learning cost. Very roughly, calibration is O(records) whereas coefficient learning is O(records * features). So the tiny add-on cost of calibration shouldn't affect anyone's evaluation of the relative merits of algorithms. NB still retains its computational advantage. One thing that is often discounted in theoretical discussions is that NB takes much less I/O than something like LR, typically in the range of 5-100x (depending on how many iterations you want to do updating your LR coefficients). If you're doing, for example, a MapReduce implementation then NB has huge computational advantages. In LR each coefficient update costs you another map/reduce pass across your entire data set (whereas NB is always done in exactly one iteration). So if NB + calibration gets you something close to LR for vastly less computation and I/O, why wouldn't you use it? Having said that, if you're talking about small amounts of data that fit into RAM and you can "just load into R", then sure use LR over NB. For that matter use a Random Forest [0]. The reason NB is still around is because it offers a point in the design space where you spend almost no resources and still get something surprisingly useful (and recalibration narrows the utility gap between NB and better methods even more). [0] And you should still consider calibrating your random forest's output.