2 ms·
The calibration cost is trivial compared to the coefficient learning cost. Very roughly, calibration is O(records) whereas coefficient learning is O(records * f
by _dps 11y ago
The calibration cost is trivial compared to the coefficient learning cost. Very roughly, calibration is O(records) whereas coefficient learning is O(records * features). So the tiny add-on cost of calibration shouldn't affect anyone's evaluation of the relative merits of algorithms. NB still retains its computational advantage.
One thing that is often discounted in theoretical discussions is that NB takes much less I/O than something like LR, typically in the range of 5-100x (depending on how many iterations you want to do updating your LR coefficients). If you're doing, for example, a MapReduce implementation then NB has huge computational advantages. In LR each coefficient update costs you another map/reduce pass across your entire data set (whereas NB is always done in exactly one iteration).
So if NB + calibration gets you something close to LR for vastly less computation and I/O, why wouldn't you use it?
Having said that, if you're talking about small amounts of data that fit into RAM and you can "just load into R", then sure use LR over NB. For that matter use a Random Forest [0]. The reason NB is still around is because it offers a point in the design space where you spend almost no resources and still get something surprisingly useful (and recalibration narrows the utility gap between NB and better methods even more).
[0] And you should still consider calibrating your random forest's output.