3 ms·
I work at Google on a product driven by ML doing ranking and regression tasks. Can confirm, very relevant. That said, ML is usually superior to the rules and he
by Xorlev 8y ago
I work at Google on a product driven by ML doing ranking and regression tasks. Can confirm, very relevant. That said, ML is usually superior to the rules and heuristics systems we've been able to come up with, so we take on the debt once we stop being able to improve our heuristics, but only once we've tried really hard at the heuristics such that we have a baseline quality bar to beat. That justifies the effort, but it's still a lot of work to be vigilant and keep an eye on shifts in signals, unintended dependencies, good metrics that mean something, etc..
- jacquesm 8y agoWhat really bugged me for a while is how unbelievably easy it was to beat a very large amount of hand tuned code using ML. Going from 92.x% accuracy to 97% accuracy even without any tweaking at all feels like cheating.
- brlewis 8y agoAre you at liberty to share, accuracy of what?
- Radim 8y agoCan't speak for OP, but such accuracy numbers often hide a 20/20 hindsight bias. After having built and run a rule-based system for a while, you always get tremendous subject matter expertise, a feel for what works. Any rewrite of the system at that point will lead to much improved accuracy. The clarity is reflected in a better choice of the input signals, features, data preprocessing, metrics, workflows… A "magic ML" (without domain understanding) beating well-tuned SME rules is a dangerous fantasy, in any non-trivial endeavour. In other words, without that clarity, you're better off gaining it first through simple iterations of rules, figuring out what matters.
- jacquesm 8y agoLego parts recognition.
- heavenlyblue 8y agoAha, and what you're actually getting is like 99% accuracy with incredibly costly false positives; as opposed to hand-tuned rules which made sure that most of the mistakes made are cheap ones. If you are smart, then what you're doing is probably easily transformable to the set of rules you had before. At that point you can compare why exactly it's so good at the metric you're measuring so there's no "cheating". Sadly most of the ML consultants just take an exemplary code from one of the tutorials and then show you the metric it generates after having run.
- srean 8y agoDo you have more information on jaquesm's ML model than what he mentioned in the comment or was it a sweeping claim about any high accuracy ML system. There is no reason why the consequences of false positives and false negatives cannot be incorporated in the model itself. In fact for certain kinds of systems such as 'alarms', or 'imbalanced classes' this is pretty standard.
- jacquesm 8y agoHe doesn't and the incredibly sure way in which he speaks of a system without any knowledge of the context or the application is an interesting study in how online conversations derail. Anyway, the misclassifications are much the same as with the original system, in fact on the same parts only with far lower incidence so to me it looks as if the ML system simply managed to extract a lot more features (and automatically) than I would have time for to do by hand, on top of that it adapts easier to new, previously unseen content because I don't need to come up with a bunch of (reliable!) rules to tell those parts apart from the previous ones (this does require a complete retraining of the net). For some subset of the problems available ML works very well indeed, for others it may be a marginal improvement and in many cases ML is just dragged in to a project even though it has no place there. If you're in the first category: consider yourself very lucky and reap the benefits.
- _Tev 8y ago> For some subset of the problems available ML works very well indeed, for others it may be a marginal improvement and in many cases ML is just dragged in to a project even though it has no place there How to recognize which problems are well-suited for ML? Are there any rules of thumb for (relative) laymen already?