3 ms·
We have found in database ML to be great to productionalize basic machine-learning methods. We use BigQuery ML and it saved us a ton of fussing with various fra
by lootsauce 7y ago
We have found in database ML to be great to productionalize basic machine-learning methods. We use BigQuery ML and it saved us a ton of fussing with various frameworks, and because we were already in BQ it saved us the moving of data from database to ML cluster and back into database. We simply author some SQL and get great results quickly.
- mlthoughts2018 7y agoThat strikes me as bad usage of ML algorithms for quite a few reasons. People who say this kind of thing usually are not ML practitioners and want to drink some Kool Aid that they can save money & avoid hiring them with techniques like this. How are you benchmarking the in-DB models against variations that can improve them? If they are basic regression models, how are you solving nuanced problems relating to coefficient estimates that cannot be solved with plug and chug libraries (such as from the article “Let’s Put the Garbage Can Regressions and Garbage Can Probits Where They Belong” by Achen)? I’d venture to guess you’re assuming one particular type of model just works, and either are unaware it is failing, or are unaware of wasted money left on the table by not improving or understanding it. This is extremely common especially with simpler use cases. For example, you might think A/B testing is easily solved as an abstract problem and you just need a framework. What could be hard? I’d argue that A/B testing is a great example of something nobody should be buying from a vendor or getting from unexamined library output, and it’s extremely easy to be losing a bunch of money on junk tests and not even understand your tests are junk. It only becomes much more true with OLS & logistic regression, decision trees, SVMs. You need customized systems operated by people who actually know statistics. Embedding this inside a DB where you cannot seriously develop software to customize the models & analysis just doesn’t work. The cost of serializing the data elsewhere is not the salient one.
- CuriouslyC 7y agoI've never heard of anyone developing IN the database, rather it's a place where you deploy code, so you can do fun things like spectral_cluster((select ...)) or gp_predict_value((select ...)) in the db, avoiding needless serialization and transfer.