4 ms·
I can't imagine it, just like I can't imagine why you'd want to do it in a DSL oriented around set theory instead of, say, the DB implementation language. I ass
by scottlocklin 6y ago
I can't imagine it, just like I can't imagine why you'd want to do it in a DSL oriented around set theory instead of, say, the DB implementation language. I assume they're thinking of deep learning? Why you'd want to do it "in database" is beyond me.
The only reason I thought of doing it was skipping the "marshall the data" step for online algorithms; performance basically. If you look at something like Vowpal Wabbit, it stores/consumes data in this ridiculous text format dating from when people invented the Support Vector Machine in the 80s. This annoys the shit out of me, as it involves creating the ridiculous 80s text format. What if it consumed database columns instead of this nonsense, and did the hashing trick and everything you needed to solve the problem? That would be cool. That, in fact, would be something like how the universe is supposed to function instead of the stunted grotesqueries we have today. You could do crap like exploratory analysis right on your database, as a query. You could even get fancy and do wackadoo online matrix decompositions while you're writing the data out in the first place (or at least when nobody's looking), and store it as metadata, meaning you know all kinds of good shit about your data even as you're writing it down. Anyway, because marketing departments keep bellowing about "deep learning" instead of the actual breakthroughs in machine learning and linear algebra of the last 20 years, nobody gave a shit about it. Even (large research group in gigantor corp) couldn't figure out a way of selling the idea. I went on to a productive career in something entirely different, and all I got out of it was the ability to make snarky comments about seemingly clueless academics.
- kohlerm 6y agoWhy would you want to do the Machine Learning within the DB anyway? To me this looks like rather a corner case. E.g. I usually want to scale my MLE infrastructure independent from the DB. To it seems much to pipe the Data into log like Kafka to compute the result, just because that is much easier to scale. Similarly for training it seem to be that using some files on an Objectstore as the input is usually much easier to scale.
- scottlocklin 6y agoI guess if you don't care about data exploration, efficiency or doing work on data much larger than memory, there's no reason to do it. As you note, most people seem to get along fine without this idea.
- newdude116 6y agoHm. Maybe because you could do it in the RAM directly in the future? Maybe you want to use an encrypted DB without decrypting it? This is research, not a vanilla solution.