5 ms·
I watched a talk by someone in the intelligence community space nearly 8 years ago talking about the data dirt that most companies and spy agencies are combing
by devonkim 7y ago
I watched a talk by someone in the intelligence community space nearly 8 years ago talking about the data dirt that most companies and spy agencies are combing through and the kind of abstract research that will be necessary to turn that into something consumable by all the stuff that private sector seems to be selling and hyping. So I think the old guard big data folks collecting yottabytes of crap across the world and trying to make sense of it are well aware and may actually get to it sometime soon. My unsubstantiated fear is that we can’t attack the data quality problem with any form of scale because we need a massive revolution that won’t be funded by any VC or that nobody will try to tackle because it’s too hard / not sexy - government funding is super bad and brain drain is a serious problem. In academia, who the heck gets a doctorate for advancements in cleaning up arbitrary data to feed into ML models when pumping out some more model and hyperparameter incremental improvements will get you a better chance of getting your papers through or employment? I’m sure plenty of companies would love to pay decent money to clean up data with lower cost labor than to have their highly paid ML scientists clean it up, so I’m completely mystified what’s going on that we’re not seeing massive investments here across disciplines and sectors. Is it like the climate change political problem of computing?
- dgacmu 7y ago> In academia, who the heck gets a doctorate for advancements in cleaning up arbitrary data to feed into ML models Well - Alex Ratner [stanford], for one: https://ajratner.github.io/ https://ajratner.github.io/ And several of Chris Re's other students have as well: https://cs.stanford.edu/~chrismre/ https://cs.stanford.edu/~chrismre/ Trifacta is Joseph Hellerstein's [berkeley] startup for data wrangling: https://www.trifacta.com/ https://www.trifacta.com/ Sanjay Krishnan [berkeley]: http://sanjayk.io/ http://sanjayk.io/
- devonkim 7y agoI was asking somewhat rhetorically but am glad to see that there’s some serious efforts going into weak supervision. At the risk of goalpost moving, I am curious who besides those in the Bay Area at the cutting edge are working on this pervasive problem? My more substantive point is that given the massive data quality problem among the ML community I would expect these researchers to be superhero class but why aren’t they?
- dgacmu 7y ago... they are? There are a lot of people tackling bits and pieces of the problem. Tom Mitchell's NELL project was an early one, using the web in all its messy glory...http://rtw.ml.cmu.edu/rtw/ http://rtw.ml.cmu.edu/rtw/ Lots of other folks here (CMU). Particularly if you add an active learning. Hard messy problem that crosses databases and ML.