4 ms·
the right title should be: "The World's Top 10 Most Innovative Companies In Big Data according to their marketing team." The article I would like to read is ab
by genofon 13y ago
the right title should be: "The World's Top 10 Most Innovative Companies In Big Data according to their marketing team."
The article I would like to read is about the real advantages they made, not just "they are applying BIG DATA to the X,Y industries". No description on the size of the problem nor the advantages or details.(I'm especially looking at you AYASDI... )
I've been to Big Data conferences where the main application was mean and standard deviation on a huge datasets, I'm just sick of companies inflating the Big Data bubble (and this comes from a "data scientist")
- michaelochurch 13y agoI've been to Big Data conferences where the main application was mean and standard deviation on a huge datasets, I'm just sick of companies inflating the Big Data bubble (and this comes from a "data scientist") "Data science" is a bizarre job description to me. Some companies' data science teams are doing work that could be done in Excel, and others are doing sophisticated machine learning research. In many companies, though, it's a watered-down version of the (extinct, sadly) R&D job description. I've heard people opine that 95% of "data scientists" have never implemented anything more sophisticated than an SGD regression, and it wouldn't surprise me. Granted, you can get some neat insights and visualizations with off-the-shelf tools, but I still consider it important to know how the algorithms actually work and what the basic assumptions are. For example, you can usually use linear regression for 2-class classification problems (as opposed to the mathematically more correct logistic model, since probabilities are in [0, 1] and linear models diverge at the boundary) get a reasonable predictive model, but it's worth knowing when (and why) that short-cut breaks down. In software, you hear "data scientist syndrome" used to describe people who have a lot of short job tenures because they (a) leave if there isn't interesting work for them, and (b) tend to be the first laid off, not because they're bad but because R&D is first to bleed when things go bad. The software industry forces you to choose between short-term job security (people doing interesting work are most exposed to organizational changes) and long-term career health (people not doing interesting work turn into dinosaurs). It stands to reason that the selection process for organizational credibility (favoring tenure) would be something other than deep knowledge of data science (which requires a stream of interesting work, and the behavioral correlates such as job volatility). It's sad that the typical corporate culture of reliable mediocrity, at the expense of excellence, has also learned to use the vocabulary of "Big Data". Reality: truly Big Data (> 10 TB) is a huge pain in the ass. It's what you have to deal with when there is too much noise, or the model's essential complexity is too high, to get a good model out of "small data".
- dannypgh 13y agoCurious- why would you ever use linear regression for binary classification? Logistic regression can be implemented as linear regression predicting the input to the logistic function.
- oldskoolbob2000 13y agoFrom my experience, it's easier to explain linear coefficients. Also, at least in R, linear regression tends to run faster than logistic regression.
- sgy 13y agoMight be a good read http://mrvar.fdv.uni-lj.si/pub/mz/mz1.1/pohar.pdf http://mrvar.fdv.uni-lj.si/pub/mz/mz1.1/pohar.pdf
- michaelochurch 13y agoLinear regression has a closed-form solution and, even when that is impractical (50k+ features, at which point you start not wanting to do the O(p^3) operations involved) and you have to use gradient descent, it still tends to run faster. If most of your data are in the p ∈ (0.2, 0.8) range, then the logistic curve is approximately linear anyway. Where linear regression fails you on a logistic problem is if you have a lot of data where the linear model will give p's outside of (0, 1). While the linear model's mathematical foundations are "incorrect"-- more precisely, it's not the maximum-likelihood model and may have a conditional likelihood of zero, that is, being impossible as "the correct" model on the data-- the truth is that, at 100k+ features, getting "the right" model is effectively impossible with most real-world data sets and you have to use simplifying assumptions and techniques (e.g. regularization, early stopping) anyway. Linear regression doesn't usually do as well as logistic regression, but sometimes it can, and most people are going to try both.
- dannypgh 13y agoI think I see the source of the confusion. Linear regression does not imply any specific method for solving for your coefficients. You can solve a linear regression by inverting your matrix or you can use gradient descent, to give two mechanisms. Yes, solving for your coefficents based on inverting the matrix can be O(n^3) (although wikipedia tells me there's ways of doing it that are O(n^2.373), but at the end of the day, it's slow for large n) but gradient descent doesn't have this problem, as gradient descent is O(n) with regards to your training set. And if you're implementing gradient descent to build a classifier, you can equivalently do logistic regression by just changing your cost function (that is, you'll be solving a linear regression to give you f(x) such that S(f(x)) is the probability you care about, with S(x) being the sigmoid function). So linear regression can perhaps be easier than logistic regression, but only when n is so small that you're solving your linear regression by inverting the matrix. Once you've switched to gradient descent (or anything that can solve for solutions to arbitrary cost functions) there is no difference.
- sgy 13y agoMean and standard deviation are [measures] to calculate data, but not [applications]. Application is making that data meaningful and organized, in order to push whatever industry forward. It's just what actually IBM, GE et. all. are doing. Still, the article doesn't have to be objective and you might be right but a little bit negative.