5 ms·
Wow! This is really cool. In statistics and machine learning this would be considered an unbalanced data set: predicitng who will start a company when the vast
by squigs25 13y ago
Wow! This is really cool.
In statistics and machine learning this would be considered an unbalanced data set: predicitng who will start a company when the vast majority of people will not is a very difficult task. It's similar to predicting who will be a terrorist (another really difficult problem).
I think the threshold they are using is way off however. Even if someone has only a 5% chance of becoming a founder (or less), that's pretty significant. I understand that would probably increase the population by many orders of magnitude, but only capturing 17% of 350 means ~60 startups will be found as a result of this program. Given that the large majority of those are likely to fail, the numbers could be better.
Some really interesting predictors might be what meetup groups does the individual belong to, what is their current job title, what is skills and connections do they have on linkedin and facebook, how many founders are they "connected" to, and who are they following on twitter.
It's also worth mentioning that this is probably biased, because the data set of individuals includes data points for founders only after they became founders. You would ideally want the data from before they became a founder. Perhaps over time this model would get better, as non-founder individuals become founders.
- kevin_morrill 13y agoCTO of Mattermark here. It is a really interesting problem, because as you say even if you boost the odds 25x they're still really low. We trained the data set on venture backed founders (e.g. Series A or beyond), which is a bit higher bar than just any founder. The hope being that once you reach Series A you're less likely to fail than just having seed funding. At some point we want to go back and look at what differentiates founders that reach seed vs. venture backing.
- JasonCEC 13y agoCan you talk a bit about your feature selection or models? I run a statistical quality control company using machine learning, and picking up on flaws with tiny probabilities (one batch in every twenty or thirty million) might benefit from similar techniques!