4 ms·
PG or Aaron: Could you de-identify YC applications data and make it publicly available for others to analyze? [EDIT:] Or maybe you could make a prediction chal
by zeratul 15y ago
PG or Aaron: Could you de-identify YC applications data and make it publicly available for others to analyze?
[EDIT:] Or maybe you could make a prediction challenge using the YC application data? You could host the challenge on http://www.kaggle.com/ http://www.kaggle.com/ or http://tunedit.org http://tunedit.org
Note to self: There is actually a very sparse body of literature that talks about data driven approach to VC, e.g., http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=4621115 http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=4621..., this one seems interesting (predicting exits): http://www.springerlink.com/content/w024ku34221u3258/ http://www.springerlink.com/content/w024ku34221u3258/
- benatkin 15y agoIt would probably take them a lot of thought to fully de-identify the data and once they did it might be disappointing. A lot of so-called "de-identified" data out there actually contains things that could be used to identify people.
- robryan 15y agoI would assume that it would be just about impossible to de-identify YC applications, as they are all fairly unique and there is generally a lot of information online about each startup. So if an app had a story about a previous product and then that story shows up again in their blog or when they are doing publicity it makes it pretty easy.
- benatkin 15y agoThey could release it like this: pick a small number of fields like number of founders and age, and release each of them as separate arrays. Ex: {"founders": [3, 2, 2, ...], "ages": [31, 22, 27, ...]} (the size of the founders array and the ages array would be different because the first is per-startup and the second is per-founder) Just to get an idea of how vague de-identified data might be. And I'm not even sure that would be de-identified enough.