4 ms·
Paul from Bayes Impact here. I appreciate the sentiment, though in all respect it does seem like most of your concerns are addressed on the website, either on t
by pyduan 12y ago
Paul from Bayes Impact here. I appreciate the sentiment, though in all respect it does seem like most of your concerns are addressed on the website, either on the fellowship page or in the others.
> unless you provide a little more information about what these "hard" problems are
The second paragraph does go briefly over the problems we are currently working on (granted, not in much detail for the sake of brevity, but enough to give an idea of what type of challenges they are). There is a little bit more information on the front page, but granted since we started Bayes Impact two months ago we haven't been able to put as much work into the website content as we'd like to.
> Honestly it reads like your offering basic in training in a a random selection of tools
This is simply not the case -- while their level of experience varies, our current fellows actually comprise some well-established data scientists in their own right. It is precisely because the problems worth solving are tough to solve that we need to round up talented individuals who are able to commit to working on social impact projects full-time and pair them up with industry and domain experts who have the domain knowledge but may not have the time.
They each bring their own set of skills -- for example, someone who built Lyft's grid optimization system might be uniquely suited to help save lives by improving ambulance and fire truck dispatch and reducing average emergency response times.
> and then hoping some non profits present a problem with nice clean data that can be solved through application of a few methods from scikit.learn
This is precisely the point of Bayes Impact and why a longer engagement model such as fellowships is needed in the space (most current data science for social good organizations work on a volunteer basis model), so we have the time to build these longer relationships with nonprofits to leverage data science even in cases where data is messy or sensitive. We go a little bit more in-depth about it on our article here: http://blog.bayesimpact.org/blog/the-bayes-impact-mission/ http://blog.bayesimpact.org/blog/the-bayes-impact-mission/
> Worse 4-6 months might not even be enough time to formulate a problem that needs a solution
This is why they're not 4-6 months, but typically 6-12. We do have a pilot 3 month program in the summer for problems that are comparatively easier to work on.
> and then hoping some non profits present a problem with nice clean data that can be solved through application of a few methods from scikit.learn
This is why we have a fellowship application page and not a project application page -- we actually tend to identify and scope projects ourselves.
On that note though, I want to point out there is no need to be so overly dismissive of the work nonprofit and civic organizations have been doing in collecting and storing clean data. For example, most fire departments we talked to had surprisingly good data, and some such as the Fire Department of New York had even started initiatives of their own to use data science to improve their processes. For example, by integrating building permit data with their own systems, they've been able to direct inspectors where fire were predicted to be more likely to occur.
One direction we've been headed towards is seeking these data-educated organizations to create pilot projects, then use the results of these as a basis to export these solutions in similar institutions whose data practices may not be as good. In that end, we are helped by some data engineers from companies like Splunk or Cloudera so we do believe in working with these organizations in the long run to bring them up to speed. This is precisely the problem we're trying to solve with our model!
> For the record I work for a non profit analyzing complex diseases
Then you might be interested in the project we are doing on Parkinson's with the Michael J. Fox Foundation! Feel free to email me for more details.
- micro_cam 12y agoI'm trying to offer constructive, if harsh, criticism based on my own experience which includes recruiting for similar positions and working with large and small 501(c)(3)'s. I don't mean to come off as dismissive but to suggest that your write up is vague to the point of being easily dismissed and provide feedback on how someone from outside your local peer group might read this. And there are organizations out there with great IT and clean data but I and most people in this field have lost months writing hideous combinations of NLP and regular expression to pull data out of old medical records and things and hand validate it or correct for batch effect in supposedly clean data. I think that fleshing out the projects and areas of investigation you guys already have lined up would go a long ways towards addressing my concerns and making the program more appealing to the typical analytical folks i've worked with. I'd also suggest focusing the intensive course on analytical methods not the tools, this is what will intrigue people with expertise. At the moment it reads like it is focused at people new the the field with no programing experience. What data sets/types are you using for the Parkinson's thing? My main focus is on analysis methods that resist the noise, imbalance, heterogeneity and other issues typical in extremely wide/multivariate genetic+clinical+proteomic studies...a few sentences about the study in the write up would have told me a lot about if my skills could be useful. (I'm not looking to relocate but I am always open to collaborations and correspondence with people working on similar things.)
- pyduan 12y agoAs I said earlier -- I definitely appreciate the sentiment, and constructive criticism is always welcome when actually substantiated. I also took your post as an opportunity to elaborate a bit more on our model so my post got longer as a result. > And there are organizations out there with great IT and clean data but (...) This argument also works the other way round -- there are organizations out there with terrible data (and this is especially common with medical data), but there are also many high impact projects for which the data does exist in a workable form that are begging to be solved (and that we are actually working on solving). We are focusing on these in the short term, while laying the groundwork for the others in the medium-long term (both through the research arm we are building, and our data engineers). There is no reason not to get the low-hanging fruit first. > I think that fleshing out the projects and areas of investigation you guys already have lined up (...) Agreed. Since we created Bayes Impact two months ago our main focus has been on building the program from scratch and working on the projects as well, so the website has unfortunately taken a backseat. Another problem is that government organizations are very sensitive about communication and we can only communicate about our projects on their timeline. This results in us not having a website as fleshed out as we'd like, but this is par for the course for a new organization. > I'd also suggest focusing the intensive course on analytical methods not the tools Ah, I just saw the paragraph you're referring to. I get how the language may be a bit confusing and will make the appropriate changes -- our goal is actually to do the opposite: we bring on individuals who already have the analytical methods but some may not have had exposure to best industry practices. Because we focus on building production systems and not just write case studies, it's important to bring them up to speed in that minor respect. This is why we can spend only a week teaching tools -- teaching analytical methods to people without the required background would likely take much longer, which is not our target audience. At a broad level we simply provide an avenue for data scientists to work on social impact problems in collaboration with domain experts, with us taking care of the overhead of scoping projects and doing the dirty work of acquiring and preparing the data as well as defining the implementation strategy. We also smooth out the edges in our Fellows' backgrounds if any but this is really not the core of the program. Fortunately the pool of applicants as well as our current fellows does not seem to echo your fears but I'll review and see which changes to the fellowship page could help remove ambiguities in the future. Hope it helps clarify. Regarding the Parkinson's project, feel free to reach out to me by email -- unfortunately we need to wait for the press release from the MJFF and the other partner before I can actually communicate about the details publicly.