3 ms·
Any serious attempt at modelling this over python would use the pydata stack (numpy, pandas, etc), which run on top of C++ anyways.
by silveraxe93 6y ago
Any serious attempt at modelling this over python would use the pydata stack (numpy, pandas, etc), which run on top of C++ anyways.
- disgruntledphd2 6y agoYeah of course, apologies if that wasn't clear. The best solution here would probably be to package up the core routines into a library and use this from either R or Python.
- silveraxe93 6y agoSorry if I came out a bit snippy out there. But yeah I assumed you meant python without numpy, etc. A lot of the criticism I saw was because the core routines did not need to be packaged up. There were a lot of common data structures reimplemented, etc. I don't think the model had many novel routines. It could be built just using industry standard and tested tools in python, R, Julia (if you really want speed) etc. But it reinvented the whole ecosystem in one big ball of C. tbh, this should have been built on STAN or similar. There's so many variables and assumptions that the output is completely dominated by the parameters chosen. Seeing the distribution of outcomes instead of a point estimate would be actually useful.
- disgruntledphd2 6y agoTrue, I think that's worth noting. To be fair though, this codebase is pretty old, and it's unlikely that the technology landscape looked much like today (especially in terms of R and Python), so I can see how they ended up here. I'd love if this was written in R, Python or Stan so I could contribute, but that's probably not the researchers focus ;) While Stan is amazing, I shudder to think as to how long this model would have taken to run using MCMC (1 week plus maybe?).