4 ms·
Excellent to see functional programming ideas like deferred computation and clean interfaces make their way into the scientific computing space. One thing thou
by kortex 5y ago
Excellent to see functional programming ideas like deferred computation and clean interfaces make their way into the scientific computing space.
One thing though: I know these are good ideas. But to someone not as familiar with these patterns, they may wonder "why go through all this trouble?"
- Frost1x 5y ago>One thing though: I know these are good ideas. But to someone not as familiar with these patterns, they may wonder "why go through all this trouble?" It's not only why, it ignores the fact that the vast majority of scientific code is written for the science. Usually you're already dealing with layers of abstract theory in the science you're working in, you often don't want to deal with additional cognitive load of then abstracting that for a computer because that's now a new problem. If you take this approach in scientific software, unless you're a commercial vendor implementing some battle tested theory with a paying market, you're going to quickly find yourself without a job--that or you're a wizard at efficiently abstracting abstract theory quickly. I've yet to see someone more efficiently abstract such an implementation compare to simply implement a working, good enough, concrete implementation. You only abstract and optimize what really has to be or will clearly result in a net savings of resources for your specific project context. The vast majority of scientific code isn't written to be robust by design its a known cost cutting measure to meet budgetary constraints, it's written for a few target goals, not generalization. I've worked in scientific computing and applied science for quite awhile and the ideas like this just don't make sense in most contexts. From a software design perspective, sure, it makes complete sense but it's just unnecessary additional complexity and overhead in most contexts. People with software backgrounds often walk into scientific computing with naivity but often good intent that practices should mirror enterprise software approaches and it's just flat out misguided. A lot of theory is iterative with a short half-life so to speak, so any code you write surrounding it as a grounding basis will often become outdated quickly. You're often not implementing a method or function like the toy problem here of computing expected value which is a common shared statistical idea that will long outlive the majority of the research you're involved in. You'll instead be writing something highly specific to some toy new theory/model that may later be found to be false or iterated on to find an all together better abstraction you need to again abstract in your software. You're going to be using these sort of clean robust numerical abstractions, like expected value, when you need them though because they already exist and you can simply call them. It's the unique aspect tightly tied to the research (which is often inherently unique in terms of software it needs) you're doing that will be slapped together rapidly. The vast majority of it will be tossed away. If you hit something successful, that's when it's time to start thinking about these ideas because now you should consider refactoring your nice generalizable sharable theory others find value in and can use to such a nice clean implementation, then throw your existing prototype into the fires of hell where it belongs. There's usually not a lot of funding for that though and that's where commercial industries come in to scoop up publications and create nice clean efficient implementations of said theory they roll into some computational package to sell to people.
- pid-1 5y agoI don't think that's exclusive to academia. Business also need to choose wisely when to make ugly POCs for prototyping and when to create robust products / libs to save money long term. It's not uncommon for labs to have frameworks and internal libs to aid prototyping and experimenting.
- sharikous 5y ago> If you hit something successful, that's when it's time to start thinking about these ideas I agree with you completely. But the article did not specify the use case of these guidelines. They are not to be applied (in my opinion) for research code when you quickly need to publish something. They can be useful however when your already proven and battle tested ideas are used by other people. For example for keeping a shared code base inside a lab, or when you want to provide a robust implementation on top of your ideas.
- mattkrause 5y agoI thought the deferred part could be better. For example, one easy win would be to replace the list comprehensions with a generator: there's no point in allocating that entire list just so that statistics.mean can iterate over it. At the opposite end, the switch to a Distribution class also enables a huge speedup: keep the sampling machinery for situations where you actually need (Cacuhy?), but while die.expected_value() returns (self.n_sides+1)/2, which is effectively free.