14 ms·
Data-Oriented Programming in Python
- tomrod 4y agoThis is a wonderfully technical article. I'd love to learn more about Python internals as a scientific coder.
- barefeg 4y agoI recommend any of the talks by James Powell at PyData. For example this one https://youtu.be/cKPlPJyQrt4 https://youtu.be/cKPlPJyQrt4 Edit: maybe this one on Numpy may be more relevant: https://youtu.be/u2yvNw49AX4 https://youtu.be/u2yvNw49AX4
- pedrovhb 4y agoThe official Python documentation is excellent, and in many ways goes beyond providing just a list of existing modules and what they do. Sometimes if I'm bored I'll actually just pull up documentation for something I'm not 100% familiar with and have a look around, and I almost always find something new and useful. A couple of interesting ones are [0][1], and [2] is a nice starting point for discovering more. Not everyone's cup of tea, but I also found it enjoyable dive into asyncio with the docs. [0] https://docs.python.org/3/howto/descriptor.html https://docs.python.org/3/howto/descriptor.html [1] https://docs.python.org/3/library/collections.html https://docs.python.org/3/library/collections.html [2] https://docs.python.org/3/ https://docs.python.org/3/
- hgibbs 4y agoI'd like to plug riptables (https://github.com/rtosholdings/riptable https://github.com/rtosholdings/riptable), which is (more-or-less) a performance upgrade to pandas.
- anigbrowl 4y agoLooks nice!
- _visgean 4y agoHmm nice article but imho skips over the biggest optimization: numpy uses BLAS libraries so stuff like > >>> multiply_by_two = homogenous_array * 2 will be calculated most of the times using a BLAS library - whichever you are using (https://numpy.org/devdocs/user/building.html https://numpy.org/devdocs/user/building.html)
- cdavid 4y agoThat article talks about DL, where blas is much less relevant. The kernels are mostly CUDA (for GPU) and similar stuff for other accelerators.
- duped 4y agoI'm curious how you would do data oriented programming in a language with no type system and no control over memory layout. And I guess the answer is "you can't, but JITs might exist someday that do it for you" But you can't wave your hands around and say compiler optimizations will fix performance problems - they can, but they're not magic, and the arrow in the proverbial knee for optimization passes are language semantics that make them impossible to realize (forcing the authors to either abandon the passes, or rely on things like dynamic deoptimization which is not free).
- sirwhinesalot 4y agoBy using only coding patterns that are known to JIT well and lower level primitive types and containers if provided by the language. Maximizing the use of packages written in native code also helps. The resulting code is even more annoying to write than using a lower level language typed language in the first place, but ecosystem access sometimes makes up for it. Hopefully tools like mypyc get better, letting well-typed python code with reasonable usage patterns be compiled to reasonably efficient native code. Last time I used it I was pleased with the performance benefits but it couldn't even compile all files in a module to a single shared library, despite this being mentioned as possible (and recommended) in the docs. Maybe I was doing something wrong, but they don't answer their github issues often, alas. Any little thing helps though, it's one thing for throwaway scripts to be inefficient, but applications? At a large scale it is a monstrous waste of time and literal energy.
- jessermeyer 4y agoThose are basically contradiction of terms. Orienting the program structure around the data necessarily requires control over memory layout and how it is interpretted.
- gnuvince 4y agoUnfortunately, the terms "data-oriented design" and "data-oriented programming" refer to two different styles of programming. Data-oriented design is the approach to programming made popular by Mike Acton's CppCon keynote—as you say, it focuses on the layout of objects in memory to make the processing of data take advantage of the underlying hardware. Data-oriented programming is a style of programming that, as far as I know, originates in the Clojure community. It emphasizes using general data structures (vectors, dictionaries) to store all data and make code more re-usable. It has nothing to do with good cache utilization, pre-fetching, or avoiding branch mispredictions. It's a shame that two styles of programming which are almost diametrical opposites share such similar names. From the look of the article, it's discussing data-oriented design, but in Python, and I agree that it's kind of a weird match.
- wheelerof4te 4y agoTo spare you a couple minutes of your life, the article is saying this: Python + C modules = Speed Nothing new here, move along.
- brilee 4y agoIt actually isn't saying that. What do you think Python is made of, under the hood? It's C modules. The argument is not that NumPy is written in C, but that it amortizes the cost of Python overhead over multiple data, rather than incur it on each datum.
- college_physics 4y ago> In practice, scientific computing users rely on the NumPy family of libraries e.g. NumPy, SciPy, TensorFlow, PyTorch, CuPy, JAX, etc.. this is a somewhat confusing statement. most of these libraries actually don't rely on numpy. e.g. tensorflow ultimately wraps c++/eigen tensors [0] and numpy enters somewhere higher up in their python integration [0] https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/framework/tensor.h https://github.com/tensorflow/tensorflow/blob/master/tensorf...
- cauch 4y agoIt's a details, but I keep seeing it: > Yet, [the scientists] struggle to move away from Python, because of network effects, and because Python’s beginner-friendliness is appealing to scientists for whom programming is not a first language. I don't believe it's the whole story. In my case, during my 13 years in academia, I saw my field going away from C++ and towards python. Not because of network effects (it was the opposite: it was more difficult to not use what everybody was using), or because scientists were not able to program (the entry language of the whole field was C++, and python arrived only because scientists with a deep knowledge of C++ started to themselves switch the core library to be usable with python). I think something that computer scientists forgot when they consider the subject is that the way computer scientists do software is just not working when you do science. In science, you use coding as an exploratory tool. You are lucky if 10% of your code ends up being used in your final publication. Because the 90% was only there to understand and to progress towards the proper direction. For this reason, things like declaring variables, which is very important when one makes a professional software, are too costly to be useful when you need to write down a piece of code that you will ever run once to check a small hypothesis, especially when you have another language not requiring it. Another aspect is that you will present your scientific results to your colleagues, not your code (they are not interested in that), and they will come up with questions or good ideas, all very good for science, but rarely compatible with the way your algorithm was built in the first place, and you will need to shoe-horn it into your code (to test it) without taking 3 weeks. In this case, python flexibility and hackability is very useful. It's also visible in the popularity of things like Jupyter notebooks (I have to acknowledge it even if I personally don't like working with such tools), which reuse a working approach similar to what was done in mathematica and matlab, that were created with the scientific workflow in mind. I'm sure python simplicity has played a role. But I have the feeling that some people are totally oblivious on the fact that there may be other reasons.
- whatever1 4y agoComputer scientists also assume that you know what inputs your program needs and what is the range of the outputs. That is out of touch with scientific research. We may change overnight completely the inputs, the core logic and the outputs. Having to babysit function signatures, manage memory and types throughout these activities is just draining.
- TheDesolate0 4y ago
- hoppla 4y agoI wonder if the concept of pointer lookup latency also applies to other languages, such as Go. I assume so though…