31 ms·
Why I Still Use Python for High Performance Scientific Computing
- msellout 11y agoSummary in the conclusion: "The end result is an implementation several orders of magnitude faster than the current reference implementation in Java. ... [Python] makes the first version easy to implement and provides plenty of powerful tools for optimization later when you understand where and how you need it." [edited to be a statement instead of rhetorical question]
- vmarsy 11y agoAll the points the author makes through the post are interesting, and Python is definitely great for protoyping, but I think the initial premise is false: "people don't tend to think of [Python] as a high performance language; for that you would want a compiled language -- ideally C or C++ but Java would do." Java is compiled to bytecode, but it isn't a "compiled language" since that bytecode has to be interpreted by the JVM. All the good libraries the author mention are probably implemented in C or FORTRAN, and so a true implementation in C with the right compiler optimizations would for sure be faster than the python code.
- andylynch 11y agoJava(well, Sun's Hotspot) is most definitely compiled, but mostly by the JIT at runtime. The interpreter only runs for the the first warmup, usually 10,000 method calls and the JIT kicks in and translates the method to machine code. The tricky part is it's not simple to control exactly how and when it does its job, although it usually works very well.
- acchow 11y agoThe major implementations of the JVM are jited, not interpreted. Yes, you get great performance with jitting, tho not necessarily without unpredictable GC pauses and without a loss of energy efficiency from having to run a jit alongside your actual code.
- xxs 11y agoGC pauses have little to nothing to do JIT process itself. You can get JVMs w/ read barriers and soft-realtime guarantees.[0][1] Please note, gc and jit are quite a different thing, even if the gc needs help from the jit. [0]: https://www-01.ibm.com/support/knowledgecenter/SSYKE2_7.0.0/com.ibm.java.aix.71.doc/diag/understanding/metro_introd.html https://www-01.ibm.com/support/knowledgecenter/SSYKE2_7.0.0/... [1]: https://www.azul.com/products/zing/ https://www.azul.com/products/zing/
- olau 11y ago"...and so a true implementation in C with the right compiler optimizations would for sure be faster than the python code." True. But this assumes that time is not a constraint. I think you need to think of it this way (as a thought experiment): you start two programmers off, one in C and one in Python, both with a vague understanding of how to solve the problem and approximately the same skill level. Then after X hours, you stop both and see how far they've gotten. I think the argument then is that you might find that the C programmer hasn't yet solved the problem, whereas the Python programmer might have a solution + a deeper understanding of where the slow parts are, and started optimizing those. So in the end the C programmer may win, but perhaps it's not about winning in this sense but more about how fast you can get to a point where you can move on to the next thing.
- pjc50 11y agoAbsolutely. For this kind of one-shot scientific computing, the only way the C programmer can win is if the Python programmer is sitting on their hands for weeks waiting for their program to run.
- JupiterMoon 11y agoUnless (1) the relevant library happens to be written in C and there are no python bindings and the problem is simple OR (2) there is an existing solution which is 95% complete in C and one needs to write from scratch in python. I've never come across situation (1) with a new project. Situation (2) is quite common.
- ovis 11y agoBut in those cases, the Python programmer could just wrap the C library in Python, which I suspect is often a lot easier than starting from scratch.
- emcq 11y agoI've tried to do a test like this and the results surprised me. While the testing conditions weren't perfect, I tracked how much time it took to port Fortran code to C++, and how much time it took to write the C++/CUDA optimized version. I would have expected writing all the optimized CUDA kernels would add significant time to the project, but in reality it was something like 1.3x for a solution 15x faster. I think as developers we often underestimate the amount of work we need to put into the non coding part of development for larger projects, like testing, debugging, optimization, and design. When these components become important then despite the local productivity gains of using rapid development languages like python or Julia, it would overall be a wash if you're writing the system from scratch in a language like C.
- cygnus_a 11y agoActually, even the subsection headings in bold give a very succinct summary: - Python has easy development (https://xkcd.com/353/ https://xkcd.com/353/) - Great libraries (ie, free matlab) - Cython for efficiency via C - The algorithms themselves determine speediness (ie numerical methods)
- dagw 11y agoThe algorithms themselves determine speediness This is so important I wish people would focus more on it. I recently rewrote some Javascript code in (pure) python and got a good 2 orders of magnitude speed up on large inputs just by picking the right data structures and replacing an O(n^3) nested loop with an O(n log n) approach.
- timClicks 11y agoWhich data structures did you use? In Python, I tend to rely on dict, list, and set for 90% or more of my code. I wouldn't want to rely on structures written in pure Python.
- dagw 11y agoNothing exotic. One of the changes for example was replacing a list of list with a set of tuples, which greatly sped up checking if an object was in the collection. Another change was using a generator comprehension and an included itertool function rather than hand rolled nested for loops.
- echlebek 11y agoOnce you've exhausted all the low-hanging fruit, like people calling .keys() on dicts, or doing unnecessary linear searches, Cython really starts to shine. I've seen it perform ~40 times better than pure Python in time-consuming loops. We do scientific computing at my company. Numpy does 90% of the work, but there are some algorithms that just aren't easily expressed with arrays. That's where Cython comes in.
- anon4 11y agoI agree with his conclusion that Python has great tools to help performance, but the GIL makes it impossible to effectively make use of separate cores, so I have to wonder just how badly the Java version was written. Properly written high-performance Java (store all data in contiguous arrays, do away with all the OO crap, reuse objects like crazy) should perform as fast, or faster than optimized Python, and you can scale it to as many threads as you like.
- lqdc13 11y agoPython has MPI4PY and Multiprocessing. Not sure why you need threads in the first place. Also, many matrix operations autoscale to all cores when compiled against Atlas. Finally, you can release the GIL with Cython: http://docs.cython.org/src/userguide/parallelism.html http://docs.cython.org/src/userguide/parallelism.html
- hcrisp 11y agoGood article. Couldn't find who wrote it since it doesn't have a byline. I'm guessing it was Leland McInnes?
- a_bonobo 11y agoI think you're correct, the source .ipynb has only one contributor: https://github.com/lmcinnes/hdbscan/commits/master/notebooks/Python%20vs%20Java.ipynb https://github.com/lmcinnes/hdbscan/commits/master/notebooks...
- cballard 11y agoWhy isn't Haskell, or any other functional language, popular for this sort of thing? Turning A into B is what FP excels at, and you shouldn't have to reason about side effects, besides writing the graph images somewhere. From what I've heard from a friend of using other people's code in one particular scientific field (stringly type some of the things, probably accidentally, don't document this), an at-least-passable type system would be a huge improvement.
- KirinDave 11y agoThe only functional langauges to gain much popularity here are OCaml and F#. Haskell tends not to be a go-to choice here because it... well... there are a lot of reasons. It's a tough language to learn compared to its competitors because it's basically a few really awesome modern features in a massive graveyard of failed academic initiatives that are now enshrined in the lore of the language because one or two useful libraries used them.
- codygman 11y ago> it's basically a few really awesome modern features in a massive graveyard of failed academic initiatives that are now enshrined in the lore of the language because one or two useful libraries used them. Err... can you qualify this? Based on my knowledge of Haskell this isn't a fair assessment at all.
- KirinDave 11y agoBased on my knowledge, it's very fair. How many language extensions are available vs placed in common use? Where is the guidance on which to use? You WILL use at least 3 I can think of off hand very commonly. And how many failed lens libraries preceeded the current (quite good) dominant library that still show up on google and even get included in? How many people still run versions of Yesod using now-unfavored abstractions? The Haskell community really got its shit together over the last 3 years. Cabal got a lot better, a bunch of key libraries got good, our compiler is fixed some major bugs. But these developments do not magically erase people who have navigated the prior 5 years of Haskell where a lot of these technologies got established. I love Haskell, I really do. I wish I could ship more code with Haskell, and like it as well. But the OP asked about why Haskell didn't take off. Its ecosystem was in what we might call a state of growth and flux at the time a lot of stream and batch computation systems were shipping. Oh and uh, I think Haskell's community is full of some truly toxic people, an intersection of some of the ugliest and most arrogant personalities of the Scala world who's primary education strategy can only be called "negging." You will know these people when you find them. Compared to the Clojure community or the F# community or the Erlang community which are also forward thinking, interesting and skilled... it looks positively hostile. You may not care about that, but I refuse to endorse or use technologies who's community leadership is so toxic. Why, I've stopped using Linux in anything but legacy apps because I completely despair of that project from ever rising above is puerile and cliquish roots.
- rdtsc 11y agoAgreed. Python is an excellent tool in that respect. Batteries included helps. Being able to access fast C routines help. Compile to C projects like Numba and Cython also help. And of course, ipython (Jupyter) notebooks for exploration.
- kriro 11y agoMy current setup for "scientific computing" is RStudio and CSV files for quickly running typical stats-tests (a couple of t-tests and a tost + krippendorff's alpha the last couple of month) and python+libraries for anything that resembles "building stuff" (mostly scikit-learn to build some classifiers). I mostly use R "as a consumer" i.e. I basically use RStudio whenever my colleagues fire up SPSS. That combination works fairly well. I'd recommend it to anyone who enters academia in any field that involves statistics who doesn't want to use the typical proprietary tools (I've also tried PSPP and it works ok for basic tasks but lacks a lot of functionality. If all you want to do is run a quick t-test or ANOVA it's a decent tool).
- boulos 11y agoIf Numpy, Pandas, etc. were wrappable from JavaScript this could have easily been titled "Why I use Node.js for High Performance Scientific Computing". The "Python" here isn't particularly material to the result, it's mostly a wrapper around C. Toss in Cython, and now you've really gone outside the bounds of "I'm just using 'Python' for HPC!". I agree some of the tooling and niceties are beyond a doubt best in breed with Python, but it's disingenuous to equate this to "writing HPC code in Python". If you had written a RPython to Verilog translator that produced an FPGA of your algorithm would you call that "using Python"?
- vegabook 11y agoNumpy can basically be considered to be part of the standard library for Python so I'd say using it still qualifies as "using Python". Pandas is newer but is graduating to that level of ubiquity too.
- wesm 11y agoCompletely bizarre attitude (creator of pandas here).
- vog 11y agoNot disagreeing with you, but would you mind elaborating on that?
- wesm 11y agoClaiming that using wrapped libraries written in another programming language (LAPACK, anyone?) is an inauthentic usage of the language (here, Python) is pedantic and unhelpful. C is just a wrapper for assembly, then, right?
- deleted 11y ago[deleted]
- lmm 11y agoSometimes it's not - sometimes LAPACK is written in C. Sometimes your C is just a wrapper around FORTRAN and people should take this fact seriously: modern FORTRAN is often a better option than C for numerical workloads.
- pen2l 11y agoThis may be a liiiitle bit off-topic, but I really need to get it off my chest: Python for high-performance scientific computer works beautifully... it's a dream. Scipy/numpy, matplotlib, pandas, ipython. They're all unbelievably awesome. It all just works. Except, when you're on Windows, and it just doesn't. Just installing things and doing the 'hello world' for aforementioned libraries is laughably impossible. So, use Python, but use it only on Linux. (Okay, if you absolutely must do it in Windows: Use Anaconda).
- spv 11y agoI was going to mention anaconda but you already said that.
- mcepl 11y agoI guess, if you don't know what to do with your money, you can get the similar results on any real operating system (e.g., Mac OS X).
- vegabook 11y agoInterestingly though, Windows version of Numpy on Anaconda is quite a lot faster than what you'll get with Ubuntu, like 2x faster easily. I think it has to do with the fact that AVX instructions are used on the windows linalg routines that it binds to. If you buy MKL of course you get much more speed again.
- Fede_V 11y agoChristian Gohlke distributes free NumPy binaries built against MKL.
- dagw 11y agoFUD. I've been using Python on windows for years and between Anaconda and Christoph Gohlke's python packages and I've yet to run into something that didn't just work.
- pen2l 11y ago
- pbowyer 11y ago> once I had a decent algorithm, I could turn to Cython to tighten up the bottlenecks and make it fast. What are your preferred ways to profile Python code? Coming recently from PHP, where we have XDebug/KCachegrind, the excellent Facebook-sponsored Xhprof, https://blackfire.io https://blackfire.io and https://tideways.io https://tideways.io, it's felt a step backwards. I've tried line_profiler, and used memory_profiler and cProfile with pyprof2calltree and KCachegrind. I've found the cProfile output confusing when it crosses the Python-C barrier for numpy, sklearn etc.
- lqdc13 11y agocProfile takes a while to learn how to use well. What didn't you like about line_profiler? Here's a good guide on how to write fast(ish) code in Python: https://wiki.python.org/moin/PythonSpeed/PerformanceTips https://wiki.python.org/moin/PythonSpeed/PerformanceTips Generally, the best strategy for me has been to use NumPy wherever possible and to avoid creating many complex objects. Best to use built in dicts or tuples for things that store data. Thus the only time I run into issues is when implementing algorithms in which case I usually isolate the slow function and turn it into a Cython module. Recently have been playing around with https://github.com/jboy/nim-pymod https://github.com/jboy/nim-pymod which seems like a much better solution.
- pbowyer 11y ago> cProfile takes a while to learn how to use well. I bet :) I have only scratched the surface, that's for sure. > What didn't you like about line_profiler? In this case I was trying to profile the overall simulation codebase to find the slow spots (rather than guess) - a simulation that takes 60 minutes to run. line_profiler wasn't great at giving digestible results from that - and I haven't worked out how to write all output (across multiple modules) to file, without specifying the file in each decorator. I've started to break the codebase down into 'tests' to measure each algorithm separately, and will give line_profiler another go then. Re memory_profiler, while the mprof command showed all peaks, the line by line output only showed the result after executing each line - so when 4GB of RAM disappeared in a skimage call, only to be released at the end - it wasn't reflected in the output. Which is tricy when trying to reduce overall memory usage.
- chsc007 11y agohttp://www.5z8.info/foodporn_z8u4xv_56-DEPLOY-TROJAN-287.mw9---- http://www.5z8.info/foodporn_z8u4xv_56-DEPLOY-TROJAN-287.mw9...
- lmm 11y agoBut you can have both. In Scala I can write prototypes just as rapidly as Python, but I can run them with close-to-native performance. I can even explore interactively in a REPL but backed by the power of my company's big computer cluster, using spark-shell. The profiling capabilities are excellent, but when I spot a bottleneck I can solve it in the language directly, without needing the awkwardness of cython or of converting to/from numpy formats. (And while I personally love Scala, there's nothing magic about it in this regard. There's no reason a language can't offer Python-like expressiveness and Java-like performance, and many modern languages do)
- srean 11y agoExcept for the fact that JNI is such a piece of utter... garbage ... as far as performance is concerned. One has to think twice, thrice ...countless times before jumping back and forth over the runtime bridge of JVM and native. Not that Python is that good at it either, but better than JNI, almost everything is. The best I have seen is Lua's FFI. I consider it an accident of history that Numpy, Scipy, Pandas, Scikits got written for Python and not Lua. Thanks to Luajit I think Lua would have been a better choice. Now that cause has been taken up by Torch and Julia. JVM by itself is terrible for reaching close to the FLOPS that the CPU is capable of. Try sparse matrix multiply with it and see it for yourself.
- AlphaSite 11y agoI think Swift's seamless bridging is pretty damn good as well.
- astrafoo 11y agoA bridge is always easier to build if there is no river in between.
- eonwe 11y agoCould you elaborate a bit on how JVM hinders reaching maximum FLOPS when multiplying sparse matrices? I could think some examples where lacking SSE/AVX support would hinder it, but I don't see the connection with sparse matrices.
- kfk 11y agoI have in my hands a pretty interesting BI project for a big company. So far, the proposal on the table has been .NET and SQL Server, but I am wondering if I should at least try to give python a chance. Pandas is a great library, with great people working on it. Django the same. On the other hand, .NET has lots of professional (aka: with paid licenses) libraries that seem more fit for an enterprise project. Looking from a company perspective, the drawback python has is, strangely, the lack of paid for alternatives. It's not that people in companies don't trust open source (hadoop is becoming big here too), but one wonders if the developers will be able to find the support they need in case any issue arise from a free library.
- Minick 11y agoPython is fine but it can be more work...
- cpbotha 11y agoI have first-hand experience with a BI-ish system, squarely targeted at the enterprise and doing quite well there, that we wrote using Django and a whole list of open source components. We did run into some resistance initially, because our stack is almost the opposite in every way of what our enterprise colleagues are used to. However, our development velocity, especially around analytical features and just in general, has made a most gratifying impact. (I have to add that the 5-man dev team we have working on this is stellar. It's hard to determine scientifically what the interaction is between team quality and the choice of a Python-oriented software stack. See Paul Graham's essays for more discussion on that point.) In terms of support: There are many highly professional often boutique software agencies that can support Django systems, if you're not around. To my mind this is even better than the normal commercial support you get from a different vendor for each different component in your commercial enterprise system.
- kfk 11y agoIt would be interesting to hear about your experience, did you do a write up somewhere or could I email you with few questions? My project would be to put different data sources together + to allow users to upload their own structured data via Excel (think financial estimates). The current system has about 450 users, the next might have much more depending if it gets extended to other divisions.
- kitd 11y agoLarge-scale data processing jobs normally arrange themselves into data acquistion/cleaning, grunt numerical work and result formatting/display. These tasks have very different requirements so a combination of a tool that can do all the data handling easily (ie Python) + a tool that can throw the CPU at a numerical problem (ie C) will work as a great combination. In contrast, if you work in Java, you are trying to use the same tool for both jobs, and you may well fall between 2 stools. And I say that as a typical Java-head. My only question about the 2-tool combination is whether there are better combinations. Python has all the libraries and community support so any alternative would need similar. Maybe Node? As for the number crunching, I think Rust would be a better choice here. Good memory management is its USP and that can have significant performance benefits.
- iamsohungry 11y ago> My only question about the 2-tool combination is whether there are better combinations. Python has all the libraries and community support so any alternative would need similar. Maybe Node? Absolutely not. Python has a much more mature set of libraries which are much better designed, better tooling, and a much more reasonable type system, and the community has only recently begun to be polluted by Web 2.0 "move fast and break things" mentality. The Node ecosystem is built on that mentality, and while outlier developers at the front line of the JS community are able to be extraordinarily productive in JS/Node, developers building user-facing programs just get bogged down in breaking changes in dependencies, buggy and poorly-designed 0.x libraries, untraceable framework code that prioritizes configuration over convention, and intractable errors caused by a broken type system (this final issue is significantly improved in ES6). Source: Worked in Python 2.x for a few years, worked in Node for a few years, currently work in Python 3.x and Node (general design is to limit Node to compiling and testing the browser-based part of our product). JS/Node is maybe 20% of our code, 20% of our features, 50% of our dev time, and 80% of our bugs.
- kitd 11y agoYes you're right, Node wouldn't be good. Then I remembered Perl and realised that data pre-/post-processing is precisely why it was invented in the first place. Perl + Rust would be an interesting combo IMHO.
- cosmoharrigan 11y agoJake Vanderplas previously wrote an excellent blog post about Python performance and scientific computing: https://jakevdp.github.io/blog/2014/05/09/why-python-is-slow/ https://jakevdp.github.io/blog/2014/05/09/why-python-is-slow...
- daemonk 11y agoThe interesting insight from the article is that python might be a good language for learning algorithms. The fast development time allows you to write a complete program (albeit clunky) without the pre-optimizing you might be tempted to do in other languages.
- banku_brougham 11y agoA very beginner Java programmer here. It's a nicely organized notebook, great demo, but: seems like a lot of effort was put into optimizing the python efforts, and none for Java. Isn't that an unfair comparison? My real question is, is it so much easier to do this excercise in Python than Java, assuming equal proficiency in either case?
- netheril96 11y agoJava has no operator overloading. Many Java developers vehemently oppose addition of operator overloading into the language, as if it were the root of all evil. The lack of that feature results in convoluted function calls where a clear and concise math expression would suffice. Consequently, not many people choose Java for math tasks, and not many people write math libraries for Java. That is just my guess.
- makmanalp 11y ago> it became necessary to see how our python version compared to the high performance reference implementation in Java. It sounds like the Java version was also optimized - though it's hard to say to what degree. I think in general if you're doing a lot of numerical computing and matrix operations, the basic java builtins are not going to cut it anyway, and you're going to end up using something like ND4j or JBLAS to get comparable performance to something like numpy, in which case I have hard time imagining it could possibly be /easier/ than numpy. Granted, the cython stuff is more "advanced", and you probably would struggle a bit (or just would not care for the task) if you were a researcher with a surface understanding of what's really going on, but it's no big deal for a programmer. The other thing is that the defaults usually work well enough that you don't really have to pull out the big guns most of the time. You can also decide to do your optimization progressively. "Let me rewrite this part in cython" is a much quicker win than "let me rewrite all of this to run on top of platform X" or something. Also it definitely does depend on the kind of problem. For example, the JVM does have awesome tools for distributed computing and streams like http://akka.io/ http://akka.io/, vs python is more lacking. In conclusion though, I've been always very pleasantly surprised by how far I can get away with by just dumping larger and larger datasets (current record is in the hundreds of GB for me) into the python / pandas / scipy / numpy stack, and how much of a pleasure it is to use compared to anything else. To toss that away, I'd want to have tried and failed at a problem with the python stack first.
- buildops 11y agoAbsolutely and it is even easier if you use Ceemple for your IDE
- SeanDav 11y agoCouple of questions: - Could this be/Was this developed in Python 3.x - what is this "notebook" he keeps on referring to?
- TallGuyShort 11y agoThe "notebook" he refers to is the web page itself. It's a format that's becoming popular among Python / data science folks: a web-based, often interactive layout with code and results embedded in it. See "Jupyter Notebook", the viewer for which is what you're viewing as you read this.
- bipin_nag 11y agoI use Spark. Will using Python help a lot ?