9 ms·
This is a weird article at this point in time. The question it addresses: "Does Python's performance matter?" Has always had the answer: "Sometimes, and you
by passive 9y ago
This is a weird article at this point in time.
The question it addresses:
"Does Python's performance matter?"
Has always had the answer:
"Sometimes, and you have options for those cases."
The OP found a "sometimes", and he's using one of those options. In this case, he's got Python for prototyping and glue, with Haskell improving performance. This is as it should be.
I don't know of any Python advocates who say it's the right tool for every part of every job. What we will say is it's usually a good "first" tool for every job. Building a system out in Python allows you to get something representative fairly quickly, which helps identify if there are areas where Python alone is not enough.
- deleted 9y ago[deleted]
- scribu 9y agoI think the article makes a lot more sense if you consider it in the context of "Python for data science". In the last few years, there's been a lot of hype about replacing other number crunching solutions (R, SPSS, even Matlab) with the Python ecosystem of tools (Pandas, SciPy, etc.).
- autokad 9y agoi dont seem to follow. if you are doing data science, all the bottle necked stuff will be running in numpy or pyspark. Choosing python over R, SPSS, Matlab usually doesnt come down to which one is faster, and R as far as i know is at least not vastly superior in speed.
- make3 9y agoor a cuda wrapper
- michaelsbradley 9y agoAs with Python, the fast libraries written for R are usually implemented with something else under the hood. Take the data.table library, for example: https://github.com/Rdatatable/data.table https://github.com/Rdatatable/data.table It's wicked fast for many kinds of tasks, but its R API is just a thin layer on top of C.
- kazagistar 9y agoThis was explicitly addressed in the article: as soon as you have to do anything which isn't a trivial numpy operation, performance goes off a cliff, and that can be a problem.
- adamson 9y agoThis also isn't necessarily true. Take the example of TensorFlow. You build a representation of the computation you want to run, and then you can run nearly the whole thing end-to-end in native C++ using Eigen data structures, with occasional shuttling of data back into PyObjects (rare) or numpy (common) for metrics tracking. Cython is a much more powerful tool than I think the author of this article realizes.
- chestervonwinch 9y agoI don't understand what you mean about "hype about replacing ... with Python". How is it hype when the majority of people already use Python (see link)? https://www.kaggle.com/surveys/2017 https://www.kaggle.com/surveys/2017
- scribu 9y agoThe fact that a majority of people use Python for data crunching today doesn't prove that there wasn't hype in the past. I'm not trying to say it's wrong, just that it's become a very visible niche for the language.
- rhodysurf 9y agoExcept Pandas and SciPy use libraries written in CXX or Fortran and not pure python so speed is not really an issue with them usually
- halbritt 9y agopandas.read_csv is kind of abysmally slow, unfortunately. There are a couple of alternatives, but nothing has really taken hold. Dask exists, but not everyone can run a distributed system to read a multi-gigabyte csv.
- gameswithgo 9y agoI would argue that performance always matters, and that Python is never the right tool for the job in an absolute sense. Python may be the right tool for the job given the options we have today but there is no reason we cannot have a language exactly as nice to use as Python is, but that also provides good performance. Languages like Nim or F# approximate that ideal, for instance. And while I realize there are high perf variants of Python, these should be the primary standard, and only path. The slow path shouldn't exist. It is a failure of our community that we allow languages to proliferate while remaining slow. This is bad because allowing slow tools to become popular means people create slow things, which wastes other peoples time and energy. Electron becoming a standard way to make cross platform desktop apps is another example. Someone is not wrong to choose electron for that job, but we the software community are wrong to have let something that inefficient become the easiest way to do that job. You cannot simply dismiss this issue by saying "Well don't use the slow software you don't like then" as many of these things become de-facto standards that you cannot avoid. Your place of work may require Microsoft Teams as the chat software, and now you are using a huge % of your laptops ram and battery for simple text transmission. Atom becomes the popular target for language plugins and ends up the only usable way to get good IDE features for your language, and you suffer the performance hit for it. We can do better!
- lmm 9y ago> It is a failure of our community that we allow languages to proliferate while remaining slow. This is bad because allowing slow tools to become popular means people create slow things, which wastes other peoples time and energy. Make it work, then make it work right, then make it work fast. I mean yes, a lot of things are slower than they should be, but the level of outright correctness bugs in software today is mindblowing. So while replacing our tools with faster tools should be something we do in the long term, I'd put a higher priority on lowering defect rates and making it easier to produce working software.
- moocowtruck 9y agoi'd like to know how we make correct software
- lmm 9y ago> This is a weird article at this point in time. It's a timely article, because a number of things have changed in recent years to make the tradeoffs around using Python quite different from what they once were. Ten years ago, Python was slower than the alternatives by a small constant factor, datasets weren't big enough for python performance to be an issue, Python had a world-class tooling/library ecosystem and higher-performance languages at a similar level of conciseness/productivity were basically unknown. Today, as the article says, things are different: Core counts are rising so practical Python performance is falling further and further behind, datasets have gotten large enough for Python performance to be an issue, Javascript has proven that it's possible to get much higher performance out of a scripting language, languages like Haskell have gone mainstream and offer a comparable-to-Python (better, in fact, given what a mess Python's packaging situation is) tool/ecosystem experience and comparable levels of productivity with much higher performance. Every tool is a "sometimes", but good engineering is knowing when a given approach moves from being the right one 90% of the time to being the right one 10% of the time.
- laike9m 9y agoAgree, I really like the why "things are different" part.
- scaryclam 9y agoThis is a pattern I've had some success with several times now. Create a quick Python implementation for parts of data pipelines and then go back and re-write in Java/Go/C/whatever the best tool is for that bit of the job later when we know where the bottlenecks are.
- deleted 9y ago[deleted]