4 ms·
I think the article makes a lot more sense if you consider it in the context of "Python for data science". In the last few years, there's been a lot of hype ab
by scribu 9y ago
I think the article makes a lot more sense if you consider it in the context of "Python for data science".
In the last few years, there's been a lot of hype about replacing other number crunching solutions (R, SPSS, even Matlab) with the Python ecosystem of tools (Pandas, SciPy, etc.).
- autokad 9y agoi dont seem to follow. if you are doing data science, all the bottle necked stuff will be running in numpy or pyspark. Choosing python over R, SPSS, Matlab usually doesnt come down to which one is faster, and R as far as i know is at least not vastly superior in speed.
- make3 9y agoor a cuda wrapper
- michaelsbradley 9y agoAs with Python, the fast libraries written for R are usually implemented with something else under the hood. Take the data.table library, for example: https://github.com/Rdatatable/data.table https://github.com/Rdatatable/data.table It's wicked fast for many kinds of tasks, but its R API is just a thin layer on top of C.
- kazagistar 9y agoThis was explicitly addressed in the article: as soon as you have to do anything which isn't a trivial numpy operation, performance goes off a cliff, and that can be a problem.
- adamson 9y agoThis also isn't necessarily true. Take the example of TensorFlow. You build a representation of the computation you want to run, and then you can run nearly the whole thing end-to-end in native C++ using Eigen data structures, with occasional shuttling of data back into PyObjects (rare) or numpy (common) for metrics tracking. Cython is a much more powerful tool than I think the author of this article realizes.
- chestervonwinch 9y agoI don't understand what you mean about "hype about replacing ... with Python". How is it hype when the majority of people already use Python (see link)? https://www.kaggle.com/surveys/2017 https://www.kaggle.com/surveys/2017
- scribu 9y agoThe fact that a majority of people use Python for data crunching today doesn't prove that there wasn't hype in the past. I'm not trying to say it's wrong, just that it's become a very visible niche for the language.
- rhodysurf 9y agoExcept Pandas and SciPy use libraries written in CXX or Fortran and not pure python so speed is not really an issue with them usually
- halbritt 9y agopandas.read_csv is kind of abysmally slow, unfortunately. There are a couple of alternatives, but nothing has really taken hold. Dask exists, but not everyone can run a distributed system to read a multi-gigabyte csv.