3 ms·
I feel like you considering Python as having few real pain points in data science as either lack of knowledge of other languages or imagination/ambition. I do w
by ddragon 6y ago
I feel like you considering Python as having few real pain points in data science as either lack of knowledge of other languages or imagination/ambition. I do work on Python for data science/engineering in a production environment and I do find many of those all the time. Python does not feel like pouring my thoughts because I cannot write Python directly without taking hours to handle a few million points of data, so I have to juggle with multiple dialects (pandas, numpy, pytorch, tensorflow), and Python ends up being one of the languages that I need some documentation at all times. And then that documentation is also not obvious what's the input, do I need to give a tuple, or perhaps a dict or even a list because typing is very recent so the ecosystem isn't up to date and sometimes I can't even type my code properly because I legitimately don't know what a library function returns. And even with mypy I keep finding errors that I can only find after deploying (I don't blame Python on this one, and Julia isn't really better, but I can still dream of something that is dynamic when I want but still safe). Also Python for a dynamic language has pretty mediocre interactive story (using ptpython since the default repl is unusable), even simple things like copy pasting to a REPL can end up being a pain because of indenting, and the repl is far away from Lisp, Clojure and even Julia and Elixir. And Elixir also makes me especially disappointed with Python's multithreading story (in Elixir it's so natural that it really is the "pouring your throughts on the screen" for distributed).
It ended up being a rant, but I have many pain points with Julia as well, but it's still a new language that has more space to evolve and find ways to solve them, and if the solution is yet another language that solves all of them, I'll quickly jump. I spend a lot of time programming, so any significant improvements on the usage of my time is worth the effort in learning.
- bonoboTP 6y ago> so I have to juggle with multiple dialects (pandas, numpy, pytorch, tensorflow) NumPy should be enough for general computation. If you need autodiff or GPU then add in PyTorch. Pandas is more about various metadata than the actual numerical computing. If you want those types of features, the complexity doesn't disappear if you go to a different language. There's an effect where a new generation of developers see complexity built by the earlier generation, say it's too complicated and mess up, we don't need all that, so start over clean and it all looks so easy. But it's deceptive, because it will get complicated again once you put in all the features but it will look familiar now to this generation of developers as they grow side-by-side with the new language/framework. After a few years the cycle repeats and a new generation says "what's all this mess, why do I need to juggle all this, I just need XY." > Also Python for a dynamic language has pretty mediocre interactive story (using ptpython since the default repl is unusable), even simple things like copy pasting to a REPL can end up being a pain because of indenting, and the repl is far away from Lisp, Clojure and even Julia and Elixir. Use Jupyter Notebooks (or IPython if you don't want to leave the shell). > disappointed with Python's multithreading story Thread pools (executors, futures etc.) and process pools (multiprocessing module) work quite nicely.
- ddragon 6y agoI'll have to disagree with your defeatist generalization here: creating something with better knowledge of the problem and a lot of hindsight will not necessarily get as complicated for the same feature set. In fact the resets are probably necessary so you can get to even higher feature sets before getting so obtuse that it's a pain to move another step. Julia doesn't even need full parity with numpy because you can trivially write your needs in straightforward Julia (in fact Julia does not have numpy, only Julia arrays). And Pytorch equivalent? Also uses Julia arrays and straightforward methods (as you can just differentiate Julia code directly). Pandas equivalent? It's a wrapper over Julia arrays. Need named tensors? Just pick another wrapper over Julia arrays and it will work on Julia's Pytorch equivalent without changing anything. I do have to deal with Jupyter Notebooks, but they do have a lot of "real pain points" as well (it's always a pain when a data scientist gives one to add for deployment, half the time it does not work because it does not keep track of the whole execution history and need a major rewrite), but the main issue is the disconnect between the exploratory code and the production code. Julia's workflow (Revise + REPL) means I still write the structured code as usual, but every change in the code reflects in the REPL, like I'm interacting with the production code directly (and the Julia language supports a lot of commands to interface with the compiler, from finding source code to dumping the full structure of any object and even all types inferred). And of course, Julia also has Jupyter in the first place as it's name suggests, and a very cool alternative for some workflows like Pluto.jl that improves on reproducibility). You might also want to try Erlang/Elixir, it's really amazing how you can trivially spawn hundreds of thousands of tasks that are incredibly robust even though it's a dynamic language (since you add not only tasks to do something but also to monitor and handle errors). And you can even connect and hot swap parts of the application live (as an example of superior interactivity). I'm not saying that Python is not good, it wouldn't get where it is if it weren't. I'm saying that things can be way better, and we already have examples of it in most particular areas, but none that covers all of Python's strengths, which is why it will still be king for the foreseeable future. But I can only hope that we will get better and better tools to tackle more and more complex problems in the most simple and optimal way that we can given all that we learned as a community.