6 ms·
I must be missing something. Modern data science workloads involve fanning out data and code across dozens to hundreds of nodes. The bottlenecks, in order, are
by kuzehanka 7y ago
I must be missing something. Modern data science workloads involve fanning out data and code across dozens to hundreds of nodes.
The bottlenecks, in order, are: inter-node comms, gpu/compute, on-disk shuffling, serialisation, pipeline starvation, and finally the runtime.
Why worry about optimising the very top of the perf pyramid which will make the least difference? Why worry if you spent 1ms pushing data to numpy when that data just spent 2500ms on the wire? And why are you even pushing from python runtime to numpy instead of using arrow?
- corndoge 7y agoGood lord, hopefully latency isn't 2.5 seconds!
- sjwright 7y agoI can’t even. How could you ever get 2500 msec on transit? That’s like circling the globe ten times.
- justinclift 7y agoMaybe a bunch of SSL cert exchanges through some very low bandwidth connections? ;) Still, it's more likely a figure used for exaggeration, for effect.
- deleted 7y ago[deleted]
- kahnjw 7y agoLatency in training literally does not matter. You care about throughput. In serving, where latency matters, most DL frameworks allow you to serve the model from a highly optimized C++ binary, no python needed. The poster you are replying to is 100% correct.
- dnautics 7y agoquote is 'data spend 2500 ms on the wire'. That's not latency. For a nice 10GbE connection, that's optimistically 3 GB or so worth of data. Do you have 3GB of training data? Then it will spend 2500 ms on the wire to distribute to all of your nodes as part of startup.
- gbrown 7y agoNot everyone operates at that scale, and not every data science workload is DNN based I agree with your general point, however, but the role I'd hope for with Rust is not optimizing the top level, but replacing the mountains of C++ with something safer and equally performant.
- MiroF 7y agoBut the title of this post is Python vs Rust, not C++ vs Rust. Maybe BLAS could be made safer but i don't think that's what's happening here
- danielscrubs 7y agoEvery small thing counts when you have big data which is exactly why you need performance everywhere, if Rust can help with that I don’t mind switching my team to that. The problem are usually when you do novel feature engineering not the actual model training. But I was a C++ dev before checking the assembly for performance optimization so I guess I have more wiggle room to see when things are not up to snuff. If I got a cent for every: -you are not better than the compiler writers you can’t improve this. Especially from the Java folks. They simply don’t want to learn shit, which is fine if they just where not so quick with the lies/excuses when proven wrong.
- kuzehanka 7y agoNovel feature engineering? Like this? https://towardsdatascience.com/python-performance-and-gpus-1be860ffd58d https://towardsdatascience.com/python-performance-and-gpus-1...
- danielscrubs 7y agoLooks good. I’ve tried Numba and that was extremely limited. Current project we can’t use GPUs for production so we can only use it for development. Not my call, but operations. They have a Kubernetes cluster and a take it or leave it attitude. We did end up using C++ for somethings and Python for most. I’d feel comfortable with C++ or Rust alone if there was a great ecosystem for DS though.
- ldng 7y agoI see a graph on ... logarithmic scale ? No unit ? I don't know what that benchmark means.
- kahnjw 7y agoThis is just not true. The python runtime is not the bottleneck. DL frameworks are DSLs written on top of piles of highly optimized C++ code that is executed as independently from the python runtime as possible. Optimizing the python or swapping it out for some other language is not going to buy you anything except a ton of work. We can argue about using rust to implement the lower level ops instead of c++. That might be sensible though not from a perspective of performance. In a "serving environment" where latency actually matters there are already a plethora of solutions for running models directly from a C++ binary, no python needed. This is a solved problem and people trying to re-invent the wheel with "optimized" implementations are going to be disappointed when they realize their solution doesn't improve anything.
- likeabbas 7y agoA big push for NN is to get them running in real time on local GPUs so we can make AI cars and other AI tech a reality. 2500ms could be life and death in many scenarios