2 ms·
How does Ray compare to spark? Is there a reason to use spark once libraries like Dask or Ray become more mature?
by kmax12 9y ago
How does Ray compare to spark? Is there a reason to use spark once libraries like Dask or Ray become more mature?
- artwr 9y agoOne of the advantages for libraries like Dask is that in the world of "many core" architecture, you incur less overhead than spark especially if you want to schedule large work on a single large machine in the cloud. This in turn enables you to transition a single workload from single machine to multi machine in a more seamless fashion. The fact that Dask also has high level collections which it knows how to parallelize is also interesting. For workloads which are more related to nd-arrays, matrices and scientific computing, my understanding is that is is more efficient than Spark. The integration with your ecosystem is also important. If you have to ingest from the (Java) big data ecosystem for instance, Spark has had a lot of work put in its integration with it, it just works for the most part.
- lmeyerov 9y agoRE:spark, we're curious about Ray mostly because of the potential for interactive-time (ms-level) compute for powering user-facing software. RE:dask, we care that Ray interops with the rest of our stack (Arrow). I haven't evaluated Ray-on-pandas, and the Ray was previously focused on powering traditional ML, so again, just first blush on the announce. I don't think anything is inherent, more about priorities and momentum. For example, Spark devs have been working on cutting latency, and Conda Inc is/was contributing to the Arrow world. I had assumed the pygdf project would get to accelerating arrow dataframe compute before others, so this announce was a pleasant surprise!