5 ms·
One of the advantages for libraries like Dask is that in the world of "many core" architecture, you incur less overhead than spark especially if you want to sch
by artwr 9y ago
One of the advantages for libraries like Dask is that in the world of "many core" architecture, you incur less overhead than spark especially if you want to schedule large work on a single large machine in the cloud. This in turn enables you to transition a single workload from single machine to multi machine in a more seamless fashion.
The fact that Dask also has high level collections which it knows how to parallelize is also interesting. For workloads which are more related to nd-arrays, matrices and scientific computing, my understanding is that is is more efficient than Spark.
The integration with your ecosystem is also important. If you have to ingest from the (Java) big data ecosystem for instance, Spark has had a lot of work put in its integration with it, it just works for the most part.