4 ms·
First of all, great to see more powertools to choose from for my ds workflow! However, I am suprised to see no mention of Dask in the article. How do these lib
by themmes 8y ago
First of all, great to see more powertools to choose from for my ds workflow!
However, I am suprised to see no mention of Dask in the article. How do these libraries compare?
- maartenbreddels 8y agoDask and vaex are not 'competing', they are orthogonal. Vaex could use dask to do the computations, but when this part of vaex was built, dask didn't exist. I recently tried using dask, instead of vaex' internal computation model, but it gave a serious performance hit. There is some overlap with dask.dataframe, I think they are closer to pandas than vaex is. Vaex has a strong focus on large datasets, statistics on N-d grids and visualization as well. For instance calculating a 2d histogram for a billion row can be done in < 1 second, which can be used for visualization or exploration. The expression system is really nice, it allows you to store the computations itself, calculate gradients, do Just-In-Time compilation, and will be the backbone for our automatic pipelines for machine learning. So vaex feels like Pandas for the basics, but adds new ideas that are useful for really large datasets.
- themmes 8y agoHow could I've missed you being the author. Thanks for your extensive answer, will definitely try the library! And thanks again for Ipyvolume, has been very useful so far.
- maartenbreddels 8y agothanks!