3 ms·
Dask and vaex are not 'competing', they are orthogonal. Vaex could use dask to do the computations, but when this part of vaex was built, dask didn't exist. I r
by maartenbreddels 8y ago
Dask and vaex are not 'competing', they are orthogonal. Vaex could use dask to do the computations, but when this part of vaex was built, dask didn't exist. I recently tried using dask, instead of vaex' internal computation model, but it gave a serious performance hit.
There is some overlap with dask.dataframe, I think they are closer to pandas than vaex is.
Vaex has a strong focus on large datasets, statistics on N-d grids and visualization as well. For instance calculating a 2d histogram for a billion row can be done in < 1 second, which can be used for visualization or exploration.
The expression system is really nice, it allows you to store the computations itself, calculate gradients, do Just-In-Time compilation, and will be the backbone for our automatic pipelines for machine learning.
So vaex feels like Pandas for the basics, but adds new ideas that are useful for really large datasets.
- themmes 8y agoHow could I've missed you being the author. Thanks for your extensive answer, will definitely try the library! And thanks again for Ipyvolume, has been very useful so far.
- maartenbreddels 8y agothanks!