3 ms·
Any plans for DataFrames to support larger-than-memory data though? JuliaDB is super useful for this I've found, basically a solid Julia alternative to Dask.
by electriccello 6y ago
Any plans for DataFrames to support larger-than-memory data though? JuliaDB is super useful for this I've found, basically a solid Julia alternative to Dask.
- StefanKarpinski 6y agoMy current thinking is that it would make sense to implement the standard DataFrame abstraction (with all the generic functionality it offers) with a couple kinds of distributed data frame. One approach is to make a distributed stack of local data frames; another is to make a single data frame where each column is a distributed vector. Not sure which is better. Fortunately, Julia makes it really easy to experiment with these kinds of things. Of course, as soon as you have a system like that you need a scheduler for work at which point you want something like Dask or Dagger (https://github.com/JuliaParallel/Dagger.jl https://github.com/JuliaParallel/Dagger.jl).
- oxinabox 6y agoWe are so strong with the more general Tables.jl abstraction. Tables.jl took us a step away from having lots of types of `AbstractDataFrame`. Subtyping `AbstractDataFrame` is hard, I have to deal with two packages that do it and it is a big and not entirely documented interface. More natural is extending Tables.jl (like DataFrames and JuliaDB does). and we continue to build more tools that are table agnostic, and have APIs that DataFrames and a distributed table package special case when they can do it more efficienctly
- g0wda 6y agoThe basic Dask-like functionality is in Dagger.jl, we have plans to move out the distributed array functionality out of it and keep just the scheduler in there. I'm working on a set of functions to work with many DataFrames (or anything) in parallel and do it out of core if possible, it's basically like JuliaDB but with a FileTree abstraction rather than a table abstraction. http://shashi.biz/FileTrees.jl/lazy-parallel/ http://shashi.biz/FileTrees.jl/lazy-parallel/
- phillc73 6y agoWould Query.jl[1] be an alternative to consider? Their README says: The package currently provides working implementations for in-memory data sources, but will eventually be able to translate queries into e.g. SQL. There is a prototype implementation of such a "query provider" for SQLite in the package, but it is experimental at this point and only works for a very small subset of queries. Still early days, but sounds like they're working on the same problem. [1] https://github.com/queryverse/Query.jl https://github.com/queryverse/Query.jl