3 ms·
> but where it generally will lose out is when you’re chaining together multiple operations which a compiler will vectorise more efficiently I think the bigger
by vlmutolo 4y ago
> but where it generally will lose out is when you’re chaining together multiple operations which a compiler will vectorise more efficiently
I think the bigger issue is remaining cache-friendly in the face of multiple operations. If you do the operations one at a time, you’re doing multiple passes over the data.
In any case, the approach that Polars[0] took seems like a pretty good solution, if difficult to implement well.
Polars is a dataframe library, and has a “lazy” API. This builds up a graph of operations, runs it through a query optimizer, and then executes them all at once. This allows for parallelization, optimizing expressions, and fewer passes over the data. The downside is that it’s nontrivial to implement a query optimizer.
[0]: https://pola.rs/ https://pola.rs/