4 ms·
Hey y'all, creator of siuba here--happy to answer any questions! One piece of context I try to bring into discussions is that the way I test and develop siuba
by closed 6y ago
Hey y'all, creator of siuba here--happy to answer any questions!
One piece of context I try to bring into discussions is that the way I test and develop siuba is by livecoding data analyses for an hour [1]. I encounter a lot of arguments like "X is possible with pandas", but when I sit down with analysts in realistic settings (e.g. time constrained) it turns out X works in more limited ways then they thought [2][3].
I'm a big fan of pandas though. It's what siuba is built on!
[1]: https://m.youtube.com/c/chowthedog https://m.youtube.com/c/chowthedog
[2]: https://mchow.com/posts/2020-02-11-dplyr-in-python/ https://mchow.com/posts/2020-02-11-dplyr-in-python/
[3]: https://siuba.readthedocs.io/en/latest/developer/pandas-group-ops.html https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
- ellisv 6y agoHow's the performance? I'm certainly willing to give up a little computational performance for being able to write my code faster.
- closed 6y agoUsing the experimental fast grouped pandas functions, it should run at the speed of optimized pandas code! Since siuba functions just run on pandas DataFrames, you can always hand tune for performance, but imo most of the time pandas code runs slow it's because of something like .agg(lambda ...) somewhere. There's an example with timings here: https://siuba.readthedocs.io/en/latest/developer/pandas-group-ops.html https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
- tpoacher 6y agoIs the pipe operator >> specific to pandas objects or is it generally applicable? I'd like to see how you implemented it!