3 ms·
It's a lovely idea to build pandas like functionality on top of NumPy's structured dtypes, but these benchmarks comparing PandaPy to Pandas are extremely mislea
by shoyer 7y ago
It's a lovely idea to build pandas like functionality on top of NumPy's structured dtypes, but these benchmarks comparing PandaPy to Pandas are extremely misleading. The largest input dataset has 1258 rows and 9 columns, so basically all these tests shows is that PandaPy has less Python overhead.
For a more representative comparison, let's make everything 1000x larger, e.g.,
closing = np.concatenate(1000 * [closing])
Here's how a few representative benchmark change:
- describe: PandasPy was 5x faster, now 5x slower
- add: PandasPy was 2-3x faster than pandas, now ~15x slower
- concat: PandasPy was 25-70x faster, now 1-2x slower
- drop/rename: PandasPy is now ~1000x faster (NumPy can clearly do these operations without any data copies)
I couldn't test merge because it needs a sorted dataset, but hopefully you get the idea -- these benchmarks are meaningless, unless for some reason you only care about manipulating small datasets very quickly.
At large scale, pandas has two major advantages over NumPy/PandasPy:
- Pandas (often) uses a columnar data format, which makes it much faster to manipulate large datasets.
- Pandas has hash tables which it can rely upon for fast look-ups instead sorting.
- meowface 7y agoThis is why you can never accept benchmarks provided solely by the software creators. Same for accepting studies about a company's product when the company's commissioned and funded the studies. It'd be cool if there were neutral third-parties, kind of like Jepsen, that any project could defer rigorous benchmarking to, perhaps in exchange for a flat fee (everyone pays the same fee, no matter how big or small they are).
- munmaek 7y agoAnd then they learn how to game the benchmarks. You just can’t win.
- skrebbel 7y agoNo, because the trick is that you're paying a knowledgeable person to run the benchmark. That person would presumably actively iterate on the benchmarks and try to detect / avoid cheating.
- kmbriedis 7y agoPeople would probably find out what hardware they use for benchmarks and optimize for that, leading to performance decrease for many othes
- deleted 7y ago[deleted]
- mhh__ 7y agoI am writing a small benchmarking site (aimed at asymptotic performance rather than singular tasks) which should hopefully have a fairly generic API/Specification, so I'll look into the more "competitive" side of it (Project vs Project as opposed "Uh oh, Commit f4r0adfja has made the build slower and [test 3] slower for large n". The logic itself is easy but I'm having to step into the dark arts of DevOps and web-cancer - should be online by the summer.
- cerved 7y agoTo be fair, it's clearly stated in the readme: "The performance claims only hold for small datasets, 1,000-100,000 numpy rows. Pandas perform better with larger data sets, the only functions that improve with a 1000x increase in size is rename, column drop, fillna mean, correlation matrix, value reads, and np calculations even out (np.log, np.exp as well as etc)"
- lmeyerov 7y agoThis is a real problem, akin to how modern JITs will have diff run modes for the same code. We're overdue for pandas-without-huge-overhead. See this thread for more on why pandas overhead adds up for real settings, and is crazy huge once you go to dask/ray/modin, and even worse, say spark: https://twitter.com/lmeyerov/status/1218093436286296069 https://twitter.com/lmeyerov/status/1218093436286296069 + https://twitter.com/lmeyerov/status/1220847229537157121 https://twitter.com/lmeyerov/status/1220847229537157121 . For interactive analytics apps, we want scale for bigger workloads, yet also less overhead to achieve < 100ms: ideally moving a slider will trigger lots of calcs over many widgets and their subcomponents, and the underlying stats code looks normal (e.g., df > sql > some imperative/functional lang.) Internally we write C-like buffer manips & hand-written GPU kernels & weird streaming code to achieve < 100ms wall-clock, and are trying to get tech like these up to snuff for interactive data software. As per the JIT comment, feels like a matter of time, so happy to see this (and the pressure it creates!)
- shoyer 7y agoMost of those disclaimers were added tonight, after seeing my comment :) https://github.com/firmai/pandapy/commit/692c968771bb19d4f12fac3d1eb76f8e611db165 https://github.com/firmai/pandapy/commit/692c968771bb19d4f12...
- missosoup 7y agoNot only that, but the benchmarks ignore the cost of translating back and forth between pandas format. By the time you've done a comprehensive benchmark, you'll realize that pandas is actually pretty well optimised and the only opportunity for speedup is writing specialised functions for a narrow set of use cases. This library is a small box of such specialised functions. It will never be able to compete with pandas in general. cuDF is a more plausible candidate for replacing pandas in performance-critical scenarios, and even cuDF explicitly aims to supplement pandas rather than replace it.