4 ms·
I would not say pandas is a recreation of SQL/Spark, they have very different use-cases in my experience. SQL/Spark is like a bulk data management tool: I use i
by tech_ken 2y ago
I would not say pandas is a recreation of SQL/Spark, they have very different use-cases in my experience. SQL/Spark is like a bulk data management tool: I use it if I need to load massive data from some remote store, perform light preprocessing, join up a couple dimension tables, etc. Having normalized and joined my data, then pandas enters as a 'last-mile' processing engine, particularly when paired with ex. SKLearn or whatever other inference lib you're using. Pandas is awesome if you need to ex. apply string manipulation to a data table or daisy-chain some complicated computations together. Honestly in my opinion the API is really nice, I came over from R tidyverse and the 'chained methods' approach to pandas let's me carry over all my old patterns and paradigms. I find it far easier to use that approach than having to write like 20 dependent subqueries or staging tables