4 ms·
I think the link I put below probably does a better job explaining what I'm looking for. This has mainly to do with dataframes (not unstructured text). Here's t
by geebee 5y ago
I think the link I put below probably does a better job explaining what I'm looking for. This has mainly to do with dataframes (not unstructured text). Here's the link again in case this thread gets long and you don't know what I'm referring to:
https://pandas.pydata.org/pandas-docs/stable/getting_started/comparison/comparison_with_sql.html https://pandas.pydata.org/pandas-docs/stable/getting_started...
Generally, I vastly prefer the SQL operations to the pandas ones, though (and this is very important) only when pandas is essentially recreating what is in SQL's sweet spot. For example, you can use pandas operations to do joins, aggregations, filters, and so forth. I would rather write that code in sql.
I would not prefer to generate summary statistics in SQL, find correlations between columns, or do other things that are in the realm of scientific or statistical programming. There will be a grey area in there, for sure. I also find that many things that require SQL trickery (such as self-joins) often have a very very simple pandas solution such was cumulative sums on a column. So I go back and forth between SQL and pandas quit a bit (as each operation returns a data frame).
Just to be clear again, some people just can't stand SQL and want to stay away from it as much as possible. Other people, like me, greatly prefer it, but even for us there are scenarios where we'd much rather use pandas than get into leetcode style SQL trickery.