4 ms·
Nah, I'll do it with SQL
by _aaed 3y ago
Nah, I'll do it with SQL
- rectang 3y agoOne nice thing about solutions like pandas or Julia is that they’re much easier to write tests for or otherwise validate. I can’t tell you how many times I’ve been handed a big ball of SQL which doesn’t behave like its author thinks it does, diverging in subtle or not-so-subtle ways.
- ivirshup 3y agoibis in python is a really nice middle ground. Nice API + in a programming language, but executes on database backends (which could be polars or duckdb on in memory arrow tables).
- akdor1154 3y agoI agree, and interestingly so does the author of the article so it's a bit weird that you're receiving downvotes. Julia's DataFrames library is more consistent than Pandas by a mile, but it's still a bit weird. JuliaDB's IndexedTable and NDTable was a really awesome API design, it's quite a pity that JuliaDB is now unmaintained. :(
- pjmlp 3y agoSame here, I don't get the point of this other that "don't want to learn SQL".
- smabie 3y agoYou're saying all Pandas usage (an incredibly popular library) is because people don't want to use SQL?
- cookieperson 3y agoIt's a broken argument with some truth too it. IE you can run a SQL query put the result in a dataframe to dump it to an interchange format. But in the same breathe... If you learn to use SQL a great deal of workloads often used via dataframes APIs kind of disappear, and in doing so learning a new language to use a new dataframes API isn't really worth it.
- pjmlp 3y agoAs far as I am aware, plenty of use cases can also be done via OLAP.
- laratied 3y ago[dead]
- cookieperson 3y agoA lot of data scientists don't know SQL and don't understand why people use it. That said, there are cases where in memory workloads and certain manipulations are less encumbered by dataframes APIs. But in a lot of cases... They get abused by people who really do need to push themselves a little bit to learn something new.
- cjalmeida 3y agoYou use both. Once your data fits comfortably in memory it's naive to try to build histograms, pivots and charts using pure SQL.
- cookieperson 3y agoSQLite is often much faster then dataframes jl and pandas.
- ayhanfuat 3y agoI highly doubt that SQLite is faster than pandas, let alone dataframes.jl, for analytical workloads.
- cookieperson 3y agoMight surprise you how many people are using pandas or dataframes for OLTP on a daily basis because they don't know better.
- cjalmeida 3y agoSQLite is good for a bunch of stuff, but it's terrible for analytic workloads. Not even in the same ballpark for pandas, let alone Julia
- agacera 3y agoSame here. But have you tried duckdb? You can do sql in the pandas dfs and it is fast af. https://duckdb.org/2021/05/14/sql-on-pandas.html https://duckdb.org/2021/05/14/sql-on-pandas.html
- cookieperson 3y agoDuckdb is sick. You can also do queries on parquet, etc.
- avnigo 3y agomydf = pd.DataFrame({'a' : [1, 2, 3]}) print(duckdb.query("SELECT SUM(a) FROM mydf").to_df()) I can see the appeal, but if you're working in Python, something doesn't sit right with me when having to write out variable names as strings. E.g., if I want to refactor the code, my LSP or parser won't pick up those references. > The SQL table name mydf is interpreted as the local Python variable mydf [...] Not only is this process painless, it is highly efficient. It might be painless and convenient at first, but I feel like this could get you in trouble down the line. Is there a way to avoid this?
- freilanzer 3y agohttps://duckdb.org/docs/guides/python/ibis.html https://duckdb.org/docs/guides/python/ibis.html