5 ms·
These tools are very new compared to CSV+pandas. And most things you want to get data out of won’t give it to you in parquet. The future is very promising, I a
by teej 4y ago
These tools are very new compared to CSV+pandas. And most things you want to get data out of won’t give it to you in parquet.
The future is very promising, I am personally very excited about DuckDB. But it’s too soon to be griping about old tutorials.
- datalopers 4y agoThese aren’t really new though. These are simply bringing old-school reliable SQL directly into Python. Hopefully Python “data” people will stop using godforsaken tools like pandas or spark and finally embrace what the rest of us have been using for decades.
- sakras 4y agoAre you kidding?! SQL is an ancient mess with different dialects and idiosyncrasies. What bringing a query engine into Python gives is a functional query DSL. DuckDB’s non-SQL interface, where you build up a query plan from nodes in Python is the future: you can give names and reuse subqueries, you can programmatically generate query trees, you get all of the power of a real programming language for analytics, and it all gets boiled down to fast implementations of relational operators.
- datatrashfire 4y agoAs someone who is fluent in Spark, pandas, and several SQL "dialects", at the end of the day I have to say I think all the query paradigms suck, SQL is just the best we have. I find it's "limitations" generally enforce structure that makes rewriting and distilling bad queries easier. Well written and formatted queries can be extremely readable and easy to navigate/grep in a way that I have never found the others to be. The flexibility of dropping in and out of a full fledged programming language and a query api, I find it tends to grow unnecessarily complex in the hands of many practitioners. Although my preference is a SQL cursor or equivalent interface in my chosen language. For some reason I find the strict separation between SQL (the declarative expression of business logic and relations) and the language (imperative or functional control) very helpful. SQL also has the benefits of portability. Almost every query computation engine supports SQL these days. While, obviously you will need to rewrite the parts that rely on unique extensions, the migration path is greatly simplified. The one thing I think pandas has going for it that I desperately wish was picked up as a new standard in SQL are aggregations for seemlessy moving between different time series frequencies. Pandas as problematic as it is, I have yet to find anything else that makes time frequency conversions as convenient and predictable.
- datalopers 4y ago> DuckDB’s non-SQL interface, where you build up a query plan from nodes in Python is the future DuckDB is a convenient drop-in for OLAP in the way that SQLite is for OLTP, it's a great library. Python's death on the vine as a data analytics platform cannot come fast enough. There's a reason people are abandoning pandas and spark/databricks in droves and fleeing to useful SQL-like tooling such as Snowflake. Once this tech bubble bursts, and the tens-of-thousands of subsidized Python-packing data scientists/engineers/analysts get laid off, the language will return to what it should be: another scripting language that generates SQL for talking to a real data engine.
- bigcat123 4y ago