3 ms·
At the large tech companies I've been working at, there are many different big data engines that speak different dialects of sql (presto / spark). People write
by captaintobs 4y ago
At the large tech companies I've been working at, there are many different big data engines that speak different dialects of sql (presto / spark). People write SQL queries in one language and want to run it in another, but it doesn't just work, there are many parts of the query that need to be manually changed in order for it to run which is tedious and error prone.
- travisjungroth 4y agoWhere I work, we handle it by having Data Scientists running experiments on our platform commit their queries as Python code to a repository of metrics.
- captaintobs 4y agoI actually built the system you work on :). It was the main inspiration for this project because PyPika is not a great experience for data scientists.
- travisjungroth 4y agoThought you might catch that! I've actually helped swap a few things to SQLGlot from pypika.
- diehunde 4y agoGot it. I'm asking because I worked on something similar. The idea was to unit test some Airflow workflows locally. The production workflows were using Hive, but having a local Hive container was too slow for tests, so we wrote a small parser to translate the Hive queries into SQLite queries at runtime. In the end, we had a decent PoC but couldn't complete it because of all the Hive features, but it was super fun.
- captaintobs 4y agoyep, we do the same but using duckdb, it's got a lot more analytical functions and closer to hive than sqlite