4 ms·
We also inspired from the same blog post ('Engineers Shouldn’t Write ETL') and built our own internal ETL tools. Our primary design goal was the system to be s
by armanboyaci 6y ago
We also inspired from the same blog post ('Engineers Shouldn’t Write ETL') and built our own internal ETL tools.
Our primary design goal was the system to be self-service for data scientists. Since our data scientists use pandas dataframes and jupyter notebooks all the time, we built the system around these two: (1) We have a library (that we call pype) acting an interface between the database and python dataframes (similar to .to_csv method), so there is no SQL queries in ETL scripts, (2) schedule (parametrized) notebooks using some special keywords.
We have a demo screencast: https://drive.google.com/file/d/1SVTduaIH_3IsJ-QoGI4mLYZE8JvCXThA/view?usp=sharing https://drive.google.com/file/d/1SVTduaIH_3IsJ-QoGI4mLYZE8Jv...
- seddonm1 6y agoLooks good. It is nice to see how much influence the 'Engineers Shouldn't Write ETL' post had! With Apache Arrow (https://arrow.apache.org/ https://arrow.apache.org/) I think the future looks very bright for both of our projects. It is important to have standard open source libraries and my early experiments have shown very good performance results.