23 ms·
Open Source Python ETL
- mitjafelicijan 2y agoThis is actually exactly what I needed for my current project!
- thibautdr 2y agoThanks for your comment, don't hesitate to share your use case! Also, you can reach out on Slack if you have any questions or need help.
- thibautdr 2y agoHi everyone, thanks for posting Amphi :) To give some context, Amphi is a low-code ETL tool for both structured and unstructured data. The key use cases include file integration, data preparation, data migration, and creating data pipelines for AI tasks like data extraction and RAG. What sets it apart from traditional ETL tools is that it generates Python code that you own and can deploy anywhere. Amphi is available as a standalone web app or as a JupyterLab extension. Visit the GitHub: https://github.com/amphi-ai/amphi-etl https://github.com/amphi-ai/amphi-etl Give it a try and let me know what you think
- slt2021 2y agoi liked the idea of leveraging jupyterlab as server. data engineers/scientists already use jupyter, so this is neat idea. custom extension for jupyterlab is a great way to leverage existing jupyterlab install base: not everyone will be willing to install and jump through hoops to install software X, but installing extension is one pip install away and no need to run separate process, since you are running inside jupyterlab server. this reminds of ALTERYX (another drag and drop ETL tool)
- thibautdr 2y agoThanks! Being based on JupyterLab also allows Amphi to benefit from the vast ecosystem of extensions already available, such as the Git extension or using different file systems (S3). Some users pointed out they were Alteryx users but liked the Python code generation from Amphi :)
- slt2021 2y agojust an idea: is it possible to code generate Airflow code? since a lot of companies use airflow as ETL orchestrator
- thibautdr 2y agoAmphi generates Python code, so you can definitely orchestrate them through Airflow but it doesn't generate "Airflow code" so to speak. Now, in the future we might develop Airflow specific workflows or maybe operators.
- johhns4 2y agoWow amazing work! How does the inputs work, are they created for you or does it support custom as well?
- OutOfHere 2y agoYou know what does not set it apart? AI-washing. Also, lying about being open source when it isn't.
- isjamesalive 2y agoTo be fair, the only place the words ‘open’ and ‘source‘ appear in the readme are once in a sub-heading, where it’s phrased ‘open-source’. It’s clearly labelled ELv2. Possibly more of a subtle miscommunication or misunderstanding than a deliberate lie.
- OutOfHere 2y agoDon't kid yourself. The title of this submission itself starts with "Open Source". Moreover, the author has made the explicit decision to not fix the readme.
- v3ss0n 2y agoWhat's the difference compare to Windmill.
- thibautdr 2y agoHi, thanks for your question. I'm not familiar with Windmill, but after checking it seems to be an open source developer platform to build applications. Amphi is a low-code tool to develop data pipelines (or ETL pipelines).
- esafak 2y agoWindmill is a low-code workflow engine: https://www.windmill.dev/flows https://www.windmill.dev/flows
- whalesalad 2y agoBeen happy with Dagster but this looks interesting.
- c0brac0bra 2y agoConsidering switching some ancient Talend and Airflow processes over to this if I can get the time.
- tayloramurphy 2y agoI'm curious as to the story of how things like this come to be. It seems like there are already a ton of "open source python ETL" tools on the market. Was this a passion project by the author? Was this born out of academia? Was there a specific problem they were trying to solve that others didn't? It's not necessary to answer these questions in the docs but it is useful for folks who may be familiar with the other options out there.
- thibautdr 2y agoThanks for your comment, those are valid points. I come from the industry, having worked for an ETL vendor for 6 years. I've personally witnessed a need for a low-code (graphical) ETL for Python environments. In short, traditional ETLs are GUI ETLs for Java environments while modern data tools are either focusing on the EL part or are code-oriented (dbt). With Amphi, I want to offer a low-code graphical alternative to develop Python-based pipelines. I also believe that modern data stack tools don't effectively address use cases for unstructured data, which is another focus of Amphi (with extensive file integration and RAG support).
- tayloramurphy 2y agoAppreciate the reply! Thanks for the context.
- iblaine 2y agoLow code ETL tools (informatica, Appworx, talend, pentaho, ssis) were the original services for ELT/ETL. A lot of progress was made to go towards ETL-as-code starting with Airflow/Luigi. Going back to low code seems backwards as this point. (I have used all of the above tools in my 15+ yr career. Code as ETL was a huge industry shift)
- deleted 2y ago[deleted]
- thibautdr 2y agoThanks for your comment! I do believe it depends on who you ask and ultimately both will co-exist. I also think low-code solutions democratize access to ETL development offering a significant productivity advantage for smaller teams. With Amphi, I'm trying to avoid the common pitfalls of other low-code ETL tools, such as scalability issues, inflexibility, and vendor lock-in, while embracing the advantages of modern ETL-as-code: - Pipelines are defined as JSON files (git workflow available) - Generates non-proprietary Python code: This means the pipelines can be deployed anywhere, such as AWS Lambda, EC2, on-premises, or Databricks.
- banku_brougham 2y agoIm very leery of low code, but I like the idea of ETL defined as configuration.
- stoperaticless 2y agoEtl as text is good, because you can save it in version control. (Is it “code” or “json” is irrelevant for the vcs) Edit: save in vcs stringly implies usability of ‘diff’ and ‘grep’
- sqlcook 2y agoYou’re missing the point of the benefits of solutions like these, and the original set of tools like the Informatica of the kind. Those tools come with limitations and constrains, like a box of legos you can build a very powerful pipeline without having to wire up a lot of redundant code as you pass data frames between validation stages. Tools like Airflow/Spark etc are great for what they are, but they don’t come with guidelines or best practices when it comes to reusable code at scale, your team has to establish that early on. You can open a pretty complicated large DAG in and right away you’ll understand the data flow and processing steps. If you were to do similar in code, it becomes a lot harder unless you comply to good modular design practices. This is also why common game engine and 3d rendering tools come with a UI for flow driven scripting. It’s intuitive and much easier to organize.
- ic_fly2 2y agoWith all the data issues strong quality and normalisation I often get the impression that enabling more people with non CS backgrounds to do this work is not necessarily a good thing. In other words, if writing python and sql is the skill requirement that stops you from making an etl pipeline, maybe do something else.
- pm90 2y agoWith this argument, Computer Science wouldn’t have progressed beyond assemblers.
- anakaine 2y agoThis is elitist and frankly, unhelpful. The answer to a skills shortage is not a practitioner lockdown, but policy, training, guidance and mentoring. If you're stuck in start up land and you have this issue, you have hired the wrong skills. If you're encountering this in enterprise land, your organisation, and potentially you depending on your position of influence, should be angling to improve compliance and literacy not through obstruction but through policy and upskilling. Failing to do so will kill your ability to innovate.
- necovek 2y agoFWIW, while I disagree with the parent comment, I don't see you arguing against it. They actually implied that you should try upskilling first — but if that fails, you shouldn't be doing ETL yourself. I mostly disagree with the parent comment because there's so many things one can easily do up to a level, and then when the going gets tough, you need to call in an expert. Eg. most people can operate a screwdriver or impact driver to fix things, but to fix some problems, you really need a trained technician (or well, an experienced DIY person, but that's not everybody). The fact that you are not strong enough to screw in an M14 bolt does not mean you should be forbidden from using an impact driver: tools are there to help you. The logic of the parent comment was seemingly that if you are not strong enough to tighten an M14 bolt, you probably don't know what you are doing regardless of the type of the bolt you are tightening, so you should simply not do it. The point I agree with in a parent comment is that not everybody can achieve a similar level of proficiency: while upskilling and improving/simplifying tools can get you most of the way there, there's always going to be that extra bit that requires a sudden, sharp jump in knowledge, smartness or experience to be able to deal with it.
- cvalka 2y agoTHIS IS NOT OPEN SOURCE!
- anakaine 2y agoOK
- maphew 2y agoIt's published on GitHub under license ELv2 - Elastic License v2. This does not meet the open source definition, so indeed it's not Open Source. ELv2 is an open source sibling though, closer than many other openish licenses: https://www.elastic.co/pricing/faq/licensing https://www.elastic.co/pricing/faq/licensing Still, Amphi should not claim to be 'Open Source'.
- runningmike 2y agoIMHO this title is primarily chosen for promotion. Not needed. But unfortunately many people, young and old, have never heard of OSI.
- anakaine 2y agoHey, I really like the design. I currently have a lot of ETL going on through various mechanisms, but the thing that is always difficult to communicate to BAs and PMs, and any other individual is a graphical "what is this thing doing and how". This is neat for those of us who are visual.
- thibautdr 2y agoThanks! Don't hesitate to give it a try and reach out if you need anything :)
- tiraz 2y agoHow does it distinguish itself from Dagster or Prefect? Both are there for quite some time, also have a GUI, but a much larger feature set.
- FranzFerdiNaN 2y agoBoth don’t have the drag and drop feature of this, you have to write python yourself.
- vekker 2y agoDoes this also manage the infrastructure side of ETL? Usually some parts in a complex ETL process take a lot more processing power, so are run on different machines. From a quick glance at this, it seems like a WYSIWYG ETL tool for running ETL jobs on one machine?
- thibautdr 2y agoThanks for your question. Amphi generates Python code using Pandas and can scale on a single machine or even multiple machines using Modin, but the process is manual for now. Future plans include deploying pipelines on Spark clusters and other services such as Snowflake.
- olavgg 2y agoThis looks visually similar to Apache Nifi.
- awesomebytes 2y agoI was not familiar with the acronym ETL and it is not explained anywhere in the website! My feedback would be to at least write it once, on the first instance so others like me will know what they are reading :)
- ghoshbishakh 2y agoIt is written on the website: Extract, transform and load. Yes, an illustrative example description would help I agree.
- thibautdr 2y agoThanks for pointing that out, it's actually mentioned (Extract, transform and load ...) in the very first sentence below the tagline, but if you didn't get it then it's not clear.
- ljouhet 2y agoDidn't understand either: Extract, transform and load... vs ETL: Extract, transform and load data... Extract, transform and load (ETL) data... Extract, Transform and Load data...
- fuzztester 2y agoOthers already replied about what ETL is. Wikipedia: https://en.m.wikipedia.org/wiki/Extract,_transform,_load https://en.m.wikipedia.org/wiki/Extract,_transform,_load I'll just add: It is a common term and practice among enterprise software users, i.e. generally medium or large companies that use packaged plus custom software for their business needs. ETL is not common among startups, because they have a different focus, infrastructure and scale.
- m463 2y agoIn computing, extract, transform, load (ETL) is a three-phase process where data is extracted from an input source, transformed (including cleaning), and loaded into an output data container. The data can be collated from one or more sources and it can also be output to one or more destinations. https://en.wikipedia.org/wiki/Extract,_transform,_load https://en.wikipedia.org/wiki/Extract,_transform,_load
- kkfx 2y agoDo not take me wrong, I appreciate and thanks anyone who contribute to FLOSS, but all low/no code approaches I see turn out to be garbage. IMVHO the reality is that people need to be trained and became capable of fishing alone instead of giving them fishes all days. ML in ETL is needed for raw initial classification of documents received in various formats from various sources, to clean-up scanned crap, no more than that, all the effort to plug LLMs was so far and i bet will be for the next 10 years a disaster. ETL is something that should not exists in a modern world because we should exchange data in usable formats instead of having to import the with all sort of gimmick, we do not have such acculturated world but at least we can try to simplify and teaching instead of adding entropy.
- jamesblonde 2y ago#dang The title needs changing - it's not open-source, it is license ELv2 - Elastic License v2.
- maleldil 2y agoWhile you're right (it's indeed not open-source), the project advertises itself as such, so the title is "right", even if it's a lie.
- lma21 2y agoisn't the code available here? https://github.com/amphi-ai/amphi-etl https://github.com/amphi-ai/amphi-etl what makes it not OSS?
- uneekname 2y agoThat is source available, not open source. The term "open-source" is widely used to describe software that is licensed using a specific set of software licenses that grant certain freedoms to users. You can read more here[0] [0] https://opensource.org/licenses https://opensource.org/licenses
- mrtranscendence 2y agoIt's open source if you're using language like a normal human being. If you're a bit of a pedant and wish everyone to adhere to definitions imposed from on high regardless of real-world usage, it's absolutely not open source.
- cvalka 2y agoIt's source available. It's definitely not open source. It's deception plain and simple.
- deknos 2y agoIs it true opensource / free software, or are there non opensource parts?
- Joeboy 2y agoSince there are "ETL" people here, I have a couple of naive questions, in case anybody can answer: 1) Are there any"standard"-ish (or popular-ish) file formats for node-based / low-code pipelines? 2) Is there any such format that's also reasonably human readable / writable? 3) Are there low-code ETL apps that (can) run in the browser, probably using WASM? Thanks and sorry if these are dumb questions.
- roenxi 2y agoThey're good questions, but they are not answerable blind. The correct choices depend too much on what problems you are trying to solve, the formats and scale of the data involved, the tolerances for downtime and what other software is being used. My advice is to avoid, in general, low code tools if you plan to have software engineers involved. And once there aren't any software engineers whatever gets built is going to be a mess by software engineering standards so just roll with it. Any tool is equally likely to hit your pain points (and generate an unmanageable mess).
- thibautdr 2y agoThanks for the great questions: 1. As far as I know, there isn't a "standard" file format for low-code pipelines. 2. Some formats are more readable than others. YAML, for example, is quite readable. However, it's often a tradeoff: the more abstracted it is, the less control you have. 3. Funny you ask, I actually tried to make Amphi run in the browser with WASM. I think it's still too early in terms of both performance and limitations. Performance will likely improve soon, but browser limitations currently prevent the use of sockets, which are indispensable for database connections, for example.
- whazor 2y ago"Python ETL", Github language statistics: TypeScript 87.1% It looks nice though.
- gregw2 2y agoIsn’t pandas centric ETL much more memory intensive and less compute efficient than using SQL?
- thibautdr 2y agoI wrote an article questioning the use of Pandas for ETL. I invite you to read it: https://medium.com/@thibaut_gourdel/should-you-use-pandas-for-etl-0ecf56c84d57 https://medium.com/@thibaut_gourdel/should-you-use-pandas-fo...
- rkozik1989 2y agoThat's kind of the tradeoff you make with any low-code/no-code technology. You leverage prebuilt components and string them together to achieve some kind of task. Which isn't most efficient thing in the world to do but it does work assuming you have enough compute resources to throw at it, and return what you generally achieve is an end product that's completed faster than the traditional development route. You could just use SQL but then you'd have to develop and test the entire infrastructure to support your component-oriented architecture from scratch, and at that point you're kind of just reinventing the wheel because that's basically just pandas with less features. Low-code is kind of just Authorware for a new generation... assuming you're old enough to remember that technology.
- C4stor 2y agoIt's a good idea, but from the docs it looks like the high level abstractions are wrong. If my data pipeline is "take this table, filter it, output it", I really don't want to use a "csv file input" or a "excel file output". I want to say "anything here in the pipeline that I will define that behaves like a table, apply it this transformation", so that I can swap my storage later without touching the pipeline. Same things for output. Personally I want to say "this goes to a file" at the pipeline level, and the details of the serialization should be changeable instantly. That being said, can't complain about a free tool, kudos on making it available !
- thibautdr 2y agoHey, not sure I get your point here. I believe the abstraction provides what you're describing. You can swap a file input with a table input without touching the rest of the components (provided you don't have major structural changes). Let me know what you meant :)
- mritchie712 2y agoIf you're looking for "open source Python ETL", two things that are better options: https://dlthub.com/ https://dlthub.com/ https://hub.meltano.com/ https://hub.meltano.com/ we[0] use meltano in production and I'm happy with it. I've played around with dlt and it's great, just not a ton of sources yet. 0 - https://www.definite.app/ https://www.definite.app/
- jdnier 2y agoAre you able to describe what makes them better? (Honest question, I'm not familiar with either or with Amphi.) It seems Definite's use case is focused on connecting to lots of data sources. For much smaller scale, how does Amphi compare?
- mritchie712 2y agomost data engineers would think of something like Fivetran when you say "ETL" (look at the ETL section here[0]). It looks like Amphi could handle some low code transformations (the "T" in ETL), but calling it ETL feels like a stretch. So to rephrase a bit, if you're looking for an open source, python based Fivetran alternatives, dlt and meltano would be my picks. 0 - https://mattturck.com/landscape/mad2024.pdf https://mattturck.com/landscape/mad2024.pdf)
- thibautdr 2y agoHey, Amphi's developer here. Those two tools are great, big fan of dlt myself :) However, Amphi is a low-code solution while those two are code-based. Also, those two focus on the ingestion part (EL) while Amphi is focusing on different ETL use-cases (file integration, data preparation, AI pipelines).
- mritchie712 2y agoI understand that. I'd change the title / H1 though, "Open Source Python ETL" doesn't describe what you're building very well. Good luck! Looks cool.
- nextworddev 2y agoIf you are enterprise, just go with Databricks lakeflow
- throwaway984393 2y ago[dead]
- mrwyz 2y agoNot open source. Misleading title.
- paulvnickerson 2y agoVery cool, thanks for sharing. Does it support the pandas-like rapidsai dask_cudf framework? (https://docs.rapids.ai/api/dask-cudf/stable/ https://docs.rapids.ai/api/dask-cudf/stable/)
- thibautdr 2y agoGreat, thanks for sharing. I was familiar with Dask and cudf separately but not this one.I was planning to implement dask support through Modin but I'll definitely take a look at dask_cudf.
- paulvnickerson 2y agoCool. We use it a lot at work for working with large data sets on a GPU cluster.
- Kalanos 2y agoReminds me of Elyra
- thibautdr 2y agoYes, there are similarities, but Elyra allows you to develop orchestration pipelines for Python scripts and notebooks, so you still have to write your own code. With Amphi, you design your data pipelines using a graphical interface, and it generates the Python code to execute. Hope that helps.
- febed 2y agoWhich open source Python based ETL tool would one recommend for someone starting an ETL project today? It’s a data volume heavy project with lot of interdependencies between import tasks.
- rldjbpin 2y agoas open source as open weights models, but will companies adopt it solely on pricing?