9 ms·
Show HN: Kestra - Open-Source Airflow Alternative
Hey HN, I'm really proud to share with you my new open source project: Kestra https://github.com/kestra-io/kestra https://github.com/kestra-io/kestra
I created a few years ago a successful open source AKHQ project: https://github.com/tchiotludo/akhq https://github.com/tchiotludo/akhq (renamed from KafkaHQ) which has been adopted by big companies like Best Buy, Pipedrive, BMW, Decathlon and many more. 2300 stars, 120 contributors, 10M docker downloads, much more than I expected.
Now let's talk about Kestra, an infinitely scalable orchestration and scheduling platform for creating, running, scheduling and monitoring millions of complex pipelines.
I started the project 30 months ago and I'm even more proud of this project that required a lot of investment and time to build the future of data pipelines (I hope). The result is now ready to be presented and I hope to get some feedback from you, HN community.
To have a fully scalable solution, we choose Kafka as our database (of course, I love Kafka if you didn't know) as well as ElasticSearch, Micronaut, ... and can be deployed on Kubernetes, VM or on premise.
You may think there are many alternatives in this area, but we decided to take a different road by using a descriptive approach (low code) to build your pipelines allowing to edit directly from the web interface and deploy to production with terraform directly. We paid a lot of attention to the scalability and performance part which allows us to have already a big production at a big French retailer: Leroy Merlin
Since Kestra core is plugin based, many are available from the core team, but you can create one easily.
More information:
- on the official website: https://kestra.io/ https://kestra.io/
- on the medium post: https://medium.com/@kestra-io/introducing-kestra-infinitely-scalable-open-source-orchestration-and-scheduling-platform-8e4d47193616 https://medium.com/@kestra-io/introducing-kestra-infinitely-...
- check out the project: https://github.com/kestra-io/kestra https://github.com/kestra-io/kestra
Your comments are more than welcome, thank you!
- emteycz 5y agoThis looks incredible! Is there a way to use this as a managed service? Are you looking for independent partners/integrators?
- tchiotludo 5y agoFor now, we don't provide a SAAS for Kestra, it's definitely on the roadmap and our next project. In a meantime, we provide different installation : - Docker compose: https://kestra.io/docs/administrator-guide/deployment/docker/ https://kestra.io/docs/administrator-guide/deployment/docker... - Kubernetes: https://kestra.io/docs/administrator-guide/deployment/kubernetes/ https://kestra.io/docs/administrator-guide/deployment/kubern... - Jar: https://kestra.io/docs/administrator-guide/deployment/manual/ https://kestra.io/docs/administrator-guide/deployment/manual... Kestra is not so complicated to be installed, for Kafka and Elasticsearch, you could use Amazon managed service or Aiven for example. But be sure that we will provide a managed service as soon as possible
- emteycz 5y agoI'm definitely trying it out on my machine this evening. What's your position on other companies providing managed service of your project?
- tchiotludo 5y agoTo be honest, we don't have think about that for now. It's a complicated subject.
- chockchocschoir 5y agoTrue, you don't have to think about it because it's obvious: Anyone is allowed to provide a managed service of Kestra since you licensed the project under Apache License 2.0. Not that complicated. emteycz: I say go for it. If you get it up and running, let me know (email in profile) as I'm interested in trying it out as a managed service.
- emteycz 5y agoI'm more interested in integrating this into the software/ops of my customers at the moment. I just don't want to run a service right now, I like consulting more. I have a friend who has a data/ops cloud service. He's been looking for something like this. However in that case Kestra would be hidden behind his abstraction, I guess.
- deleted 5y ago[deleted]
- minroot 5y agoWhy Java?
- tchiotludo 5y ago- For performance mostly, Kestra rely a lot on Java thread to be able to handle a very large workload - Because the application is built on top of Kafka, and Kafka Streams that is only available on Java - Because the java ecosystem is very large and there is a lot good library to handle a lot of workload - Because I love strong typing and the language (but no matter for the user, just a personal pleasure :D)
- Hypocritelefty 5y ago
- Hypocritelefty 5y ago
- mekster 5y agoWhy is everyone ok with logs being dumped like it's in a trash can? I see partial structures and then JSON string as is and then some long blob of string no one can understand what it is with no new lines. What devs want are pretty simple, structured log with table layout without repeating the column names on every row to make it look insanely verbose for any human to consume. I'm picking up bits of open source apps to build a decent solution with Vector (which has awesome Vector remap language to parse strings into structured data if it isn't already) and throw it into ClickHouse to view it from Metabase. Apparently, Kibana, Graylog or even Grafana are pretty bad at displaying logs to even feel tiny comfortable reading it every day. Logging is such a crucial part of developer life and not sure why that there aren't any sane open source solutions.
- tchiotludo 5y agoSorry I don't understand, we display with pretty print, see here : https://kestra.io/assets/img/05.8b5545ef.png https://kestra.io/assets/img/05.8b5545ef.png It's not as json ? or I don't understand where you see that.
- nerdponx 5y agoI think the point is that logs should have a well-defined structure that can be easily parsed and queried by other programs.
- mekster 5y agoThis is what I see on the first page of the log in your demo. https://imgur.com/a/S1QkzuG https://imgur.com/a/S1QkzuG It can look a bit better with more config for these 3 other tools but these are pretty much what they're and the readability isn't any better than tail/grep. - Graylog : https://adamtheautomator.com/wp-content/uploads/2022/02/image-162.png https://adamtheautomator.com/wp-content/uploads/2022/02/imag... - Grafana : https://grafana.com/static/assets/img/blog/logcontext_explore.png https://grafana.com/static/assets/img/blog/logcontext_explor... - Kibana : https://blog.ip2location.com/wp-content/uploads/2018/12/logstash-filter-ip2location-screenshot.png https://blog.ip2location.com/wp-content/uploads/2018/12/logs... Comapred to displaying structured logs in Metabase. (This isn't text logs but you get the point.) - https://www.predictiveanalyticstoday.com/wp-content/uploads/2017/06/Metabase-1000x547.jpg https://www.predictiveanalyticstoday.com/wp-content/uploads/...
- wokwokwok 5y agoYou’re basically pitching this as a more complicated version of airflow that does basically the same thing, but slightly differently, and scales better? … but your core dependencies are a Kafka cluster and an elastic search cluster which are both a pain in the ass to scale; so really, could you run this seriously without a really expensive hosted cloud instance of both of those? This kind of wording: > Since the application is a Kafka Stream, the application can be scale infinitely Is a major turn off to me. Kafka cannot scale infinitely. Nothing can. In fact, Kafka can be a pain in the ass to scale. In makes me question some of the other commentary on the project.
- mountainriver 5y agoYeah and ES doesn’t scale forever, also running both of those is incredibly computationally expensive. You would really need the right use care
- tchiotludo 5y agoAgree that both are expensive to scale on multiple node. But keep in mind, you can use it with a single node (like others do with a database like mysql). Just don't go multiple node if not needed by the project. But when you will need to, with Kestra you can go multiple node and scale.
- jturpin 5y agoElasticsearch is not a pain in the ass to scale, it is one of the easiest databases to scale. Kafka is medium, since they ditched Zookeeper.
- chockchocschoir 5y agoEasy/hard is depending on the experience of the user. Someone with a lot of experience with Elasticsearch will have a easy time scaling Elasticsearch and hard time scaling Kafka, and vice-versa. Better to compare how complex they are to scale in terms of actions required.
- 5y ago
- sirjaz 5y agoAre there any plans for a desktop app for Kestra or the ability to support Windows Server outside of docker?
- tchiotludo 5y agoThe support of windows server seems to be easy I think. Since it's java behind, most of the api is working on windows. Just need to create a custom task for windows, added in the backlog : https://github.com/kestra-io/kestra/issues/519 https://github.com/kestra-io/kestra/issues/519 For the desktop app, I don't know, build one with electron can be simple, but a full app is not on the roadmap for now. What is your usages ?
- dantetheinferno 5y agoWhy is this better than Airflow, or Prefect, or Dagster?
- idomi 5y agoOr Ploomber?
- tchiotludo 5y agoAirflow have design issue and performance issue, If you want to have some details, you can find some reason on this article: https://kestra.io/blogs/2022-02-22-leroy-merlin-usage-kestra.html https://kestra.io/blogs/2022-02-22-leroy-merlin-usage-kestra.... For other workflow engine (dagster, prefect, ...), we decided to use a complete different approach on how to build a pipeline. Since others decide to use python code, we decided to go to descriptive language (like terraform for example). This have a lot of advantages on how the developer user experience is: With Kestra, you can directly the web UI in order to edit, create and run your flows, no need to install anything on the user desktop and no need a complex deployment pipeline in order to test on final instance. Other advantage is that it allow to use terraform to deploy your flows, typical development workflow are: on development environment, use the UI, on production deploy your resource with terraform, flow and all the others cloud resource. After, it will be really nice to have some independent performance benchmark. I really think Kestra is really fast since it was based on a queue system (Kafka) and not a Database. Since workflow are only events (change status, new tasks, ...) that is need to be consume by different service, database don't seems to be a good choice and my benchmark show that Kestra is able to handle a lot of concurrent tasks without using a lot of CPU.
- rajandatta 5y agoWe all may have questions for you on some of your descriptions and choices. None of that should take away from the fact that this is a pretty impressive stage for a 30-month open source project. Have not seen the participants - how many contributors do you have?
- tchiotludo 5y agoThe project start as a side project (yet another side project I do the night and weekend) but was quickly promoted and used in a French Big Retail Company. This one trust on the project and decide to go production with Kestra. So they decide to inject some resource in order to develop some features that need and that is missing. But basically, not so much people for now. We are trying to start a community around the product and started to communicate around the product since few weeks only, I hope community will follow us! And I hope to succeed like on my other open source project: https://github.com/tchiotludo/akhq https://github.com/tchiotludo/akhq
- speedgoose 5y agoCan you run software containers as steps ?
- tchiotludo 5y agoYes, of course! You have 3 solutions for that: - you can use this task using runner:DOCKER property and choose the image: https://kestra.io/plugins/core/tasks/scripts/io.kestra.core.tasks.scripts.Bash.html https://kestra.io/plugins/core/tasks/scripts/io.kestra.core.... - you can also use PodCreate to launch a pod on a kubernetes cluster: https://kestra.io/plugins/plugin-kubernetes/tasks/io.kestra.plugin.kubernetes.PodCreate.html https://kestra.io/plugins/plugin-kubernetes/tasks/io.kestra.... - you have also CustomJob from VertexAI on GCP to be able to launch a container a ephemeral cluster (with any CPU / GPU): https://kestra.io/plugins/plugin-gcp/tasks/vertexai/io.kestra.plugin.gcp.vertexai.CustomJob.html https://kestra.io/plugins/plugin-gcp/tasks/vertexai/io.kestr...
- speedgoose 5y agoGreat! That makes Kestra more useful than Dagster for me.
- sackerhews 5y agoCool. But please fix that light gray text on white background in the demo (or make it even paler for more avant-garde :)
- tchiotludo 5y agoDo you have a screenshot please ? I didn't notice where. Thanks
- chockchocschoir 5y agoI just noticed the title says "Open-Source Airflow Alternative" but Airflow is already Open-Source, so shouldn't you describe it as just "Airflow Alternative"? Otherwise you make it sound like Airflow isn't Open-Source but this is.
- tchiotludo 5y agoThe title was changed by moderators and I can't edit it anymore :'(
- tlrobinson 5y agoThese kinds of tools seem to be meant to scale up well, but are there good ones that “scale down” to small projects too?
- tchiotludo 5y agoYou can easily scale down for small project using a mono node using a simple docker compose setup: https://kestra.io/docs/getting-started/ https://kestra.io/docs/getting-started/ Working well on a standard laptop easily
- awild 5y agoAirflow is relatively easy to set up once you have the hang of it. At its most basic it needs three containers (server, sql, executor), and your dag definitions which are very straightforward python code.
- crubier 5y agoLooks cool ! How does it compare to temporal.io in your experience ? I’m evaluating options at my current company, between airflow and temporal.
- tchiotludo 5y agoTemporal.io is a really cool framework for building business process like managing microservice workflow (like paiement workflow: user pay, we call the shipping microservice, the billing microservice, ...) and good fit to handle individual event (lots of individual events). Kestra (and so airflow) is more a workflow manager to handle data pipeline like moving large dataset (batch) between different source and destination, do some transformation inside database (ELT) or with Kestra you are also able to transform the data (ETL) before save it to external systems. This lead Kestra (and so airflow) to have a lot of connectors to differents systems (like SQL, NOSQL, Columns database, Cloud Storage, ...) that is ready to use out of the box. temporal.io, since it's first design to handle microservice (proprietary & internal service) don't have this connector out of the box, and you will need code all this interaction. So my opinion: Building data pipeline interacting with many standard systems will be done easily & quickly with Kestra (or airflow) Handling internal business process of micro service will done easily with temporal.io
- jusonchan81 5y agoNetflix Conductor is great alternative to Temporal. There is a fully managed offering for this as well. The biggest advantage is that it’s quite simple to understand and has great visualization of flows.
- hackerdad 5y agoHave you tried Netflix Conductor (https://github.com/Netflix/conductor https://github.com/Netflix/conductor) - if you are evaluating between Airflow - this could be a great alternative - scales well and gives you option to write your workflows in code as well as config.