5 ms·
My name is Ask and I'm the co-creator of this project, along with Vineet. Yes, we're definitely interested in supporting Redis Streams! Faust is designed to s
by asksol 8y ago
My name is Ask and I'm the co-creator of this project, along with Vineet.
Yes, we're definitely interested in supporting Redis Streams!
Faust is designed to support different brokers, but Kafka is the only implementation as of now.
- simonw 8y agoWell I love Celery, so I'm even more excited to learn more about Faust now :)
- asksol 8y agoI'm looking at Redis Streams and it seems to lack the partitioning component of Kafka. One of the properties that we rely on is the ability to manually assign partitions such that a worker assigned partition 0 of one topic, will also be assigned partition 0 of other topics that it consumes from. Simplicity is of course a goal, but this may mean we have to sacrifice some features when Redis Streams is used as a backend. I'm glad you like Celery, this project in many ways realize what I wanted it to be.
- welder 8y agoIf you liked Celery 3.x but don't like all the bugs in Celery 4.x, check out these alternative Python task queues: https://github.com/closeio/tasktiger https://github.com/closeio/tasktiger (supports subqueues to prevent users bottlenecking each other) https://github.com/Bogdanp/dramatiq https://github.com/Bogdanp/dramatiq (high message throughput)
- asksol 8y agoIn Faust we've had the ability to take a very different approach towards quality control. We have a cloud integration testing setup, that actually runs a series of Faust apps in production. We do chaos testing: randomly terminating workers, randomly blocking network to Kafka, and much more. All the while monitoring the health of the apps, and the consistency of the results that they produce. If you have a Faust app that depends on a particular feature we strongly suggest you submit it as an integration test for us to run. Hopefully some day Celery will be able to take the same approach, but running cloud servers cost money that the project does not have.
- adambard 8y agoI'd love to see some docs/examples around supporting pluggable brokers; Faust looks like exactly what I want from a python stream processor, but for various technical and operational reasons it would be nice to be able to provide a RabbitMQ/AMQP broker.
- asksol 8y agoWe added two levels of abstractions for this purpose: - A Stream iterates over a channel - A Channel implements: `channel.send` and `channel.__aiter__`. - A topic is a "named" channel backed by a Kafka topic - Further the topic is backed by a Transport - Transport is very Kafka specific To implement support for AMQP the idea is you only need to implement a custom channel. If you open an issue we can consider how to best implement it.
- SpaceManNabs 8y agoHey, I am a huge user of Apache Flink in personal and work projects. Skimmed through your readme, but still have the following question. What are the advantages of using Faust feature-wise or otherwise? Are you guys planning on having feature parity with Flink (triggers, evictors, process function equivalent, etc.)? I can definitely see its use in having better compatibility with ML/AI env and the other tools in python toolkit. But specifically regarding ML/AI, I can export those things into the JVM usually. And with regards with Flask/Django, I can use scala http4s. Sorry if that seemed like a lot! Trust me very excited about this project!
- sitkack 8y agoYou are the first Flink user I have come across! Could you briefly explain why it is so compelling over its competitors?
- SpaceManNabs 8y agoWhen it comes to batch, I still think Apache Spark and the standard scientific computing stack in Python are king. But when it comes to streaming, I don't think Spark Structured Streaming was very close, although Spark 2.3.0 might be sufficient for your use cases, and I use spark structured streaming only when I have to deploy spark pipeline/ml. But, besides the features I listed, I think Flink is just better written for streaming. For example, defining checkpoints and state management in Flink is so much more expressive and easier to write (and perhaps more performant from what I saw at Flink Forward). Flink Async is amazing!!!! The Flink table/sql API is not as nice as Spark SQL for batch right now, but it is getting there. Flink also seems to perform better than Spark Structured Streaming when you turn on object reusability, at least according to Alibaba and Netflix. And then, the features I listed are amazing. Evictors and triggers are super helpful. Process Functions let me write custom aggregations without a lot of mental overhead. Time Characteristics are really nice in Flink. Flink has a lot of nice connectors and stuff written already (although Parquet files are troublesome though if they come from Spark because of the way that Spark handles it read and write support classes). For example, you have to write your own kinesis connector in spark structured streaming (or use the one put out by databricks on their blogs). Sorry if this seems all over the place. I just posted what came to mind. And I should add, Databricks is doing a lot of work to make Spark Structured Streaming really comparable, so who knows what stuff looks like in 2-3 years. Like I said above, I use Spark Structured Streaming (2.3.0+) for some very specific ML stuff that my Flink environment cannot handle nicely yet. My only complaint is that there isn't a lot of community written code out there, but the Data Artisans team is super active, helpful, and nice.
- dvlsg 8y agoAny interest in pulsar support?