3 ms·
Also might want to check out http://concord.io http://concord.io, it's a bit more work to set up, but it's much faster than most stream processing systems
by pastaking 10y ago
Also might want to check out http://concord.io http://concord.io, it's a bit more work to set up, but it's much faster than most stream processing systems
- reubano 10y agoHow does concord differ from the others? spark/storm/flink/etc...? Aside from being written in C that is.
- agallego 10y agoeng at concord here. Really cool API, you should port this to concord! =) i'd say major diff is dynamic topology. So during the pipeline execution you can add/remove workers for any stage. Also each stage/operator can be written in any programming language. Storm/Flink/SparkStreaming/etc... all have much higher level API's. We built the execution engine first, these great things (DSL, etc) should come soon. For example this API would be easy to support to execute on top (the pipe abstraction that is) Here is an example of a DSL we prototyped in a couple hours.
- agallego 10y agoerr. missing link: https://github.com/jjmalina/concord-python-dsl https://github.com/jjmalina/concord-python-dsl
- reubano 10y agoPretty neat! I'm guessing that concord isn't limited to just map/reduce... correct? I think building in integrations to other systems in the stream processing ecosystem is key. First on the list is to integrate with popular workflow schedulers so that you can design topologies which riko would then parse. Next up I think would be supporting custom sources/sinks such as Twitter, HDFS, RDMS, etc. What exactly would be involved in "porting" riko to concord and what would the advantages be for doing so?
- agallego 10y agoInteresting, will keep an eye out as you mature it, looks like an awesome DSL for us to fill the runtime. Advantages are: 1. mesos integration, with that comes containerization support, multi tenancy, QoS, proper pipeline supervision, etc 2. Scheduling of pipelines. i.e.: Schedule them on 100 computers. 3. the outputs of your DAG could be consumed by other systems immediately and even written in different programming languages. So your python DSL could be the source to the Scala DSL at some point. so language interop 4. Available KV storage 5. Tracing (Zipkin) - ala Google Dapper. 6. Fast networking - C++ backed runtime is 1 order of magnitude faster than the python one. What would be involved, is not much from what I can tell. Each concord 'operator' is like a networked function. so given a DAG, you could generate many operators internally, or literally write them to a file, i.e.: operator_one.py etc. The code generation or internal scheduling would be the glue that's needed. if you ever become interested, ping me! would love to collab alex@concord.io
- bimil 10y agomany differences from spark/storm/flink, but the most notable difference of concord is its dynamic topology model - deployment, scaling, and changes to your topologies all can be done in runtime, w/o restart of the full topology, which is required for spark/storm/flink.