3 ms·
I gave the paper a full read through last night. Essentially this is a language agnostic framework for building data processing systems that are highly-avalibl
by algorithmsRcool 8y ago
I gave the paper a full read through last night.
Essentially this is a language agnostic framework for building data processing systems that are highly-avalible, distributed, topologically static (no dynamic scaling), and features exactly once processing.
You define a message handler that will always produce the same output sequence given the same input sequence and the framework provides delivery, serialization, buffering, durability and transparent recovery. They even provide a nice way of wrapping non deterministic behavior so that you can seamlessly continue even if you fail in the middle of processing a message.
That being said, the really are sloppy with their performance numbers, the comparison to gRPC isn't really fair at all due to their dynamic batching. And the code examples in the paper have some really silly errors.
But the paper is still a great introduction to reliable stream processing and basic strategies for delivering exactly once delivery.
Also, there is some interesting code that the paper glosses over in this repo: https://github.com/Microsoft/CRA https://github.com/Microsoft/CRA
- rrnewton 8y agoThanks for the feedback on the draft. Regarding "topologically static", the system doesn't assume a fixed set of communicating endpoints (like MPI ranks). It will all you to dynamically add new participants to the network. Why is the gRPC comparison not fair? Shouldn't they do dynamic batching also? It can be done without unduly affecting latency. (I did my phd in stream processing studying this proposition, and Jonathan Goldstein and others have demonstrated the same thing in Trill.) In the case of AMBROSIA, our latency increase vs gRPC is not because of batching, but is because on waiting for the log to persist in georeplicated storage.
- huukhiem 8y agoHi, Did you have a chance to read about Google's Dataflow paper[1] and their new Streaming Engine[2]? From a layperson's perspective, it seems like they are tackling some of the same ideas (separation of state && computation, applying optimisation techniques used in functional world, etc). I'd be interested in learning where and how AMBROSIA differ! [1]: https://storage.googleapis.com/pub-tools-public-publication-data/pdf/43864.pdf https://storage.googleapis.com/pub-tools-public-publication-... [2]: https://cloud.google.com/blog/products/data-analytics/introducing-cloud-dataflows-new-streaming-engine https://cloud.google.com/blog/products/data-analytics/introd...