3 ms·
Which are the limitations shared by both a single node and a distributed engine? I am confused... There are ways to achieve fault-tolerance, even for a central
by grt17 8y ago
Which are the limitations shared by both a single node and a distributed engine? I am confused...
There are ways to achieve fault-tolerance, even for a centralised system (e.g, by maintaining an active/passive replica). In cases like this, where there is such a great gap between the performance of current systems, you can always "waste" another node(s) for fault-tolerance and still operate with less cost, if you want.
I agree that some types of computations (e.g., multiple/distributed sources) may not be benefited by a centralised approach and I definitely don't claim that this is a solution for everything. However, the point is to criticise the design choices that we make for a streaming system.
Streaming support, even for popular systems today, is something like an extension (sometimes a hack) on top of the core of the system, hidden beneath multiple layers of abstraction. In addition, modern systems try to do many things at the same time (support AI, batch & stream processing, connectors to publish-subscribe systems, multiple wire protocols) and they end up doing most of them poorly.
- cangencer 8y agoI meant that products like Esper, StreamBase, InfoSphere have been around for a long time, which have a _very_ rich set of features [1], and are mostly designed around single process usage. Lot of the type of queries they support are not possible to implement in a performant way in a distributed system. Though nowadays Esper claim to have horizontal scalability - it was originally designed as a single threaded system. They do also have a passive/active type solution as you mentioned. Stream processing frameworks originally evolved to offer "big scale" through data partitioning compared to the traditional CEP systems. But CEP engines have been able to deal with windowing and similar concepts since many years ago - the main difference of the stream processing frameworks _is_ the distribution and scalability aspect. My point was that the systems linked in the original article seem to match closely to the limitations of what distributed stream processing frameworks are able to do, but only run on a single node. [1] http://www.espertech.com/esper/ http://www.espertech.com/esper/
- scott_s 8y agoI assume that by "InfoSphere" you mean what used to be called IBM InfoSphere Streams, but what is now just IBM Streams. I do research and development on IBM Streams (see a sibling comment), and I can say with certainty that it was designed as a distributed, parallel stream processing system from the start. It is a more general computing platform than CEP engines - in fact, we actually implemented a CEP engine as a Streams operator. Product documentation for the operator: https://www.ibm.com/support/knowledgecenter/SSCRJU_4.2.0/com.ibm.streams.toolkits.doc/spldoc/dita/tk$com.ibm.streams.cep/op$com.ibm.streams.cep$MatchRegex.html https://www.ibm.com/support/knowledgecenter/SSCRJU_4.2.0/com... Academic paper: Partition and Compose: Parallel Complex Event Processing. Martin Hirzel. DEBS 2012. http://hirzels.com/martin/papers/debs12-cep.pdf http://hirzels.com/martin/papers/debs12-cep.pdf