2 ms·
It depends on what you're doing. In C you're going to be hampered by, again, the ecosystem. There is nothing comparable in C as performant as all of the stream
by RhodesianHunter 3y ago
It depends on what you're doing. In C you're going to be hampered by, again, the ecosystem.
There is nothing comparable in C as performant as all of the streaming tooling built up around Kafka for example.
So your can benchmark percentages all you want, but there are a lot of things you just can't do in C that you can on the JVM (without 100 developers and three years)
- morelisp 3y ago> There is nothing comparable in C as performant as all of the streaming tooling built up around Kafka for example. This is an ill-chosen example. librdkafka is as or more performant than the Java stuff. As a server, Redpanda blows Kafka out of the water (but this is rarely any bottleneck with brokers).
- RhodesianHunter 3y agoYes, if you take the time to hand roll a replacement like Redpanda did it will be faster (though to your point that seems fairly pointless given the constraints imposed by IO) I was referring more to the general ecosystem of Flink/Spark/KStreams/KTable etc.
- agallego 3y agothis seems truthy but isn't in practice. a lot of work, my perf optimization team @ redpanda does (yes we have a full team chasing tail latencies) is spent on CPU optimization, debouncing, amortizing costs, metadata lookups, hash tables, profilers, etc. so there is a lot of additional work after the IO layer which a decent async eventing thing can get you to get good perf.
- RhodesianHunter 3y agoWhat I mean to say is that in my experience the time spent round-tripping to Kafka is more than the time it takes Kafka to do whatever I'm asking it to do. So at least for my use-cases a faster Kafka would be of no benefit. Now if it means I can run fewer smaller brokers that's awesome.
- agallego 3y agodef. that should be the case, we have iot companies pushing us on a single pthread a a few megs of ram.
- morelisp 3y ago> Flink/Spark/KStreams/KTable etc. Then it's still a bit of an untruth since all of those use RocksDB for their high-performance storage layer, which isn't in Java. You could maybe (maybe!) argue for some ergonomics (though especially hard to argue for Spark!) but performance still goes plainly to C (/ C++).
- RhodesianHunter 3y agoWe're not talking about the systems you build on top of here. We're talking about the code. You write yourself. The fact that rocks is written in c is irrelevant to the fact that in order to use the ecosystem I'm describing you need to be on the JVM or your experience will be at best subpar.
- morelisp 3y agoSure, if you ignore offloading all the hard work to a C library, we can pretend any language is as fast as any other. It's a boring conversation then though - might as well use Python... Not to mention you still pay a real performance penalty for being unable to make direct use of RocksDB features, because Kafka is shoving a million "abstraction" and cache layers in there and you just want your your damned merge operator to apply. I don't agree Kafka off the JVM is subpar. Kafka Streams has a lot of sharp edges from "abstractions" that aren't, Spark is a godawful developer experience and tuning it for performance is black magic for 99% of data engineers, and Flink keeps dithering about how consistent/stable baseline features like queries will be. You can stand up very good apps for a huge set of common cases just as quickly in Go or C++, and not have to deal with Kafka Streams's funny opinions about threads/tasks/consumers.