3 ms·
This smells like a job for Apache Kafka [1], I've yet to use it personally but its feature set appears to hit the mark though it lacks SQL. The application desc
by hartror 12y ago
This smells like a job for Apache Kafka [1], I've yet to use it personally but its feature set appears to hit the mark though it lacks SQL. The application described sounds like it uses something similar to event sourcing [2] which people have used Kafka for successfully. If you're not familiar with Kafka there is a very good interview with Jun Rao [3] on se radio.
[1] http://kafka.apache.org/ http://kafka.apache.org/
[2] http://martinfowler.com/eaaDev/EventSourcing.html http://martinfowler.com/eaaDev/EventSourcing.html
[3] http://www.se-radio.net/2015/02/episode-219-apache-kafka-with-jun-rao/ http://www.se-radio.net/2015/02/episode-219-apache-kafka-wit...
- misframer 12y agoKafka is great, but it doesn't solve the problem the author describes since you can't specify an index. Data arrives in (timestamp, metric) order, but it needs to be indexed by (metric, timestamp).
- jat850 12y agoWould Apache Samza fold in here better, perhaps?
- misframer 12y agoI don't know much about Samza, but I don't think stream processors are what we (I work with the author) are looking for. We don't really have a lot of "stream processing" to do, and aggregate functions are usually computed on-demand. Also, still have to put results from the stream process somewhere, right? Back into Kafka is something people do, but we still need indexing capabilities. As the post mentions, we use MySQL for time series storage, but we also use Kafka in front as a durable log.
- jat850 12y agoYou're welcome to email me if you like - it's in my profile. Kafka and Samza are intended (generally speaking) to go hand in hand. Samza is a re-imagined datastore that Kafka can shuttle data into. I've been investigating Samza quite heavily specifically for time series data storage. I'd be happy to share thoughts.
- andrioni 12y agoWhile I'm not using Samza, Spark Streaming also works pretty nicely in this case, although it is not so focused on keeping state (it can, though, using the checkpointing system and the `updateStateByKey` transformation) and thus might not perfectly stable if you require to handle failures without reprocessing.
- otterley 12y agoKafka is a message bus (for data in transit), not a store for data at rest. Sure, you could abuse it as an append-only log-structured primary data store, but then again, when you have a hammer, everything looks like your thumb, I guess. Also, as another poster said, it has no indices, just an offset pointer.