4 ms·
Martin Kleppmann seems to point out technologies for problems of similar patterns already exist - https://twitter.com/martinkl/status/1039938408393662465 http
by fullmetaleng 8y ago
Martin Kleppmann seems to point out technologies for problems
of similar patterns already exist - https://twitter.com/martinkl/status/1039938408393662465 https://twitter.com/martinkl/status/1039938408393662465
- tinco 8y agoThose are streaming/pubsub services though, this actually claims to be a store. I feel that's an important difference. Do people just point their system journal at Kafka and wait for something to break? At my previous job we built something similar to this out of rabbitmq and mongodb. I always wondered what the other big log companies used. Mongodb seemed like a pretty good fit, but a pure append only database might be even better. Trimming performance in MongoDB was subpar so we worked around it by creating a new collection for each day, trimming became a simple operation of dropping a collection at the end of each day.
- manigandham 8y agoAll of those are similar systems and have persistence. I'm not sure what distinction "streaming" makes but they also all support multiple publishers and subscribers. Some only use local storage on the nodes while others can tier out to cold storage like S3. MongoDB is a full OLTP document store so it won't match the write throughput and pubsub features of these focused systems. RabbitMQ on the other hand has performance limits but is meant for complex service-bus style routing and RPC uses, but I recommend using NATS for that now.
- EdwardDiego 8y ago> Those are streaming/pubsub services though, this actually claims to be a store. I feel that's an important difference. > Do people just point their system journal at Kafka and wait for something to break? Kafka can be used as a data store if you like, so long as you're happy with the data management and access patterns it gives you - it is, after all, optimised for large sequential reads. LogDevice looks to be very similar for most use cases to Kafka, hell, they even use RocksDB, which is used by stateful operations in Kafka Streaming, and of course, Zookeeper. Where it differs is that it looks like it was designed for you to be able to work against a single "cluster" that could well be running across multiple data-centres. Which is very much a Facebook problem to solve. So yeah, Kafka was a distributed log built for LinkedIn size problems, LogDevice is a distributed log built for Facebook sized problems. Most of us don't have Facebook sized problems.
- Serow225 8y agoWhat's a good distributed log for 10-dev sized companies? :)
- jacobr1 8y agoAWS Kinesis + s3
- Serow225 8y agoThanks! I'm guessing you're referring to Kinesis Streams? Is there an OOB solution to persist the records past the default 168hours, or is this something that you have to build out yourself following some pattern?
- mhotchen 8y agoYou can configure it to output to S3 and the mechanism for that is easy to configure and has fault tolerance.
- simonrobb 8y agoOKLog, Humio, and Splunk are all worth checking out.
- atombender 8y agoOKLog has been abandoned by the author (the project is now read only on GitHub). Humio is not self-hosted or open source, so not really a fair comparison. It also seems targeted towards operational logs, i.e. system logging, traffic logging, auditing. Not things like data pipelines. Kafka and friends can be used for that kind of log, but they are more like databases; they use the term "log" in the sense of sequential and append-only. Same goes for Splunk, which does have a self-hosted version, but is extremely expensive, last I checked. The SaaS version is also extremely expensive.
- 8y ago
- deleted 8y ago[deleted]