8 ms·
Jepsen: Bufstream 0.1
- philprx 2y agoI didn't find the GitHub project for bufstream... Any clue?
- aphyr 2y agoAck, pardon me. That should be fixed now!
- xiasongh 2y agoFound this on their website https://github.com/bufbuild/buf https://github.com/bufbuild/buf
- bpicolo 2y agoThey seem to have pivoted from protobuf tools to kafka alternatives. I don't think bufstream is OSS (yet). Or at least, they have very much de-emphasized their original offering on their site.
- perezd 2y agoNope! We're still heavily investing in scaling Protobuf. In fact, our data quality guarantees built into Bufstream are powered by Protobuf! This is simply an extension of what we do...Connect RPC, Buf CLI, etc. Don't read too much into the website :)
- bpicolo 2y agoGood to know. Good proto tooling is still high value :)
- mdaniel 2y agoI don't think bufstream itself is open source but there's https://github.com/bufbuild/bufstream-demo https://github.com/bufbuild/bufstream-demo which may be close to what you want (but is also unlicensed, bizarrely)
- perezd 2y agoThat's correct. Bufstream is not open source, but we do have a demo that you can try. I've asked the team to include a proper LICENSE file as well. Thanks for catching that!
- pmdulaney 2y agoWhat is this software used for? Instrumentation? Black boxes?
- smw 2y agoIt's a kafka clone. Kafka is a durable queue, mostly.
- deleted 2y ago[deleted]
- sureglymop 2y agoWhat is a durable queue and why is it needed (instead of a traditional relational db)?
- cyberax 2y agoRDBs suck for many-to-many high-availability messaging.
- toast0 2y agoTraditional relational dbs can be good at durability (if properly configured), but it a queue might be designed differently than a table. You might want to be more specific about how messages are assigned to clients than what would be convenient in a relational database.
- rad_gruchalski 2y agoKafka is an append-only log, not a queue.
- zbentley 2y agoAssuming successful implementation of KIP-932, it will soon be a queue as well. https://cwiki.apache.org/confluence/display/KAFKA/KIP-932%3A+Queues+for+Kafka https://cwiki.apache.org/confluence/display/KAFKA/KIP-932%3A...
- refset 2y ago> The Kafka transaction protocol is fundamentally broken and must be revised. Ouch. Great investigation work and write-up, as ever!
- c2xlZXB5Cg1 2y agoNot to be confused with https://www.warpstream.com/ https://www.warpstream.com/
- perezd 2y agoCorrect. WarpStream doesn't even support transactions.
- jcgrillo 2y agoNeither does any other Kafka protocol implementation, evidently ;)
- _8mlz 2y agoZing! Can’t lose if you don’t play ;)
- jcgrillo 2y agoThis cuts both ways, choosing to not implement flawed portions of the spec could be seen as a good thing. I've always been a bit suspicious of the value of "bug for bug compatibility". You don't actually need transactions in Kafka in normal operation IME. I've never tried to use "streams" before and have never encountered a case where I thought they were a good trade. Better to implement that kind of stuff in a way I can control.
- williamdclt 2y agoI'm very surprised by this: > [with the default enable.auto.commit=true] Kafka consumers may automatically mark offsets as committed, regardless of whether they have actually been processed by the application. This means that a consumer can poll a series of records, mark them as committed, then crash—effectively causing those records to be lost That's never been my understanding of auto-commit, that would be a crazy default wouldn't it? The docs say this: > when auto-commit is enabled, every time the poll method is called and data is fetched, the consumer is ready to automatically commit the offsets of messages that have been returned by the poll. If the processing of these messages is not completed before the next auto-commit interval, there’s a risk of losing the message’s progress if the consumer crashes or is otherwise restarted. In this case, when the consumer restarts, it will begin consuming from the last committed offset. When this happens, the last committed position can be as old as the auto-commit interval. Any messages that have arrived since the last commit are read again. If you want to reduce the window for duplicates, you can reduce the auto-commit interval I don't find it amazingly clear, but overall my understanding from this is that offsets are committed _only_ if the processing finishes. Tuning the auto-commit interval helps with duplicate processing, not with lost messages, as you'd expect for at-least-once processing.
- aphyr 2y agoIt is a little surprising, and I agree, the docs here are not doing a particularly good job of explaining it. It might help to ask: if you don't explicitly commit, how does Kafka know when you've processed the messages it gave you? It doesn't! It assumes any message it hands you is instantaneously processed. Auto-commit is a bit like handing someone an ice cream cone, then immediately walking away and assuming they ate it. Sometimes people drop their ice cream immediately after you hand it to them, and never get a bite.
- dangoodmanUT 2y agoThis, it has no idea that you processed the message. It assumes processing is successful by default which is cosmically stupid.
- williamdclt 2y ago
- bobnamob 2y ago> We would like to combine Jepsen’s workload generation and history checking with Antithesis’ deterministic and replayable environment to make our tests more reproducible. For those unaware, Antithesis was founded by some of the folks who worked on FoundationDB - see https://youtu.be/4fFDFbi3toc?si=wY_mrD63fH2osiU- https://youtu.be/4fFDFbi3toc?si=wY_mrD63fH2osiU- for some of their handiwork. A Jepsen + Antithesis team up is something the world needs right now, specifically on the back of the Horizon Post Office scandal. Thanks for all your work highlighting the importance of db safety Aphyr
- bobnamob 2y agoFurthermore, I'm aware of multiple banks currently using Kafka. One would hope that they're not using it in their core banking system given Kyle's findings Maybe they'd be interested in funding a Jepsen ~attack~ experiment on Kafka
- kevstev 2y agoAs someone who was very deep into Kafka in the not too distant past, I am surprised I have no idea what you are referring to- can you enlighten me?
- diggan 2y agoRead the "Future Work" section in the bottom of the post for the gist, and also section 5.3.
- kevstev 2y agoI see. I never trusted transactions and advised our app teams to not rely on them, at least without outside verification of them. The situation is actually far worse with any client relying on librdkafka. Some of this has been fixed, but my company found at least a half dozen bugs/uncompleted features in librdkafka, mostly around retryable errors that were sometimes ignored, sometimes caused exceptions, and other times just straight hung clients. Despite our company leaning heavily on Confluent to force librdkafka to get to parity with the Java client, it was always behind, and in general we started adopting a stance of not implementing any business critical functions on any feature implemented in the past year or major release.
- diggan 2y ago> While investigating issues like KAFKA-17754, we also encountered unseen writes in Kafka. Owing to time constraints we have not investigated this behavior, but unseen writes could be a sign of hanging transactions, stuck consumers, or even data loss. We are curious whether a delayed Produce message could slide into a future transaction, violating transactional guarantees. We also suspect that the Kafka Java Client may reuse a sequence number when a request times out, causing writes to be acknowledged but silently discarded. More Kafka testing is warranted. Seems like Jepsen should do another Kafka deep-dive. Last time was in 2013 (https://aphyr.com/posts/293-call-me-maybe-kafka https://aphyr.com/posts/293-call-me-maybe-kafka, Kafka version 0.8 beta) and seems like they're on the verge of discovering a lot of issues in Kafka itself. Things like "causing writes to be acknowledged but silently discarded" sounds very scary.
- aphyr 2y agoI would love to do a Kafka analysis. :-)
- jwr 2y agoI'm still hoping Apple (or Snowflake) will pay you to do an analysis of FoundationDB…
- tptacek 2y agoI do too, but doesn't FDB already do a lot of the same kind of testing?
- kasey_junk 2y agoThey are famous for doing simulation testing. https://antithesis.com/ https://antithesis.com/ Have recently brought to market a simulation testing product.
- SahAssar 2y agoI think they do similar testing, and therefore it might be even more interesting to read what Kyle thinks of their different approaches to it.
- didip 2y agoHas Kyle reviewed NATS Jetstream? I wonder what he thinks of it.
- aphyr 2y agoI have not yet, though you're not the first to ask. Some folks have suggested it might be... how do you say... fun? :-)
- speedgoose 2y agoIf you are looking for fun targets, may I suggest KubeMQ too? Its author claims that it’s better than Kafka, Redis and RabbitMQ. It’s also "kubernetes native" but the open source version refuses to start if it detects kubernetes.
- SahAssar 2y ago> It’s also "kubernetes native" but the open source version refuses to start if it detects kubernetes. I thought you were kidding, but this is crazy. https://github.com/kubemq-io/kubemq-community/issues/32 https://github.com/kubemq-io/kubemq-community/issues/32 And it seems like you cannot even see pricing without signing up or contacting their sales: https://kubemq.io/product-pricing/ https://kubemq.io/product-pricing/
- mathfailure 2y agoThis is just pure gold of an anecdote :-)
- Bnjoroge 2y agohas warpstream been reviewed?
- aphyr 2y agoNope. You'll find a full list of analyses here: https://jepsen.io/analyses https://jepsen.io/analyses
- Kwpolska 2y agoI’m looking at the product page [0] and wondering how those two statements are compatible: > Bufstream runs fully within your AWS or GCP VPC, giving you complete control over your data, metadata, and uptime. Unlike the alternatives, Bufstream never phones home. > Bufstream pricing is simple: just $0.002 per uncompressed GiB written (about $2 per TiB). We don't charge any per-core, per-agent, or per-call fees. Surely they wouldn’t run their entire business on the honor system? [0] https://buf.build/product/bufstream https://buf.build/product/bufstream
- c0balt 2y agoBased on the introduction > As of October 2024, Bufstream was deployed only with select customers. my assumption would be an honor system might be doable. They are exposing themselves to risk of abuse of course but it might be a worthy trade off for getting certain clients on board.
- perezd 2y agoThat's correct. We hop onto Zoom calls with our customers on an agreed cadence, and they share a billing report with us to confirm usage/metering. For enterprise customers specifically, it works great. They don't want to violate contracts, and it also gives us a natural check-in point to ensure things are going smoothly with their deployment. When we say fully air-gapped, we mean it!
- mathfailure 2y agoA program is either opensourced or not. When its sources aren't available - one should never trust "it doesn't phone home" claims.
- kentonv 2y agoIf a company makes unambiguous claims in their advertising which turn out to be false, they will get sued and maybe even fined by regulators. Much of the world does in fact operate on this kind of trust.
- 2y ago
- zbentley 2y agoErratum: > Transactions may observe none, part, or all Should, I think, read: > Consumets may observe none, part, or all
- aphyr 2y agoBoth are true, but we use "transactions" for clarity, since the semantics of consumers outside transactions is even murkier. Every read in this workload takes place in the context of a transaction, and goes through the transactional offset commit path.
- zbentley 2y agoAh, got it; I was assuming that “transactions” was referring to the transactions mentioned as the subject of the previous sentence, not the transactions active in consumers observing those. My mistake!
- bayareacommie 2y ago[dead]
- kiitos 2y agoGreat work as always. After reading thru the relevant blog posts and docs, my understanding is that Kafka defines "exactly-once delivery" as a property of what they call a "read-process-write operation", where workers read-from topic 1, and write-to topic 2, where both topics are in the same logical Kafka system. Is that correct? If so, isn't that better described as a transaction?
- aphyr 2y agoKafka actually does call these transactions! However (and this is a loooong discussion I can't really dig into right now) there's sort of two ways to look at "exactly once". One is in the sense that DB transactions are "exactly once"; a transaction's effects shouldn't be duplicated or lost. But in another sense "exactly once" is a sort of dataflow graph property that relates messages across topic-partitions. That's a little more akin to ACID "consistency". You can use transactions to get to that dataflow property, in the same sort of way that Serializable transaction systems guarantee certain kinds of domain-level consistency. For example, Serializability guarantees that any invariant preserved by a set of transactions, considered purely in isolation, is also preserved by concurrent histories of those transactions. I think you can argue Kafka intends to reach "exactly-once semantics" through transactions in that way.