6 ms·
Gazette: Cloud-native millisecond-latency streaming
- danthelion 2y agoGazette is at the core of Estuary Flow (https://estuary.dev https://estuary.dev), a real-time data platform. Unlike Kafka, Gazette’s architecture is simpler to reason about and operate. It plays well with k8s and is backed by S3 (or any object storage).
- Onavo 2y agoInteresting, are there any open source alternatives to tinybird? https://www.tinybird.co/ https://www.tinybird.co/
- vpol 2y agoI've wrote one, but it's not public/production-ready yet. Built on top of clickhouse as well. Basically pipe is just a collection of WITH statements with some template processing.
- mrbluecoat 2y agoNot an exact match, but https://github.com/soketi/soketi https://github.com/soketi/soketi might work for your needs (API-compatible with https://pusher.com https://pusher.com )
- jauntywundrkind 2y agoI feel a bit paralyzed by Fear Of Missing Io_Uring. There's so much awesome streaming stuff about (RisingWave, Materialize, NATS, DataFusion, Velox, neat upstarts like Iggy, many more), but it all feels built on slower legacy system libraries. It's not heavily used yet, but Rust has a bunch of fairly high visibility efforts. Situation sort of feels similar with http3, where the problem is figuring out what to pick. https://github.com/tokio-rs/tokio-uring https://github.com/tokio-rs/tokio-uring https://github.com/bytedance/monoio https://github.com/bytedance/monoio https://github.com/DataDog/glommio https://github.com/DataDog/glommio Alas libuv (powering Node.js) shipped io_uring but disabled it latter. Seems to have significantly worn out the original author on the topic to boot. https://github.com/libuv/libuv/pull/4421#issuecomment-2225860128 https://github.com/libuv/libuv/pull/4421#issuecomment-222586...
- hnav 2y agoio_uring is a low level abstraction and is generally a wash against epoll. Really won't make a difference for these kinds of applications, especially not for client nodes.
- 10000truths 2y agoio_uring allows for async reads and writes to disk without forcing a thread pool or direct I/O. That alone makes it much more scalable for workloads that touch both the network and disk.
- immibis 2y agoDoesn't the kernel use a thread pool to process the requests in the ring, because the kernel is still designed around blocking disk I/O?
- JoshTriplett 2y agoNo, most operations in the ring directly work asynchronously. The thread mechanism only exists as a fallback for combinations of operations and system configurations (e.g. filesystems) that don't support asynchronous operation.
- mightyham 2y agoI don't know anything about the internals of io_uring and am genuinely curious how it works. Saying it "directly works asynchronously" doesn't mean anything though. When circular buffer requests are processed what thread is processing the request, how is that thread managed, and how does it manage blocking/unblocking when communicating with the storage device?
- JoshTriplett 2y agoInternally, many parts of the Linux kernel operate asynchronously: they queue up a request with some subsystem (e.g. a hardware device), and get an event delivered when the request is completed. In such cases, io_uring can enqueue such a request, and complete it when receiving the event, without needing to use a thread to block waiting for it. See, for instance, https://lpc.events/event/11/contributions/901/attachments/786/1661/io_uring-BPF.pdf https://lpc.events/event/11/contributions/901/attachments/78... slide 5 (though more has happened since then). io_uring will first see if it has everything needed to do the operation immediately, if not it'll queue a request in some cases (e.g. direct I/O, or buffered I/O in some cases). The thread pool is the last fallback, which always works if nothing else does. https://lwn.net/Articles/821274/ https://lwn.net/Articles/821274/ talks about making async buffered reads work, for instance.
- abrookewood 2y agoMore details viewable here: https://gazette.readthedocs.io/en/latest/ https://gazette.readthedocs.io/en/latest/
- xyst 2y agoWhere can I get nanosecond latency streaming?
- immibis 2y agoA wire
- actionfromafar 2y agoA rather short wire. Short enough for Grace Hopper to give away to students during lectures.
- ramon156 2y agoThe wire
- Groxx 2y agoThe screen in front of your face. That'll give you like 3ns latency.
- modeless 2y agoIt is frequently faster to send an IP packet to another continent than to change a pixel on the screen.
- CyberDildonics 2y agoNo it isn't. John Carmack found a tv ten years ago that had 200-300ms of latency due to all its post processing and wrote an essay about it. That doesn't mean that it is "frequently" faster to send packets to other continents than change pixels on screens. It doesn't even apply to modern tvs set up for latency, let alone computer monitors.
- modeless 2y agoYes, it really is. The problem is it takes a lot more than one frame for most modern software to change a pixel on the screen. I'm sitting in Hawaii on wifi right now and the first random US mainland server I pinged responded in 120ms, which means sending only took 60ms. Now say you're running a 30 Hz game with 2 frames of input lag, and you've already lost before even considering the input lag of the monitor itself. There are just so many ways to accidentally get many frames of input lag. OS window compositors generally add a whole frame of input lag globally to every windowed app. Anything running in a browser has a second compositor in between it and the display that can add more frames. GPU APIs typically buffer one or two frames by default. And all of that is on top of whatever the app itself does, and whatever the monitor does (and whatever the input device does if you want to count that too).
- oatmeal_croc 2y agoWhat's the use case for millisecond-latency streaming? HFT? Remotely driving heavy machinery? Anything else?
- freeqaz 2y agoCollaborative systems come to mind. If you edit a document and want to subscribe to changes from other nodes it is valuable to have very low latency.
- mbrock 2y agoI think it's less about guaranteed 1ms real time transactions and more about, like, it's just fast enough that you most likely don't have to worry about it introducing perceptible lag? I'm working on a streaming audio thing and keeping latency low is a priority. I actually think I'll try Gazette, I just saw it now and it was one of those moments where it's like wait I go to Hacker News to waste time but this is quite exactly what I've been wanting in so many ways. I'll use it for Ogg/Opus media streams, transcription results, chat events, LLM inferences... I really like the byte-indexed append-only blob paradigm backed by object storage. It feels kind of like Unix as a distributed streaming system. Other streaming data gadgets like Kafka always feel a bit uncomfortable and annoying to me with their idiosyncratic record formats and topic hierarchies and whatnot... I always wanted something more low level and obvious...
- rswail 2y ago> wait I go to Hacker News to waste time but this is quite exactly what I've been wanting in so many ways. This has happened so many times for me that I don't consider the time "wasted". I try to make sure I separate the a) "this is interesting personally", and b) "this is interesting professionally" threads and have a bunch of open tabs for a) that I can "read later". But the items in (b) I read "now" and consider that to be work, not pleasure.
- kcb 2y agoA millisecond is an eternity in HFT.
- amluto 2y agoFrom reading the docs, this has an IMO surprising design decision: the “journal” is a stream of bytes, where each append (of a byte string) is atomic and occurs in a global order. The bytes are grouped into fragments, and no write spans a fragment boundary. This seems sort of okay if writes are self-delimiting and never corrupt, and synchronization can always be recovered at a fragment boundary. I suppose it’s neat that one can write JSONL and get actual JSONL in the blobs. But this seems quite brittle if multiple writers write to one journal and one malfunctions (aside from possibly failing to write a delimiter, there’s no way to tell who wrote a record, and using only a single writer per journal seems to defeat the purpose). And getting, say, Parquet output doesn’t seem like it will happen in any sensible way.
- jgraettinger1 2y ago:wave: Hi, I'm the creator of Gazette. > But this seems quite brittle if multiple writers write to one journal and one malfunctions (aside from possibly failing to write a delimiter, there’s no way to tell who wrote a record, and using only a single writer per journal seems to defeat the purpose). Yes, writers are responsible for only ever writing complete delimited blocks of messages, in whatever framing the application wants to use. Gazette promises to provide a consistent total order over a bunch of raced writes, and to roll back broken writes (partial content and then a connection reset, for example), and checksum, and a host of other things. There's also a low-level "registers" concept which can be used to cooperatively fence a capability to write to a journal, off from other writers. But garbage in => garbage out, and if an application correctly writes bad data, then you'll have bad data in your journal. This is no different from any other file format under the sun. > there’s no way to tell who wrote a record To address this comment specifically: while brokers are byte-oriented, applications and consumers are typically message oriented, and the responsibility for carrying metadata like "who wrote this message?" shifts to the application's chosen data representation instead of being a core broker concern. Gazette has a consumer framework that layers atop the broker, and it uses UUIDs which carry producer and sequencing metadata in order to provide exactly-once message semantics atop an at-least-once byte stream: https://gazette.readthedocs.io/en/latest/architecture-exactly-once.html https://gazette.readthedocs.io/en/latest/architecture-exactl...
- mrbluecoat 2y ago> the broker is pushing new content to us over a singled long-lived HTTP response Any plans to support websocket? https://gazette.readthedocs.io/en/latest/brokers-tutorial-introduction.html#streaming-reads https://gazette.readthedocs.io/en/latest/brokers-tutorial-in...