5 ms·
Java FFM zero-copy transport using io_uring
- jeffreygoesto 10mo ago27us roundtrip is not really state of the art for zero copy IPC, about 1us would be. What is causing this overhead?
- rohanray 10mo agoIt's not a local IPC exactly. The roundtrip benchmark stat is for a TCP server-client ping/pong call using a 2 KB payload; TCP is although on local loopback (127.0.0.1). Source: https://github.com/mvp-express/myra-transport/blob/main/benchmarks/src/jmh/java/express/mvp/myra/transport/benchmark/RealWorldPayloadBenchmark.java https://github.com/mvp-express/myra-transport/blob/main/benc...
- jstimpfle 10mo agoAsking for those who, like me, haven't yet taken the time to find technical information on that webpage: What exactly does that roundtrip latency number measure (especially your 1us)? Does zero copy imply mapping pages between processes? Is there an async kernel component involved (like I would infer from "io_uring") or just two user space processes mapping pages?
- foltik 10mo ago27us and 1us are both an eternity and definitely not SOTA for IPC. The fastest possible way to do IPC is with a shared memory resident SPSC queue. The actual (one-way cross-core) latency on modern CPUs varies by quite a lot [0], but a good rule of thumb is 100ns + 0.1ns per byte. This measures the time for core A to write one or more cache lines to a shared memory region, and core B to read them. The latency is determined by the time it takes for the cache coherence protocol to transfer the cache lines between cores, which shows up as a number of L3 cache misses. Interestingly, at the hardware level, in-process vs inter-process is irrelevant. What matters is the physical location of the cores which are communicating. This repo has some great visualizations and latency numbers for many different CPUs, as well as a benchmark you can run yourself: [0] https://github.com/nviennot/core-to-core-latency https://github.com/nviennot/core-to-core-latency
- jstimpfle 10mo agoI was really asking what "IPC" means in this context. If you can just share a mapping, yes it's going to be quite fast. If you need to wait for approval to come back, it's going to take more time. If you can't share a memory segment, even more time.
- foltik 10mo agoNo idea what this vibe code is doing, but two processes on the same machine can always share a mapping, though maybe your PL of choice is incapable. There aren’t many libraries that make it easy either. If it’s not two processes on the same machine I wouldn’t really call it IPC. Of course a round trip will take more time, but it’s not meaningfully different from two one-way transfers. You can just multiply the numbers I said by two. Generally it’s better to organize a system as a pipeline if you can though, rather than ping ponging cache lines back and forth doing a bunch of RPC.
- znpy 10mo agoIt may or may not be good, depending on a number of fact. I did read the original linux zerocopy papers from google for example, and at the time (when using tcp) the juice was worth the squeeze when payload was larger than than 10 kilobytes (or 20? Don’t remember right now and i’m on mobile). Also a common technique is batching, so you amortise the round-trip time (this used to be the cost of sendmmsg/recvmmsg) over, say, 10 payloads. So yeah that number alone can mean a lot or it can mean very little. In my experience people that are doing low latency stuff already built their own thing around msg_zerocopy, io_uring and stuff :)
- hinkley 10mo agoio_uring is a tool for maximizing throughput not minimizing latency. So the correct measure is transactions per millisecond not milliseconds per transaction. Little’s Law applies when the task monopolizes the time of the worker. When it is alternating between IO and compute, it can be off by a factor of two or more. And when it’s only considering IO, things get more muddled still.
- znpy 10mo ago> io_uring is a tool for maximizing throughput not minimizing latency. some features are explicitly designed to minimize latency. I'm thinking of the IORING_SETUP_IOPOLL and IORING_SETUP_SQPOLL flags for io_uring_setup . I'm not making that up, the manpage says that: https://manpages.debian.org/unstable/liburing-dev/io_uring_setup.2.en.html https://manpages.debian.org/unstable/liburing-dev/io_uring_s...
- deleted 10mo ago[deleted]
- blibble 10mo agoindeed, you can get a packet from one box to another in 1-2us
- foobar10000 10mo ago[dead]
- steeve 10mo agowith io_uring? How? I tried everything in the book
- deleted 10mo ago[deleted]
- rohanray 10mo agoIt's not a local IPC exactly. The roundtrip benchmark stat is for a TCP server-client ping/pong call using a 2 KB payload; TCP is although on local loopback (127.0.0.1). The payload is encoded using myra-codec FFM MemorySegment directly into a pre-registered buffer in io_uring SQE on the server. Similarly, on the client side CQE writes encoded payload directly into a client provided MemorySegment. The whole process saves a few SYSCALLs. Also, the above process is zero copy. Source: https://github.com/mvp-express/myra-transport/blob/main/benchmarks/src/jmh/java/express/mvp/myra/transport/benchmark/RealWorldPayloadBenchmark.java https://github.com/mvp-express/myra-transport/blob/main/benc... P.S.: I had posted this as a reply to jeffrey but not able to see it. Hence, reposting as a direct reply to the main post for visibility as well. Disclaimer: I am the author of https://mvp.express https://mvp.express. I would love feedback, critical suggestions/advise. Thanks -RR
- refulgentis 10mo agoPretty much what NateB said* - but that might leave you at "what's wrong with that? that's how I could get it done" There's WAY too much content, way too many names and stuff that feels subtly off. I'm 37, been on this site for 16 years. I'm assuming target audience here is enterprise Java developers, which isn't my home, so I'm sure I'm missing some stuff is idiomatic in that culture. But the vast, vast amount of things that are completely unfamiliar tells me something else is going on and it's not good. Like I bet this is f'ing cool, otherwise you wouldn't put in the effort to share it. But you're better off having something super brief** in a GitHub README than a pseudo-marketing site that's straining to fit a cool technical thing into the wrong template. * https://news.ycombinator.com/item?id=46255661 https://news.ycombinator.com/item?id=46255661 ** what you wrote is great! "The payload is encoded using myra-codec FFM MemorySegment directly into a pre-registered buffer in io_uring SQE on the server. Similarly, on the client side CQE writes encoded payload directly into a client provided MemorySegment. The whole process saves a few SYSCALLs. Also, the above process is zero copy." -- then the site looks like it wants to sell N different products and confusing flowcharts, but really, you're just geeked out and did something cool and want to share the technical details. So it's designed for the wrong audience.
- owl_might 10mo ago
- nateb2022 10mo agoThis looks like most of it was vibecoded. Unnecessary comments like: clientChannel.configureBlocking(false); // Non-blocking client can be found throughout the source, and the project's landing page is a good example of typical SOTA models' outputs when asked for a frontend landing page.
- szundi 10mo agoWhat really matters though is the quality of the human review.
- krisgenre 10mo agoOkay, but is that a bad thing?
- sgammon 10mo agoIf the author doesn't understand their own code, I probably won't
- another_twist 10mo agoVibe coding doesnt mean the author doesnt understand their code. Its likely that they don't want carpal tunnel from typing out trivial code and hence offload that labor to a machine.
- TheGuyWhoCodes 10mo agoIn my opinion adding kryo in the benchmark is somewhat disingenuous as it does not require a message schema definition while MyraCodec/SBE/FlatBuffers do. The only thing that says is schemeless and is zero copy is Apache Fory which is missing from the benchmark.
- rohanray 10mo agoI had added Kryo since that seems to be the fastest Java serialization library which does not use sun.misc.unsafe. Thanks for sharing Apache Fory! Will try to add that to the benchmark as well.
- DarkmSparks 10mo agoMost of it seems to be 404ing now
- rohanray 10mo agoOh! That shouldn't be the case :( Please let me know if you are still facing 404. I just checked and no alerts from my monitoring yet. Thanks for letting know though!
- DarkmSparks 10mo agoE.g. https://www.mvp.express/docs/v0.1.0/examples/kvstore https://www.mvp.express/docs/v0.1.0/examples/kvstore Under examples
- exabrial 10mo agoImpressive. I'm sure the numbers will continue to improve as both the FFM and this project mature. Java Native databases or KVP stores would be good usage targets IMHO
- rohanray 10mo agoI have been planning on JIA Cache as a distributed caching system built with off heap memory DS & Flyweight for readers to achieve zero copy. I think a KV store will come out as a byproduct while developing JIA Cache