9 ms·
Leveraging Rust in our Java database
- LispSporks22 3y agoI needed to call Rust from Java a few years ago. The approach was build a straightforward .so and call into it via JNA. It seemed way less complicated
- amunra__ 3y agoThe rust-maven-plugin we wrote indeed also supports JNA for these simple use cases. https://github.com/questdb/rust-maven-plugin https://github.com/questdb/rust-maven-plugin Compared with JNA, JNI is indeed more complex, but it's faster and has more features. It also solves the problem of calling Java from Rust.
- amun_ra 3y ago[dead]
- wmiel 3y agoI'm not sure I understand why they're using java if they avoid GC? Doesn't sound like the best fit, especially the foreign memory in java isn't too pleasant to work with.
- MattPalmer1086 3y agoYeah, sounds like a lot of effort. The only thing they say that explains it is they end up with a single jar file, whose only dependency is the JRE. So I guess they get platform independence and easy installation.
- nu11ptr 3y agoKeep in mind the JNI libs are platform specific. That means available platforms are a function of what the JRE runs on AND they have built the shared lib for (and bundled in to the jar)
- _ZeD_ 3y agoyou know what fatjars are?
- jerrinot 3y agoQuestDB engineer here: We use jlink to create images for selected platforms. This means not even JRE is a dependency: You unpack a tarball and you are good to go. See: https://questdb.io/docs/get-started/binaries/ https://questdb.io/docs/get-started/binaries/
- MattPalmer1086 3y agoWell, technically it is a dependency in that you can't target platforms there isn't a JRE for. But I take your point about simplifying installation.
- jabradoodle 3y agoIt's not that uncommon of move, to choose jumping through hoops with Java over writing C. Might become less common now Rust is teaching the level it is.
- cmrdporcupine 3y agoHaving worked on writing DB internals in both Rust and in other languages, I can say that there's huge time-saving advantages to having something higher-level & garbage collected at the layer of the query parser/analyzer/compiler. The borrowing/ownership semantics can get really snaky when dealing with complicated expression trees, iteration patterns, etc. It's fairly hard to write ergonomic interfaces for more complicated iteration patterns in Rust while still respecting safety. That's actually fine and by design, and it's possible with a lot of effort and thought but this is not as much of a concern in e.g. Java. E.g. skim the discussion on this proposed "cursor" API for Rust's stdlib BTree: https://github.com/rust-lang/rust/issues/107540 https://github.com/rust-lang/rust/issues/107540 (And while Rust's enum-based algebraic types & pattern matching are nice, they're actually fairly limited when compared to what you can find in e.g. Scala or F#, Haskell, etc.) But I think there's also huge win in doing something like the pager/buffer pool/storage/data structure/indexes layer in Rust. For safety and efficiency reasons.
- ilikerashers 3y agoI remember reading that the founder was working in low-latency Java development with London investment banks for years. I guess it's what he knew. Also, Rust is a hard language to start a company with so I wouldn't be surprised if this is more of a product maturity thing.
- skippyboxedhero 3y agoSurely it is at least as hard to find people who know how to write Java without GC? Presumably you can't use Hotspot so you have to write your own VM too?
- nhourcard 3y agoFolks with a background in electronic trading (FX, hedge funds, trading firms etc) are familiar with zero-gc Java. London / NY / HK are a good pool of talent in that respect
- skippyboxedhero 3y agoYep, I ended up checking their code base. I know a little bit of Java but didn't realise that sun.misc.Unsafe existed, so it does actually look fairly straightforward to create code outside of GC (for anyone reading questdb/std/Unsafe.java seems to be where some allocations are handled). A pain I am sure, but way more manageable than I thought.
- Alupis 3y agoIt's not necessarily outside of GC, it's that they make great efforts to avoid GC. Such as instantiating all objects at startup and holding references to avoid GC, never using new keyword, avoiding objects in favor of primitives, avoid exceptions, etc. These systems often restart or do a full "stop the world" GC once per day. The system is quite different than what most are used to, especially during a trend towards increasingly Functional styles with immutability as a default, etc. Peter Lawrey[1] has some great posts/talks about his experiences in HFT. [1] https://github.com/peter-lawrey https://github.com/peter-lawrey
- usrusr 3y agoSeamless integration with parts that don't need to be GC-free comes to mind: they are not building an application, they are building a building block. And that building block can be used both in applications that do require the latency guarantees of GC-free as well as in applications that don't. Another class of applications would be ones that alternate between phases of unpredictable latency (like bootup or reconfiguration) and low-latency operation.
- cmrdporcupine 3y agoThey're not the only ones to have done this in this space. VoltDB (Michael Stonebreaker of Postgres [among other things] fame) did this -- low or no-GC style Java, effectively non-idiomatic Java, but taking advantage of the Java runtime in other ways. Others have done the same. And as others have pointed out, there's things outside the DB domain in high frequency trading and the like that have done this as well. There are advantages to Java: mature runtime, large talent pool out there, good tooling (still haven't seen anything as good as JMX for any other runtime). And if there's any language whose GC could be tuned to be "responsible", it'd be the JVM; there's been more GC R&D in the JVM than in any other runtime. I worked at RelationalAI (another DB vendor) for a bit, and their DB is all written in Julia, another garbage collected language... and the GC in Julia is what I'd characterize as ... immature... for that kind of application. I would have loved to have access to the JVM's GC there. Also this looks to be more of an analytical, column oriented, database. So I can imagine they're optimizing more for throughput than transactional latency. (I could be wrong, correct me, Quest folks...) And choice of Java likely has to do with when they began working on the project and what was out there at the time. It's the real world of software eng. We work with the tools and people we have because shipping a product on time and bringing in $$ is more important than anything else. I don't know when they got started, but Rust has only matured to "mainstream" stability/acceptance in the last 2-3 years. Finally, DBs often have a very layered architecture and theyt could easily compartmentalize pieces such that latency sensitive bits could be done in native Rust. They're not apparently doing this, but I could see them doing things like moving the page buffer or column indices or storage engine over to Rust over time for performance benefits. All power to them, it's great to see them working with Rust. (aside: my email history looks like I spoke to a recruiter there at some point, maybe, but didn't interview? I think if I'd known they were playing with Rust I would have given that more attention...)
- nhourcard 3y agoAlso this looks to be more of an analytical, column oriented, database. So I can imagine they're optimizing more for throughput than transactional latency. Yes that is the case
- Joe8Bit 3y ago"Hard to write but easy for customers to deploy" is my guess. There are a bunch of very high-performance computing use-cases in finance (quant, HFT) that Java gets used for pretty routinely. That's a very attractive market to build primitives like databases for, they have very deep pockets, but you need to play in their ecosystem. I've seen this "GC-less" Java in those use-cases quite a bit. From a conceptual design POV it's likely not the best approach, but there's a lot of sunk cost in that eco-system and a lot of trust and expertise where "Choosing a better language" is often several orders of magnitude more expensive.
- nu11ptr 3y agoI mean no disrespect to the authors, but this seems like an extremely painful way to write an application. How common is it to write Java apps this way? It seems like there must have been a better language choice than trying to work against the GC and rewriting parts of the standard lib? And now adding in Rust via JNI? It just feels very painful to me.
- giancarlostoro 3y agoLooks like they have an interesting range of customers[0] so my take on "why even add the Rust" is because their customers are already using Java, a total rewrite might be considered irresponsible as it would be incompatible with their existing customer base. I do have to wonder though, if some serious Java refactoring in any way, would have helped at all. How many code smells do they have going on in their codebase? Or is it to the best of their knowledge a really clean Java codebase? Note Discord themselves has used Rust for bottlenecks with Erlang/Elixir/Beam[1]. [0]: https://questdb.io/customers/ https://questdb.io/customers/ [1]: https://discord.com/blog/using-rust-to-scale-elixir-for-11-million-concurrent-users https://discord.com/blog/using-rust-to-scale-elixir-for-11-m...
- nu11ptr 3y agoI'm not suggesting they rewrite. I'm questioning if Java was ever the best choice for this application in the first place. They are using Java as if it were C++. Perhaps it would have been better to write it in C++ in that case? I'm not drawing conclusion, as I'm not in their domain and have not given this a lot of thought, but it just strikes me as a particularly shaky foundation.
- vbezhenar 3y agoJava brings some good things over C++. For example memory-safe VM, easier language with way less gotchas, better tooling, awesome IDEs. I'd say, if C++ performance is not essential, Java might actually be a good choice. It's very fast and GC issues could be worked around.
- amunra__ 3y agoI'm the author of the blog post. The focus of the article really is about JNI in Rust. I see most questions are about "Why did you not use X language instead?", so let me try and address this. To answer the "Why not just Rust", I should first mention that Rust was still in its early days (before 1.0), and it was a risky bet to choose an emerging language. The project was started by Vlad (our CEO) who had a background writing high performance Java in the trading space. The Zero-GC techniques - whilst uncommon in open source software - are mature and a staple of writing high performance code in the financial industry. The product evolved organically, feature after feature. I personally joined the team from a C and C++ background, having previously moved from a project that suffered from minute-long compile times from single .cpp files due to template overuse. Whilst I do miss how expressive high-level C++ can be, Java has really good tooling support. When writing systems-style software most of what matters in terms of performance is how we call system calls, manage memory and debug and profile. This is an area where Java really shines. Don't get me wrong: In the absolute sense I think C++ tools tend to be better (Linux Perf is awesome!), but Java tooling is _there_. IntelliJ makes it trivially easy to run a test under the debugger reliably and consistently. It's equally easy to run a profiler and to get code coverage. The same tools work across all platforms too, might I add. It's not necessarily better, but it's easier. Turns out that while a little quaint, using Java turned out to be a pretty good choice in my opinion in practice. Times have moved on. The Rust community really cares about tooling, and it's one of the reasons why we've picked it over expanding our existing C++ codebase: We just want to get stuff done and have enough time left in our dev cycle to properly debug and profile our code.
- belter 3y agoThanks for the interesting post. Do you plan to maybe use in the future JEP 442? https://openjdk.org/jeps/442 https://openjdk.org/jeps/442
- amunra__ 3y agoWhen the time is right. There's finally new APIs coming in the Java space that will make native-code interop easier and more reliable. Our open source database edition can also be used embedded though, so we can only upgrade at the pace of our customers and because of that we still are compatible all the way down to Java 8. Were it not for this detail, we'd probably consider it a lot sooner.
- crustycoder 3y agoNowadays there's a much better technology than JNI for doing this sort of thing in Java. https://docs.oracle.com/en/graalvm/jdk/21/docs/reference-manual/native-image/ https://docs.oracle.com/en/graalvm/jdk/21/docs/reference-man...
- karussell 3y agoNative image is unrelated to this topic here (hence the downvotes, I guess). Still GraalVM could be an interesting solution as Rust seems to be supported: https://www.graalvm.org/latest/reference-manual/llvm/Compiling/ https://www.graalvm.org/latest/reference-manual/llvm/Compili...
- crustycoder 3y agohttps://github.com/oracle/graal/blob/master/substratevm/src/com.oracle.svm.tutorial/src/com/oracle/svm/tutorial/CInterfaceTutorial.java https://github.com/oracle/graal/blob/master/substratevm/src/... https://github.com/oracle/graal/blob/master/substratevm/src/com.oracle.svm.tutorial/native/cinterfacetutorial.c https://github.com/oracle/graal/blob/master/substratevm/src/... And in any case: JEP draft: Prepare to Restrict The Use of JNI https://openjdk.org/jeps/8307341 https://openjdk.org/jeps/8307341 JEP 442: Foreign Function & Memory API (Third Preview) https://openjdk.org/jeps/442 https://openjdk.org/jeps/442
- klauserc 3y agoVery interesting to see how the Rust-JNI interface gets used in a production environment (e.g., the topic of unifying logging typically doesn't come up in "your first Rust-JNI app" tutorials, for instance) I do have one question around the assignment to `static mut CALL_STATE`. Don't you need some form of synchronization/memory fence/memory barrier to make sure that other threads see that assignment? On x86/x64 it probably doesn't matter (total store order), but other architectures are less lenient.
- amunra__ 3y agoWe have an initialisation step as soon as we load the jni lib that takes care of this. Given that this gets done before any other threads are started, I don't think there'd be an issue. Good point :-)
- pron 3y ago> We seldom use the new keyword and objects are designed to be pooled and reused. Just note that depending on the selection of the GC, this kind of usage may make the GC work more than when allocating new objects, not less. In particular, with the newer GCs -- G1 and ZGC -- mutating existing objects may be more costly than allocating new ones depending on circumstances. In general, the new GCs are optimised to work the least and give the best performance when the allocation rate is neither too high nor too low; the new GCs also reuse memory better than object pools. Reusing objects also precludes scalarization optimisations, i.e. not every `new Foo` actually results in a heap allocation, and can be optimised to work directly in registers. So while on very old JVMs (such as Java 8) a "zero allocation" strategy may result in better performance and in "zero GC", on newer JVMs it may result in worse performance and more GC work (in fact, it will almost surely not yield zero GC). While it depends on many variables, I would advise against a zero allocation strategy on newer JVMs as the default path toward better performance or even better latency. It's an approach that seems to be very strongly coupled to the way the JVM was designed over a decade ago, but a lot has changed since then. Additionally, Java now offers manual memory management and efficient FFI that are significantly better than what JNI offered: https://openjdk.org/jeps/442 https://openjdk.org/jeps/442
- coldbrewed 3y ago> In general, the new GCs are optimised to work the least and give the best performance when the allocation rate is neither too high nor too low; the new GCs also reuse memory better than object pools. This a little bit surprising to me that low object creation can degrade GC performance; what's the failure mode for G1/ZGC in this scenario?
- Groxx 3y agoIn many cases, because it defeats generational collection. It pushes everything into longer generations because they hang around longer. Doing more young generation collection is sometimes cheaper in aggregate (more frequent but far smaller and usually much more efficient) than adding more data to the older generations (less frequent and more costly, longer pauses, for object pools it happens on all of it even when none of it is currently used, etc). But as doctorpangloss said: so many caveats. There's ample evidence that it is both better and worse, it depends on lots of details. The main thing you can confidently claim is that it is not the majority of code, so most language optimizations will choose to improve straightforward and common stuff at the cost of this niche. Not always, but there is definitely more energy in improving the 90%+ cases and that adds up over time. Squeezing out the last bits of performance requires constant upkeep.
- exabrial 3y ago> JNI Have you guys benchmarked FFI in Java 21 (preview, now release) yet? :) Yes I know it's super new, but I'm curious if there is a benefit in terms of: 1. performance 2. ease of maintenance 3. ease of finding production problems
- exabrial 3y agoI've never heard of QuestDB until this post, but I very much like what you guys are doing. InfluxDB 1.x had a chance to be great: the wrote a time series database that used a SQL-like dialect, they had a really nice alerting platform, they had a really nice visiualization tool. Then they abandoned all of that to jump on the hype train of "Write a new programming language!" which was the hot thing like 5 or 6 hype cycles ago. We've _never_ upgraded to their 2.x product because it literally threw away our investment. I think if you guys get pick up where they departed you'll be tremendously successful.
- nhourcard 3y agothanks for the kind words! We want to stick with SQL - having done a few extensions to make it easier to work with time series data such as SAMPLE BY, LATEST ON, etc. Window functions that the product has been lacking for some time are next to bridge the gap vs other more mature platforms while offering something very new and unique on the performance side, especially ingestion related.
- jurgenkesker 3y agoHow are you feeling about the upcoming InfluxDB 3? They moved to Rust and support InfluxQL again.
- exabrial 3y agoI'll have to take a look. If Kapacitor is back, then I'm in.