7 ms·
In short, you need to write Java like C++ for low latency.
by InGodsName 8y ago
In short, you need to write Java like C++ for low latency.
- qalmakka 8y agoThis. What is the point of using a higher level, garbage collected language instead of C++ or Rust if you're going to spend the same amount of time fighting against the VM or the Garbage collector to make it actually work the way you want it to? It feels like it's just for the sake of avoiding any kind of exposure to a newer language.
- misja 8y agoI can think of a couple of reasons; Java code is portable across OS'es; also while this code reads like C++, it might be part of a larger system with less performance-critical idiomatic Java.
- Cyph0n 8y agoI'm not sure that portability is a good reason to use Java in this case. Firstly, these trading firms probably have tight control over their server architectures. Secondly, once you start using low-level Java idioms in order to optimize your code, you've started approaching the point where differences in performance between JVM implementations across platforms become noticeable.
- tpetry 8y agoI am waiting for an article on hn how a high frequency trading vompany has rewritten their stack in rust because of the threading features.
- Gladdyu 8y agoUnlikely any time soon - most HFTs will have massive, battle-tested C++ code bases already. The data ownership model in HFT is usually very clear - receive packet, pass through some processing pipeline, emit packet, on the same thread. If you want to send data into another part of the system (eg. because different threads handle different securities), posting a message onto their event queue works well.
- snaky 8y agoNot happening. Rewriting some hot path parts of the stack in Verilog is a usual thing, and rewriting some other parts in Coq is the future.
- twic 8y agoIs it possible to synthesize hardware from Coq?
- mruts 8y agoNot entirely on topic, but Jane Street (HFT ETF market maker) wrote an OCaml to FPGA compiler https://www.janestreet.com/tech-talks/ocaml-all-the-way-down/ https://www.janestreet.com/tech-talks/ocaml-all-the-way-down...
- snaky 8y agoYes. https://deepspec.org/entry/Project/Kami https://deepspec.org/entry/Project/Kami http://conal.net/blog/posts/haskell-to-hardware-via-cccs http://conal.net/blog/posts/haskell-to-hardware-via-cccs
- Thaxll 8y agoThe threading feature that are not stable / experimental? There is not way someone is moving to Rust for this.
- liotier 8y agoMaybe one already has a heap of Java code and the latency-sensitive part is only a tiny fraction where optimization will be focused.
- danielshaya 8y agoCompletely agree with that 90% of the code I usually work with is not of the 'hot path'.
- kazinator 8y agoToo bad the 90/10 rule doesn't apply to size. Whereas 90% of the time might be spent in 10% of the code image, no 10% chunk of the image can ever be found that contributes to 90% of its bloat. 1% of the image contributes 1% to the bloat, 10% to 10% and so on. :)
- kitd 8y agoThe GC isn't the feature you want though, it's the JIT compiler. JIT compilers optimise based on the actual runtime characteristics of the running code, not the best guesses that a static compiler has to make. The effect can be significant. It's no surprise that the Java leaders in Techempower benchmarks match or surpass the C++ ones.
- zozbot123 8y agoYou can use profile-guided compilation with an AOT compiler, too-- Or you could just add optimization hints to the performance-critical parts of the code. Either way, the compiler doesn't need to "guess"
- gpderetta 8y agoAlso one of the advantages of JIT is that it can take advantage of whatever dedicated instruction is available on the current CPU, while an AOT compiler (discounting runtime dispatching) need to target the lowest common denominator. In this case though you would know exactly the target hardware.
- adrianN 8y agoOn the other hand, the JIT has to balance the runtime of its optimization passes with the expected improvement in execution speed of the code. An AOT compiler can afford much more expensive optimizations.
- kod 8y agoI've talked with HFT firms that spent effort fighting the JIT because they cared more about the latency of rarely reached code paths than the common path.
- twic 8y agoJava doesn't beat C++ in raw computational benchmarks. The JIT compiler is a marvel, but it doesn't beat peer AOT compilers. The Techempower benchmarks mostly show how good the entrant's web-serving stack is. Java has a couple of decades of multiple skilled, well-established groups competing to build the fastest web stack. C++ has Boost.Beast, which, although its developers are smart, wise, widely respected, sexy, and moderators of a Slack i frequent, can't compete in terms of resources.
- simion314 8y agoYou heavy optimize only the critical/hot parts, similar how C developers sometimes optimize the code by using assembly in a critical section.
- jlg23 8y agoThe point is development time. Even in time critical applications the majority of the code is not time critical, you only optimize the small part that really matters.
- sbanach 8y agoWith Java vs C++, the experience seems similar. In C++, memory allocation is slow, so you worry about memory allocation up front. In Java, GC is a pain so you worry about memory allocation up front. As Daniel points out, there are a few extra tricks available to native developers (eg vector instructions), so there's a small factor speed-up in best-practice C++ vs best-practice Java, but this is closing (eg the JIT increasingly uses vector instructions). Overwhelmingly the difference between a slow system and a fast system is appropriate choices of algorithms and architecture, not the language.
- bilbo0s 8y ago>In C++, memory allocation is slow, so you worry about memory allocation up front. In Java, GC is a pain so you worry about memory allocation up front... This. A lot of people seem to be willfully dismissive of this reality when they are fanboying for their favorite languages. Ask for memory in C++, you'll take a HUGE hit getting it. Ask for memory in Java, and you'll get it lightning fast, but you'll take a huge hit getting rid of it. Either way, you should carefully plan memory usage up front. Basically, one road leads to death and despair, the other leads to disease and destruction... pray you choose wisely.
- gpderetta 8y agoAny moderately performance sensitive code in C++, assuming it needs allocations at all will use a dedicated allocator. While it is hard to beat existing general purpose allocators in all the scenarios, it is easy for specific use cases.
- kjeetgill 8y agoTrue but these approaches in both C++ and Java start to converge: Arena allocation vs Pooling/unsafe. The question is, how annoying is your performance insensitive code?
- SeanLuke 8y agoPareto's law. You optimize the few parts that matter speedwise, and the rest you can do in ordinary "higher level, garbage collected" fashion. In C++ you can't do that. This is the standard argument for Lisp type-optimization (DECLAIM, THE, etc.) as well.
- kazinator 8y ago> In C++ you can't do that What?
- agumonkey 8y agoMany people are saying that cpp benefits are cancelled by all the pointer/template voodoo required to get them. Now I'd be happy to see high perf low latency talks about rust, go, ocaml, haskell, sbcl, whatever. I believe limit-cases like these are always very very instructive.
- SomeHacker44 8y agoAs a very long time Haskell user, I wonder how much latency management would be controlling for the lazy evaluation thunks being evaluated at certain deterministic times. In my gut I would not choose Haskell for my first choice for a low latency language. I would definitely have looked at Rust before Haskell.
- agumonkey 8y agoOf course, but still, it's interesting to see how far lazy fp can be reasoned about on this topic. I always love optim talks whatever the language (granted its not a clusterfk interpreter).
- CyberDildonics 8y agoHaskell abstracts away order of operations. Both high performance and low latency benefit from the ability to be very specific about as much as possible and Haskell take the opposite approach. This may be good for data flow and throughput, but as soon as you need to control explicit order, state and resource usage Haskell is going to be working against you.
- AnimalMuppet 8y agoIt seems to me that part of high-performance computing is optimizing memory layout (so you can vectorize, and to avoid cache misses). I have a hard time seeing that even in Java; I have a really hard time seeing that in Haskell. Am I missing something?
- marcosdumay 8y agoAs is usual in Haskell for this kind of question, you can theoretically capture the information at the type level, and there are good odds somebody is working in a GHC extension for that. But on practice, I don't think you are missing anything.
- dilap 8y agoMemory safety is a pretty huge one! Better tooling, too.
- deleted 8y ago[deleted]
- SuspiciousSwan 8y agoThe advantage over C++ isn't about the language itself. If your datastores all have functioning java connections, and your data modeling is already all done in java, it can be worth the pain to do performance critical code in java. I would rather do it in C++, but sometimes the most rational path forward is ugly. Rust's future is far from certain. I would feel irresponsible using it on a project that wasn't self contained.
- pron 8y agoHave you watched the talk? The point is that it requires less effort overall, and if you need to go significantly beyond that, C++ can't save you anyway and you'll use FPGA.
- kazinator 8y agoC++ can be a higher level, garbage-collected language. http://www.stroustrup.com/bs_faq.html#garbage-collection http://www.stroustrup.com/bs_faq.html#garbage-collection
- pkolaczk 8y agoOnly theoretically. It is not possible to implement a GC as efficient as Java GCs for C++. Pointer arithmetic, placement new, and other low level features make it hard to determine pointers from non-pointers and to move things in memory.
- kazinator 8y ago"It is not possible" claims are very hard to substantiate. Java is just someone's C++ program. > "Pointer arithmetic, placement new, and other low level features make it hard to determine pointers from non-pointers and to move things in memory." Real applications in a garbage collected language are mixtures of the managed code, plus unmanaged components (like foreign libraries). Garbage collected C++ can work in much the same way: there is a framework for managed code that uses GC, and then there are unsafe parts that are analogous to foreign code. (Except, the managed code can much more easily and efficiently interact with this code than the typical FFI.) There is more than one way to integrate garbage collection into C++, in any case.
- kazinator 8y ago> move things in memory. Not all garbage-collection schemes move things in memory; only copying and/or compacting garbage collectors do.
- jonas21 8y agoThere's a good answer to your question in the talk here: https://youtu.be/BD9cRbxWQx8?t=1025 https://youtu.be/BD9cRbxWQx8?t=1025