11 ms·
I discovered something amazing when working with some people who were writing HFT software. Why do you need 1TB of RAM in these machines? Because when you're J
by oppositelock 6y ago
I discovered something amazing when working with some people who were writing HFT software.
Why do you need 1TB of RAM in these machines? Because when you're Java based, you want to avoid stop-the-world GC pauses. These trading systems only have to be up from 9:30AM-4:30PM EST, so they simply disable GC altogether! At the end of a trading day, restart the app or reboot the system.
- anthonypasq 6y agowhy not just use C++ or something and never deallocate memory then?
- layer8 6y agoMalloc is usually slower than GC-based allocation (which basically just increments a pointer). Of course, one can emulate this via custom allocators in C++. My guess is that Java development is just easier and quicker, and the JIT may even result in more effective optimizations than AOT compilation.
- mathieubordere 6y agoin C++ they could also try using Profile-guided optimization to optimize hot code paths.
- amelius 6y agoIn C++ you can link with your own version of malloc (one which just returns the next consecutive block, and can be implemented simply by adding the allocated size to a pointer).
- oppositelock 6y agoYou generally choose the language which has the libraries and ecosystem which solves your problem. For instance, you'd be silly to use anything but Java to work with Hadoop. This is the pragmatic choice.
- ptx 6y agoBut as soon as you make use of almost any library from the ecosystem, you get allocations. Doesn't that defeat the purpose? Why not isolate the Java-using part in a separate process from the non-allocating fast part?
- dcolkitt 6y agoEven if you completely disable GC, Java is still allocating tons of ephemeral objects on the heap. Which in turn is leading to expensive and unpredictable page faults. In contrast, C++'s default is to allocate objects on the heap unless you knowingly call new/malloc. Of course, it's possible to write Java in such a way to minimize this type of heap-thrashing. But by that point, you're already doing the equivalent of C++'s manual memory book-keeping anyway. It's not like "just provision a ton of memory and turn off GC" is a free lunch.
- fh973 6y agoThat's covered in Java by escape analysis and allocation on the stack.
- bluestreak 6y agoWe have written a database in zero GC java and one thing I have not seen any evidence of "escape analysis". @State(Scope.Thread) @BenchmarkMode(Mode.AverageTime) @OutputTimeUnit(TimeUnit.NANOSECONDS) public class EscBenchmark { Rnd rnd = new Rnd(); public static void main(String[] args) throws RunnerException { Options opt = new OptionsBuilder() .include(EscBenchmark.class.getSimpleName()) .warmupIterations(5) .measurementIterations(5) .forks(1) .addProfiler(GCProfiler.class) .build(); new Runner(opt).run(); } @Benchmark public int testEscapeAnalysis() { int[] tuple = {0, 2}; // esc analysis? where are you? return tuple[rnd.nextPositiveInt() % 2]; } } And the output of GC profiler: Benchmark Mode Cnt Score Error Units EscBenchmark.testEscapeAnalysis avgt 5 8.234 ± 0.029 ns/op EscBenchmark.testEscapeAnalysis:·gc.alloc.rate avgt 5 2647.216 ± 9.275 MB/sec EscBenchmark.testEscapeAnalysis:·gc.alloc.rate.norm avgt 5 24.000 ± 0.001 B/op EscBenchmark.testEscapeAnalysis:·gc.churn.G1_Eden_Space avgt 5 2643.140 ± 177.137 MB/sec EscBenchmark.testEscapeAnalysis:·gc.churn.G1_Eden_Space.norm avgt 5 23.963 ± 1.613 B/op EscBenchmark.testEscapeAnalysis:·gc.count avgt 5 157.000 counts EscBenchmark.testEscapeAnalysis:·gc.time avgt 5 103.000 ms
- iAm25626 6y agoseems pragmatic; remind me of this old story: https://devblogs.microsoft.com/oldnewthing/20180228-00/?p=98125 https://devblogs.microsoft.com/oldnewthing/20180228-00/?p=98...
- amelius 6y agoThis is similar to how missiles don't need GC. https://devblogs.microsoft.com/oldnewthing/20180228-00/?p=98125 https://devblogs.microsoft.com/oldnewthing/20180228-00/?p=98...
- deepGem 6y agoI guess what you mean is the no-op garbage collector which is available in Java 11 http://openjdk.java.net/jeps/318 http://openjdk.java.net/jeps/318 Even this isn't fail proof right ? Last-drop latency improvements. For ultra-latency-sensitive applications, where developers are conscious about memory allocations and know the application memory footprint exactly, or even have (almost) completely garbage-free applications, accepting the GC cycle might be a design issue. There are also cases when restarting the JVM -- letting load balancers figure out failover -- is sometimes a better recovery strategy than accepting a GC cycle. In those applications, long GC cycle may be considered the wrong thing to do, because that prolongs the detection of the failure, and ultimately delays recovery. So essentially you need to know the memory footprint of your applications. To err on the side of caution just get as much memory that is available in the market. At this point like another reply mentions here aren't you just better off writing C++ code ?
- rch 6y agoDepends if your returns are based on writing faster code or shipping correct code faster. Once the advantages of the former are eliminated, optimize for the latter.
- aasasd 6y ago> * I guess what you mean is the no-op garbage collector which is available in Java 11* I think I heard of the same pattern before 2018, so I guess there are other ways to do it. Perhaps Java has some option like memory usage value at which the GC is run.
- hinkley 6y agoI wonder if it would be cheaper for them to have a hot standby system, and swap them every few hours.
- benjaminjackman 6y agoSpeaking from experience with JVM HFT applications (we used Scala). There are a lot of tricks though to not require 1TB. And allocation in general is a bad idea even if you don't collect because it scatters stuff all over memory and messes up cache locality. You really, really don't want to allocate in a performance sensitive jvm application if you can avoid it. It's the opposite of a lot of what I was told and taught (e.g. never do object pooling), but empirically, in my experience, allocations are the biggest slowdown. You can get an application a lot faster just by opening up the memory allocation tab in a jmc flightrecording and refactoring the biggest allocators, usually there is a lot of easy to optimize low hanging fruit that will give good performance improvements, even better than focusing on hot spots in code (in my personal experience). By far the biggest allocator in trading is going to be marketdata and calculations on it. For reading marketdata from the exchange it's best to leave raw data in memory and access it with a ByteBuffer / sun.misc.unsafe. Under this pattern classes have 1 value, the memory address to pass into sun.misc.unsafe, then everything from there on is done with offsets onto that address. For calculations it's better to write things as static functions, or use object pooling. In the course of optimizing a trading engine I wrote lots and lots of code to get allocations down to zero. It's definitely doable, but best done from the start, I refactored an existing trading engine to do that, it was not very fun.
- Androider 6y ago1TB of RAM is cheap though, versus spending engineering hours. Reminds me of the classic WTF "That would've been an option too" https://thedailywtf.com/articles/That-Wouldve-Been-an-Option-Too https://thedailywtf.com/articles/That-Wouldve-Been-an-Option...
- MangoCoffee 6y agoagree. hardware is getting cheaper and cheaper
- ksec 6y agoExcept Memory / DRAM.
- throwaway10M 6y agoRAM is not the only aspect of a Garbage Collector.... Over a decade ago, we used a third-party pre-trade risk system that was implemented in Java. Since it was a "service" (we connected to a TCP port), the underlying tricks they used to make it "fast" were transparent to us... until it was not. They highly tuned their GC to where there was seldom any GC. One day, the third-party made a change to a supervisory service to generate more periodic monitoring emails. The file handles from this actor were apparently not GC'd and held by the process until the system ran out of file handles. That made the service stop working properly. But, in addition to alerting, that service had a more important job: it was a post-trade risk "watchdog" to the pre-trade risk gateways, which we sharded across. The various pre-trade risk gateways, upon not hearing from the watchdog authority then 1) began cancelling outstanding orders and 2) not allow new orders. We saw this happening haphazardly over 10 minutes and had little time and capability to recover; this also happened near the end of the trading day when some orders (MOC, LOC) are not cancellable. This was across our whole organization so it affected many strategies. So it wasn't as simple as "stop all" and although we had many checks and recovery procedures, this was a pretty special Chaos Monkey. We ended up with a >$1B basket of random stocks that cost >$10M to liquidate over the next few days. $10M evaporating in 10 minutes because of "GC optimizations" and poorly considered OS settings. The #1 risk in algo/HFT trading is not financial, but operational. [Some of the technical details may be slightly off, as it was third-party; my outlook is pieced together from post-mortems. Also, no other party (e.g. broker, SIPC, market participant, the third-party) was financially affected besides our firm.]
- aasasd 6y agoGame devs face similar challenges, especially on phones. Though they have to be less radical in their solutions—which afaik usually boil down to ‘preload the level into memory, use object pools, and don't do any new allocations’.
- mindentropy 6y agoWhy don't they use fixed size arrays like in embedded systems and not allocate memory dynamically?
- The_rationalist 6y agoHow was these achieved before Epsilon? http://openjdk.java.net/jeps/318 http://openjdk.java.net/jeps/318 Also ZGC / Shenandoah can help a lot otherwise