16 ms·
C++ patterns for low-latency applications including high-frequency trading
- jeffreygoesto 2y agoReminds me of https://github.com/CppCon/CppCon2017/blob/master/Presentations/When%20a%20Microsecond%20Is%20an%20Eternity/When%20a%20Microsecond%20Is%20an%20Eternity%20-%20Carl%20Cook%20-%20CppCon%202017.pdf https://github.com/CppCon/CppCon2017/blob/master/Presentatio...
- munificent 2y agoThis is an excellent slideshow. The slide on measuring by having a fake server replaying order data, a second server calculating runtimes, the server under test, and a hardware switch to let you measure packet times is so delightfully hardcore. I don't have any interest in working in finance, but it must be fun working on something so performance critical that buying a rack of hardware just for benchmarking is economically feasible.
- a_t48 2y agoThe self driving space does this :)
- nine_k 2y agoDelightfully hardcore indeed! But of course you don't have to buy a rack of servers for testing, you can rent it. Servers are a quickly depreciating asset, why invest in them?
- a_t48 2y agoYou'd want it to be the exact same hardware as in production, for one.
- CyberDildonics 2y agoWhy would replaying data for testing be "Delightfully hardcore indeed!". That's how people program in general, they run the same data through their program. Servers are a quickly depreciating asset, why invest in them? I don't think they are a quickly depreciating asset compared to the price of renting, but you would want total control over them in this scenario anyway.
- nine_k 2y agoI thought that the hardcore part is taking the data from the switch to account for the network latency.
- munificent 2y ago> Why would replaying data for testing be "Delightfully hardcore indeed!". Replaying data isn't hardcore. Buying a dedicated server and running it through a dedicated switch just to gather precise timing info is.
- sneilan1 2y agoI've got an implementation of a stock exchange that uses the LMAX disruptor pattern in C++ https://github.com/sneilan/stock-exchange https://github.com/sneilan/stock-exchange And a basic implementation of the LMAX disruptor as a couple C++ files https://github.com/sneilan/lmax-disruptor-tutorial https://github.com/sneilan/lmax-disruptor-tutorial I've been looking to rebuild this in rust however. I reached the point where I implemented my own websocket protocol, authentication system, SSL etc. Then I realized that memory management and dependencies are a lot easier in rust. Especially for a one man software project.
- JedMartin 2y agoIt's not easy to get data structures like this right in C++. There are a couple of problems with your implementation of the queue. Memory accesses can be reordered by both the compiler and the CPU, so you should use std::atomic for your producer and consumer positions to get the barriers described in the original LMAX Disruptor paper. In the get method, you're returning a pointer to the element within the queue after bumping the consumer position (which frees the slot for the producer), so it can get overwritten while the user is accessing it. And then your producer and consumer positions will most likely end up in the same cache line, leading to false sharing.
- sneilan1 2y ago>> In the get method, you're returning a pointer to the element within the queue after bumping the consumer position (which frees the slot for the producer), so it can get overwritten while the user is accessing it. And then your producer and consumer positions will most likely end up in the same cache line, leading to false sharing. I did not realize this. Thank you so much for pointing this out. I'm going to take a look. >> use std::atomic for your producer Yes, it is hard to get these data structures right. I used Martin Fowler's description of the LMAX algorithm which did not mention atomic. https://martinfowler.com/articles/lmax.html https://martinfowler.com/articles/lmax.html I'll check out the paper.
- hi_dang_ 2y ago
- gedanziger 2y agoVery cool intro to the subject!
- globular-toast 2y agoIs there any good reason for high-frequency trading to exist? People often complain about bitcoin wasting energy, but oddly this gets a free pass despite this being a definite net negative to society as far as I can tell.
- arcimpulse 2y agoIt would be trivial (and vastly more equitable) to quantize trade times.
- FredPret 2y agoYou mean like settling trades every 0.25 seconds or something like that? Wouldn't there be a queue of trades piling up every 0.25 seconds, incentivizing maximum speed anyway?
- foobazgt 2y agoUsually the proposal is to randomize the processing of the queue. So, as long as your trades get in during the window, there's no advantage to getting in any earlier. In theory the window is so small as to not have any impact on liquidity but wide enough to basically shut down all HFT.
- affyboi 2y agoHow would that work? Would you randomly select a single order posted, go by market participant (like randomly select some entities that posted a trade in this window), and would you allow prices to move during this window?
- FredPret 2y agoNon-bitcoin transactions are just a couple of entries in various databases. Mining bitcoin is intense number crunching. HFT makes the financial markets a tiny bit more accurate by resolving inconsistencies (for example three pairs of currencies can get out of whack with one another) and obvious mispricings (for various definitions of "obvious")
- munificent 2y ago> The noted efficiency in compile-time dispatch is due to decisions about function calls being made during the compilation phase. By bypassing the decision-making overhead present in runtime dispatch, programs can execute more swiftly, thus boosting performance. The other benefit with compile-time dispatch is that when the compiler can statically determine which function is being called, it may be able to inline the called function's code directly at the callsite. That eliminates all of the function call overhead and may also enable further optimizations (dead code elimination, constant propagation, etc.).
- xxpor 2y agoOTOH, it might be a net negative in latency if you're icache limited. Depends on the access pattern among other things, of course.
- munificent 2y agoYup, you always have to measure. Though my impression is that compilers tend to be fairly conservative about inlining so that don't risk the inlining being a pessimization.
- rasalas 2y agoIn my experience it's the "force inline" directives that can make this terrible. I had a coworker who loved "force inline". A symptom was stupidly long codegen times on MSVC.
- foobazgt 2y agoMy experience has been that it's rather heuristic based. It's a clear win when you can immediately see far enough in advance to know that it'll also decrease the amount of generated code. You can spot trivial cases where this is true at the point of inlining. However, if you stopped there, you'd leave a ton of optimizations on the table. Further optimization (e.g. DCE) will often drastically reduce code size from the inlining, but it's hard to predict in relationship to a specific inlining decision. So, statistics and heuristics.
- astromaniak 2y agoJust in case you are a pro developer, the whole thing is worth looking at: https://github.com/CppCon/CppCon2017/tree/master/Presentations https://github.com/CppCon/CppCon2017/tree/master/Presentatio... and up
- twic 2y agoMy emphasis: > The output of this test is a test statistic (t-statistic) and an associated p-value. The t-statistic, also known as the score, is the result of the unit-root test on the residuals. A more negative t-statistic suggests that the residuals are more likely to be stationary. The p-value provides a measure of the probability that the null hypothesis of the test (no cointegration) is true. The results of your test yielded a p-value of approximately 0.0149 and a t-statistic of -3.7684. I think they used an LLM to write this bit. It's also a really weird example. They look at correlation of once-a-day close prices over five years, and then write code to calculate the spread with 65 microsecond latency. That doesn't actually make any sense as something to do. And you wouldn't be calculating statistics on the spread in your inner loop. And 65 microseconds is far too slow for an inner loop. I suppose the point is just to exercise some optimisation techniques - but this is a rather unrepresentative thing to optimise!
- nickelpro 2y agoFairly trivial base introduction to the subject. In my experience teaching undergrads they mostly get this stuff already. Their CompArch class has taught them the basics of branch prediction, cache coherence, and instruction caches; the trivial elements of performance. I'm somewhat surprised the piece doesn't deal at all with a classic performance killer, false sharing, although it seems mostly concerned with single-threaded latency. The total lack of "free" optimization tricks like fat LTO, PGO, or even the standardized hinting attributes ([[likely]], [[unlikely]]) for optimizing icache layout was also surprising. Neither this piece, nor my undergraduates, deal with the more nitty-gritty elements of performance. These mostly get into the usage specifics of particular IO APIs, synchronization primitives, IPC mechanisms, and some of the more esoteric compiler builtins. Besides all that, what the nascent low-latency programmer almost always lacks, and the hardest thing to instill in them, is a certain paranoia. A genuine fear, hate, and anger, towards unnecessary allocations, copies, and other performance killers. A creeping feeling that causes them to compulsively run the benchmarks through callgrind looking for calls into the object cache that miss and go to an allocator in the middle of the hot loop. I think a formative moment for me was when I was writing a low-latency server and I realized that constructing a vector I/O operation ended up being overall slower than just copying the small objects I was dealing with into a contiguous buffer and performing a single write. There's no such thing as a free copy, and that includes fat pointers.
- mbo 2y agoOut of interest, do you have any literature that you'd recommend instead?
- nickelpro 2y agoOn the software side I don't think HFT is as special a space as this paper makes it out to be.[1] Each year at cppcon there's another half-dozen talks going in depth on different elements of performance that cover more ground collectively than any single paper will. Similarly, there's an immense amount of formal literature and textbooks out of the game development space that can be very useful to newcomers looking for structural approaches to high performance compute and IO loops. Games care a lot about local and network latency, the problem spaces aren't that far apart (and writing games is a very fun way to learn). I don't have specific recommendations for holistic introductions to the field. I learn new techniques primarily through building things, watching conference talks, reading source code of other low latency projects, and discussion with coworkers. [1]: HFT is quite special on the hardware side, which is discussed in the paper. The NICs, network stacks, and extensive integration of FPGAs do heavily differentiate the industry and I don't want to insinuate otherwise. You will not find a lot of SystemVerilog programmers at a typical video game studio.
- winternewt 2y agoI made a C++ logging library [1] that has many similarities to the LMAX disruptor. It appears to have found some use among the HFT community. The original intent was to enable highly detailed logging without performance degradation for "post-mortem" debugging in production environments. I had coworkers who would refuse to include logging of certain important information for troubleshooting, because they were scared that it would impact performance. This put an end to that argument. [1] https://github.com/mattiasflodin/reckless https://github.com/mattiasflodin/reckless
- ibeff 2y agoThe structure and tone of this text reeks of LLM.
- ykonstant 2y agoI am curious: why does this field use/used C++ instead of C for the logic? What benefits does C++ have over C in the domain? I am proficient in C/assembly but completely ignorant of the practices in HFT so please go easy on the explanations!
- jqmp 2y agoC++ is more expressive and allows much more abstraction than C. For a long time C++ was the only mainstream language that provided C-level performance as well as rich abstractions, which is why it became popular in fields that require complex domain modeling, like HFT, gamedev, and graphics. (Of course one can debate whether this expressivity is worth the enormous complexity of the language, but in practice people have empirically chosen C++.)
- apantel 2y agoAnyone know of resources like this for Java?
- poulpy123 2y agothe irony being that if something should not be high frequency, it is trading