7 ms·
Parsing Decimals four times faster
- solmag 5y agoIt's a nice article, but don't they run against Amdahl's law? I mean they're parsing JSON from some web api?
- Bootvis 5y agoIsn't Amdahl's law only applicable to systems that some parallelism?
- edflsafoiewq 5y agoAmdahl's law basically says when you speed up part of a task, the overall speedup is limited by the part you didn't speed up. The speedup doesn't have to come from parallelism.
- CUViper 5y agoIt can be applied more generally. If you're focusing on improving the performance of something that only takes 10% of your time, then the other 90% will still dominate.
- Bootvis 5y agoAha, I see. If you’re in a race where every microsecond counts you would still care about there 10% after you’ve made sure that the 90% can’t be further optimized.
- marginalia_nu 5y agoAhmdals law deals with throughput or start-to-stop time for processing a given work, not necessarily latency.
- funcDropShadow 5y agoPossibly, but if you are processing live market data, you are in the arena of almost hard real-time systems. I.e. you effectively have a time budget to process every single tick/message. Then, even low-level optimization can make a large difference. E.g. if you can make some helper method use less memory, i.e. it uses less of your caches, that can make a big difference on your business logic. Amdahl's law is not at all suited to understand such non-linear non-local inter-dependencies.
- dahfizz 5y agoThe key, I think, is that everyone else is parsing the same JSON from the same web api. If you can do it faster than them, its still an advantage.
- andylynch 5y agoI started reading this thinking it would be about FIX tag/value. Now I’m more than a little shocked multiple someones are publishing real time data in JSON for algos?? I thought this was mostly a solved problem with OUCH, ITCH, FIX SBE and friends. But great to see this writeup, getting these details right can make a big difference.
- Bootvis 5y agoJSON is common in crypto what Cantor appears to be trading.
- gfd 5y agoDoes crypto allow colocation like stock trading? If it's gonna be some websocket endpoint, what's the point of optimizing microseconds?
- lordnacho 5y agoThis is a good point. You expect more jitter on a public network. My guess is they've come from traditional finance, where performance porn is just irresistible. Some exchanges do offer colo.
- Bootvis 5y agoIt kind of does. Most exchanges are hosted in the cloud and I don't think AWS offers colocation similar to the traditional exchanges. I bet Cantor ensures that the machines they are using are close to those used by the exchanges. HFT on a traditional exchange will be faster but that's not the competition. The competition in crypto faces the same problems so you just need to be faster than them. Of course, if the whole process has too much uncontrollable noise (jitter) due to cloud specific reasons it probably doesn't matter. I hope they managed to control this before doing this optimization :)
- ajoseps 5y agoone approach I've heard of some places using is to figure out where the exchange is hosted, spin up multiple instances to test the latency to the exchange, then choose the lowest latency instance. Not sure how often one would need to redo this process though.
- loeg 5y agoLovely concrete and succinct example of a variety of strategies to attack performance problems. Very cool.
- lordnacho 5y agoAwesome write-up. I love that they're so open about sharing the speed improvement with the open source community. I've worked in several trading firms where it's just not thinkable to do something like that.
- SloopJon 5y agoThe rust_decimal crate seems to use C#-style decimals, which always struck me as wasteful. Is there any particular advantage to that format over IEEE decimal floating point?
- pclmulqdq 5y agoWithout hardware support, IEEE 754 decimal floating point can have some very sharp edges. I'm sure that the C# decimal format avoids these by spending more space. I'm personally surprised that this firm isn't using "integer # of exchange ticks" as their storage format for prices...
- cantortrading 5y agoInteger # of exchange ticks would be nice, but that covers a very wide integer range across a lot of products and changes without warning (not to mention, can be inconsistently applied...). Using base-10 decimals is a good compromise between the perfect solution and something that 'just works'. It also means we don't have to go rewrite everything and instead can just do some optimization work when needed.
- dhosek 5y agoI find it fascinating that something like parsing decimals is not a long-solved problem (and kind of connected was the recent post on emulating the IBM system 360 where decimal parsing was baked into the hardware.
- pvg 5y agoI think it is, for the most part, a long-solved problem and both your example and the article outline it, in a somewhat circuitous and circumscribed way - it's a combination of 'make computers so much faster, it doesn't matter' and 'don't use decimals'. This writeup is about a very specialized case in which they're both using decimals (down to the internal representation) and it's not quite fast enough for their purposes. Every x86 computer, incidentally, still carries vestigial hardware support for BCD. https://stackoverflow.com/questions/33182491/why-bcd-instructions-were-removed-in-amd64 https://stackoverflow.com/questions/33182491/why-bcd-instruc...
- deleted 5y ago[deleted]
- pmontra 5y agoGreat article but I've got one question. If they have to read the assembly created by the compiler to understand how to write Rust code that the compiler could optimize, why didn't they directly write assembly code?
- dtgriscom 5y agoPortability?
- shadowofneptune 5y agoIt's much easier to read assembly code than to write it. You're only looking for the necessary performance in a few hot-spots. Everywhere else the convenience of a high-level language is welcome. With the tail-call optimization, it's really only a fact of how programming languages are implemented that you need to look at the assembly. In a language that guarantees tail calls, or has a special tail-call keyword, you could forget about what's going on at the instruction level.
- dspillett 5y agoPortability: one you hand-code assembler for an architecture you lose compatibility for others. Unless you provide multiple code paths to account for this, but that increases testing and maintenance requirements. Future optimisations using better on-chip features, or other new techniques: new instructions that reduce the task might get used by the compiler but it won't update your section in hand-coded assembler. Maintainability: you don't want to force larger parts of your workforce/community to be familiar with assembly if it is not useful day-to-day, or lock your assembly capable people to maintaining those hand-written-assembly portions.
- pmontra 5y agoUnderstood. About maintainability I guess that that code must be marked as do not modify, or the compiler will walk out of the optimization path. Future developers on that project must know very well why that code is written like that. Furthermore any new compiler or CPU could break the optimizations.