3 ms·
This makes me wonder how, in mission critical and real time systems like LSE, they maintain a balance between being ACID and fast. Does any of these transaction
by ashish01 16y ago
This makes me wonder how, in mission critical and real time systems like LSE, they maintain a balance between being ACID and fast. Does any of these transactions ever touch the disk ? If not, how do they handle machine failures ?
- SpikeGronim 16y agoI think systems like this use a lot of RAM. You can afford it when you're a stock exchange. You can get high availability from a RAM store if you replicate it over multiple machines and react quickly to machine failures. Then you keep a commit log on disk that's append only and avoids seeking.
- wglb 16y agoIf there is a database, it would be far down the chain. Imagine machines with 72g of memory whose only function is a buffer. (I would love to get my hands on one of those suckers and outfit it with Redis.)
- gxti 16y agoNo, the core of a stock exchange resides entirely in RAM. Everything is reset once a day (overnight for stocks, immediately after regular hours for futures), and brokerages are responsible for resubmitting long-lived "good 'til cancelled" orders each day before the session begins. Despite being RAM bound it's easy to parallelize because each stock is its own isolated exchange, so they can be distributed across hardware in proportion to the average volume in each stock. In other words, it's localized but embarrassingly parallel. The only traditional databases are for reporting purposes e.g. who traded with whom and are append-only as far as the core exchange code is concerned. The reporting database is coupled through the same firehose feed that you can get as an exchange member, albeit with more redundancy. There are also access control systems that only come into play when you first open a session, and lots of other moving parts each with their own task. For example, if you lose your connection to the exchange you can ask a certain (non-core) server to replay all the events that happened from a given point in time in order to catch up. These, too, would be listening to the main event stream and spooling to disk, but the central exchange processes are RAM-only. How do they handle machine failures? That's harder to speculate on from the outside, but if I were them I'd be running the same exchange in parallel on 2-3 machines. I don't know how it could be done without serializing the incoming order flow to make sure that each machine sees the same order of events, but it's not impossible.