4 ms·
> the flag word was out of new bits for flags, so an engineer reused a bit from a deprecated flag Plus some silent update failures meant that new-feature order
by jjoonathan 3y ago
> the flag word was out of new bits for flags, so an engineer reused a bit from a deprecated flag
Plus some silent update failures meant that new-feature orders sent to out-of-date servers transformed into old-feature orders. Boom!
> Adding risk checks to the last stage of an order’s life became universal in the industry
I wonder what these "fast-twitch" sanity checks / circuit breakers look like. Whenever I try to model risk, things get complicated quickly -- but presumably simple heuristics must exist if they became universal in the industry.
- kmeisthax 3y agoIt could be as simple as... - Do we have accounting for what trading strategy generated this order? - Will this order immediately lose us money? (e.g. are we buying out-of-the-money options) - Did we accidentally set the PhysicallyDeliver flag? - Have we hit our organization's margin limits? Any situation that no reasonable trading strategy would put you in, or that would otherwise be outright illegal, is a good thing to put in risk checks for.
- twic 3y agoWhat became universal in the industry is an item on the checklist saying that you have safety checks to prevent this. How seriously that item is taken, i suspect, varies quite a bit. In my experience, you aim for multiple extremely simple checks, with minimal logic and minimal calculation, so you can have confidence they won't have surprising behaviour in an unusual situation like this. The classic example is an order count limit - initialise a counter to some value at startup, and every time the machine sends an order, try to decrement it. When it hits zero, it can't send orders any more. Just throw an exception or return early or something. You display the value of the counter to human operators, and give them a button which resets it to the initial value. In normal operation, you are sending orders at a steady trickle, and humans will have to press the button every now and then. If something goes insanely wrong, as here, the counter will run down quickly, and then the humans hopefully won't push the button, because something is obviously wrong. It's a very crude safety, but it is a simple one. Another is a limit on message rate. You could use a token bucket filter. Does not affect normal operation, but stops a machine which is spraying out excessive orders. You could have it so that if the bucket runs out, it turns off until a human explicitly turns it back on. You have limits on net position too, to stop you running up huge positions in anything, but those are higher-level, and not quite the same kind of last-ditch safety check. I don't really know that either of these would have helped in Knight Capital's situation, because the precise mechanics of the "power peg" aren't clear. It sounds like a kind of explicitly-managed iceberg order, which these safeties would have caught. But another writeup [1] says it was a testing tool, not intended to be used on a real exchange at all, in which case who knows. [1] https://www.henricodolfing.com/2019/06/project-failure-case-study-knight-capital.html https://www.henricodolfing.com/2019/06/project-failure-case-...
- pclmulqdq 3y ago(Author) As far as I know, Power Peg was indeed intended to essentially be a manual iceberg order from the time before that was an order type on the exchange (with slightly different semantics). Rereading the source you quoted, it definitely wasn't a "buy high sell low" system, even if it was never used in prod.
- pgwhalen 3y ago> I wonder what these "fast-twitch" sanity checks / circuit breakers look like. Whenever I try to model risk, things get complicated quickly -- but presumably simple heuristics must exist if they became universal in the industry. You're right, they are very simple. Think things like orders per second, quantity of order, price of order, notional ordered over time, etc. You basically want to ensure things aren't "too big" or "too fast" as simply as possible. Other types of risk (portfolio risk, greek risk, etc.) are handled in different ways, upstream of these final checks.
- KMag 3y agoI was a Core Strat for Goldman's Algorithmic Trading Platform (ATP) at the time of the Knight collapse. ATP had an "upstream compliance layer" and a "downstream compliance layer" even prior to this incident. > I wonder what these "fast-twitch" sanity checks / circuit breakers look like. Leaving out any company secrets, things were structured basically how you'd expect from any software architecture design class assignment. Specific components perform sanity checks at specific levels. Some components keep track of order state. Other components making trading decisions. Still other components are tasked with getting data from place to place and translating message formats. The "compliance" rules were actually a mixture of exchange regulations, extra constraints from the Compliance department, and other sanity checks. Basically, most of the sanity checks were called "compliance rules", and from an engineering point of view it made sense to treat all of the sanity rules the same, regardless of which entity came up with the rule. ATP is a framework for execution algos: some other algorithm or person makes the big-picture decisions for big orders, and ATP takes those big-picture "parent" orders and breaks them up into smaller "child" orders at various price points at various times based on various parameters/hints annotated on the parent orders. The other key components of the trading system are the market data feeds, the exchange/venue connectors, the Smart Order Router (SOR) and the Order Management System (OMS). The OMS keeps track of the parent order state and the relationship between parent and child orders. The OMS communicates the child orders to the SOR, which (unless the child order is annotated explicitly with an exchange) distributes the orders across the various trading venues. The exchange/venue connectors allow the SOR to communicate with the exchange. Some places combine several of these components into single processes, but it's a pretty typical execution architecture. I've heard it's pretty common to either have the OMS and execution algo engine in the same process, or else have them communicate via shared memory. If I were designing a system from scratch, I'd probably have a TWAP-only algo engine using shared memory to communicate with the OMS, and then build all other execution algos (TWAP, Arrival, etc.) from TWAP orders. For instance a parent order might be for "Buy 10,000 of 1299.HK (the RIC for AIA LTD.), limit price 72.00, get done as close to the current market price as possible, finish by 15:00:00 but never trade more than 10% of the market volume in 1299.HK". An ATP engine subscribes to notifications for state changes for all orders tagged in the OMS with its particular EngineID. Leaving out the details of how the order gets to the OMS and how the parent orders get tagged for a particular engine, the ATP engine sees new parent orders assigned to it. The Upstream Compliance Layer then performs sanity checks (including some fat-finger checks for limit prices too far from the previous day's adjusted closing price) and either accepts responsibility for executing the parent order and tells the OMS to change Status from Pending to Accepted, or else tells the OMS to change status from Pending to Rejected (along with a short text description of the rejection reason). Assuming the parent order is accepted, the OMS sends a message to the client (via a series of upstream systems handling client connectivity) informing them that the order is accepted. In the case of targeting arrival price, it has a partial differential equation model of price impact, and it solves this PDE to minimize total estimated price, and decides to split out the first child order(s) to the exchange. Even though the parent order is BUY 10,000 1299.HK @ 72.00, maybe the model indicates it's optimal to split out BUY 300 1299.HK @ 68.70 and BUY 100 1299.HK @ 70.10 and wait for the rest. Within ATP, the Downstream Compliance Layer performs sanity checks on the proposed child order(s). One of the basic checks is that the executed quantity plus the total quantity across all outstanding child orders won't exceed the quantity of the parent order. In the case of Hong Kong, this includes checks that the child orders are within the exchange's circuit breaker limits for up/down percentage from the previous day's close, checks that the execution algo doesn't have excessive numbers of child orders in the market, rate-limiting child order creation, etc. (In Hong Kong, exchange connectivity is priced by the transaction-per-second, so tons of one-lot child orders eat up tons of transactions from your quota if you're re-pricing or cancelling them. One of the large multinational banks had an algo go crazy with tons of small mispriced child orders in Hong Kong sometime around 2010-2012, and due to the number of transactions-per-second they had purchased, it took them over half an hour to get everything cancelled, bleeding money the whole time.) If the proposed child orders pass the Downstream Compliance Layer, then ATP sends messages to the OMS to create the proposed child orders. The OMS then performs its own sanity checks, including again checking that the total quantity across the parent's child orders plus the already executed quantity doesn't exceed the parent order quantity. If the OMS's sanity checks pass, then the child orders are created and the SOR is notified. ATP then waits to see at least 1,000 shares of 1299.HK trade before splitting another 100 share child order (in order to keep to the 10% max participation rate), its plan based on the partial differential equation solution also restricts when and at what prices it places orders. As I remember, in most markets, it's sufficient to perform a short-sell locate in the Upstream Compliance Layer, getting a stock loan for the quantity of the parent order. However, I seem to remember Japanese regulators requiring locate checks in the Downstream Compliance Layer to check every outgoing child order, resulting in some pain to run the Japanese compliance checks quickly. I guess this results in more fair distribution of locates among clients in the case of hard-to-short stocks, but it's a pain for efficient implementation.