12 ms·
The $440M software error at Knight Capital (2019)
- bfm 4y agoThe OP details how poor software engineering practices brought down a 1.4B market marker with 1400 employees in 2012. Some of the issues mentioned include: - Keeping synthetic test data generation as part of a production build. - Keeping dead code for years. - Re-purposing a feature flag. - Refactoring without regression tests. - Manual deployments without peer reviews. They forgot to update one of their servers with the new code. - Automated alerts sent via email were ignored. - Rolled back to a version of the code running on the server they forgot to update, making things worse. - Rushing out a release without proper software engineering hygiene. The article suggests improvements that could have prevented the chain of events. For those here who are in HFT circles, have things improved after the Knight Capital Group debacle? edit: formatting
- idohft 4y agoHard to speak for HFT in general. Like in software, different firms have different levels of hygiene. About half of your bullet points were true of my previous employer, at my time of leaving.
- rebelos 4y agoSome of this is unforgivable, but reflecting on it I also realized that software engineering at quant firms has an almost impossible mandate. You want something akin to the extreme rigor of mission critical software (airplanes, cars, NASA, etc), while also remaining nimble enough to modify strategies as market conditions rapidly evolve.
- SilasX 4y agoSame is true for blockchain smartcontracts, which have similar catastrophic consequences.
- ChrisClark 4y agoThat truly is scary to me. I can easily* write advanced Solidity and could try to make something big. But I won't, because I know I would not be able to handle the stress and responsibility. One tiny logic error and millions lost. Thanks but no thanks. *The fact I believe I could easily do it is probably exactly why I'd end up making some huge mistake. ;)
- wombatpm 4y agoI’m sure that any big contracts are written only by CMM Level 5 organizations using formal methods and provably correct with certification that less than 1 line in a million contains an error. It’s all spelled out in the SOW attached to the RFP.
- astrange 4y agoI don't think I could, since the language was designed by amateurs after no research into safe programming, and is only being improved one gigantic loss at a time. Maybe if you could translate from Coq.
- nradov 4y agoWhy unforgivable? It's only numbers in an account. No one died.
- bfm 4y agoIt is challenging, although, with financial markets, it seems like it would be simpler to have some automatic anomaly detection mechanism to unplug or slow things down to prevent further damage.
- WJW 4y agoThere are a lot of preventative measures they could have taken, starting with just not leaving in dead code and paying attention to automated alerting. But the moral of the story is that they got away with it for so long that nobody cared about it anymore. After all, if it were truly a big deal why hadn't it broken years earlier. Then when the technical debt finally got called it bankrupted the entire firm in one go. Most of us (hopefully) have less devastating technical debt to deal with, but it is still a cautionary tale about what could happen if you ignore it for too long.
- posterboy 4y agoThat's a weird statement. The extreme rigor on the one hand seems to require a value judgement of the real benefits to HTF that I'm not willing to make. The remaining nimble'ity, on the other hand, is an odd word to use over agility or old fashioned responsibility. The benefit is proportional to it, but not exclusively. The rapidly evolving market conditions concern regular trade too. Swift reactions are expected in any other systems application. "almost impossible" is a weasel word. It's almost impossible to win except for the last man standing, is that it? And there's no practical upper limit to nimble'y, though conservative estimates indicate that less work is more. What's missing is the perverse incentives, corrupt policies, sociopathic leadership, ...
- deleted 4y ago[deleted]
- kevstev 4y agoI worked in algo trading for years, eventually got out because quite frankly the level of risk I was carrying on my shoulders everyday for what I was being paid were just way out of whack, I at least personally never got the huge pay days that people talked about until after I left finance for more pure tech. Interestingly, I worked at Knight and my team pioneered trying to blow up the firm, but that was in 2004, and things were much friendlier- instead of front page news, it was a small blurb on page 3 of the markets section of the WSJ. Anyway, I still have friends in that business. It hasn't really changed, they have too few people covering systems that are quite complex and while there are checks and such, no one really understands things entirely from end to end in detail that can prevent all problems. I will never invest directly in an investment bank- either through carelessness or maliciousness I could have easily caused a 9 figure loss, if not more, and there were probably a thousand other people in the same position. When I read the detailed writeup around this a few years back, I think by far the biggest issue was reusing a tag that had been previously used to denote which strategy to use. I understand why they may have chosen to do so, at the Big Bank I was working at, getting a new fix tag to be passed through all the layers properly would involve at least two other teams and coordinating releases and probably several weeks worth of meetings. If you just reuse an old value you can avoid all that since everything is already set up.
- fullsend 4y agoI appreciate your comment about pay. Recruiters will often tell me "it's finance so of course the pay will be substantial." Then when we get to talking numbers they're like "300k a year". Oh, you mean the going rate at a FAANG? And I have to move to New York or Chicago, work more hours, and actively work for people who I know are taking home paychecks with 7+ zeroes on them? Come on. Sometimes it's 400 plus bonus or whatever, which is based on fund performance and yada yada. But it feels way off. I had heard so much about the staggering paydays at these places but it seems you need an ML PHD or some trading chops to be part of that.
- gjs278 4y ago
- 4y ago
- bnastic 4y agoI remember the Knight Cap event, I was working on order routing at the time. Things have changed a lot since 2012, and at the same time haven’t. Circuit breakers and position monitoring are no.1 in any sane market making firm. What happened then I can’t imagine happening now (accumulating a huge position for, what was it, 30 minutes? With nobody killing the algos within a couple of minutes?). On the other hand, the perfect world of “code hygiene” and 100% test coverage will never exist in this world, things will slip and they do frequently. What’s better, externally, is the availability of good tools for development and change reviews (bitbucket taking hold, for example), automated deployments, containers, testing frameworks and similar. This type of software, end to end, is incredibly complex and difficult to reason about when unexpected happens (there was a TTL misconfig for multicast and we never got such and such update? Well, no one thought of that!), esp these days with the influx of ML algos for price generation.
- benjaminwootton 4y agoI worked in a lot of front office groups in investment banking. The short spell I did in HFT had great software development and DevOps practices.
- aledalgrande 4y agoThis is all basic stuff I look to set up in every team, and it's crazy given how these firms work directly with tons of money that they don't have an even higher standard. Guess I wasn't wrong turning down these roles.
- pclmulqdq 4y agoI used to work in HFT. I have seen highly variable practices in this case, including a "mini-knight" incident in the single-digit millions due to tech debt and poor test coverage. However, the most useful change that has resulted from the KCG debacle was adding several layers of kill switches, a dedicated ops team to watch trading and flip the kill switches, and embracing devops automation. There is a much more serious focus on having a defense in depth, and making sure that problems like this are noticed before they become an issue. Rollbacks are no longer the first action when something goes wrong: the kill switch comes first. Dead code, tech debt, repurposed flags, and spotty test coverage are everywhere still.
- commandlinefan 4y ago> poor test coverage Yet you don't have to hang around here long to be told that "Unit Testing is Overrated": https://tyrrrz.me/blog/unit-testing-is-overrated https://tyrrrz.me/blog/unit-testing-is-overrated
- aaronharnly 4y agoI’m curious about the “repurposed flags” part. I wouldn’t think of flags as expensive / effortful to make more of, but clearly they must be if people are tempted to reuse them. Can you help me understand what is meant by a flag in this context, and why it would be repurposed?
- isogon 4y agoRepurposing flags not always well-motivated, but one legitimate reason to do this is the memory (and particularly cache) footprint. Often flags are local to a particular object. If there are lots of such objects, you want each to take as little space as possible. You should check out the contortions linux devs go through to make struct page small [0]. This is important, because there is one such struct per page of physical memory. The memory use is a near-constant percentage of your total memory, and you wouldn't want it to be any larger than necessary. Even when there are not a lot of these objects, in low-latency software it's important to hit the cache. Your program should always just be as compact in memory as possible. Semantically flags are booleans (is proposition P true of this object). They are stored compactly as bitsets, often implicitly, say: #define FLAG_1 0x01 #define FLAG_2 0x02 /* ... */ #define FLAG_8 0x80 struct order { u32 qty; u16 id; u8 type; u8 flags; }; This struct will fit into 8 bytes. This is great, as you probably won't waste space to alignment in many cases -- 8 is a good multiple. But if you wanted to add FLAG_9 here, your flags would become a u16, and your struct would, frustratingly, stop fitting into 8 bytes. To avoid this, one might repurpose flags. Another example of this is intrustive flagging, using, for example, the high or low bits of a pointer aligned to 2^n bytes. If you run out of bits there, not much you can do. [0] https://github.com/torvalds/linux/blob/master/include/linux/mm_types.h#L72 https://github.com/torvalds/linux/blob/master/include/linux/...
- 22SAS 4y agoCurrently work at an HFT firm. Most of the firms invest well into good DevOps, Trading Systems and SRE teams, to ensure that everything from installing a trading server at the colocation facility, to CI/CD and making changes to the systems configs, is done well. There are also guards in place to ensure that if the system seems to make trades that are way too odd then pull the plug and go down immediately. Also, any code that does not need to be there, is promptly removed right away. Where I work at, we have a few people from KCG i.e what was formed after Knight Capital merged with GETCO, after this incident. Sometimes this incident is bought up, although none of them I think ever worked for Knight Capital before this incident.
- bob1029 4y agoRepurposing feature flags is some kind of next dimension horror for me. We've got quite a few of these to deal with, and if someone started changing what they mean we'd be fucked super fast. Simply suggesting that we alter the meaning of an existing FF would result in the resignation of a non-zero number of project managers on my team. Rolling back code is another thing I have no tolerance for anymore. The only option we entertain these days is a roll-forward. If your software takes so long to iterate/build that you need to go back to and old version in an emergency, you need to review your languages/tools/frameworks/processes. We maintain a contractual obligation to our customers for same-day code updates (in cases of production/regulatory emergencies) because we have enough confidence in our processes.
- sokoloff 4y agoYou’re likely using human-readable names for the flags and shipping a multi-MB payload of JS and JSON. An HFT firm is likely bit-packing flags so they can send an 8 byte payload rather than 10 and might be using an FPGA hanging right off the PHY to figure out “is this message even interesting to me?” Your feature flag might be “SHOW_STRIKETHROUGH_PRICING_ON_CROSSSELL_OFFERS”; theirs is a bit mask macro to pick off the 5th bit from the 7th byte. (Why do they care? Because if they allow themselves to get fat and slow, a competitor will take the money.) Roll-forward only, same day SLA is probably right for your business, but isn’t for a company that could have their systems dusting off $1M every handful of seconds that bad code is running. Different business problems call for different technical approaches. You should no more adopt theirs than they should yours.
- hn_go_brrrrr 4y agoRolling back isn't about avoiding rebuilds, it's about restoring to a known-good state. Making an emergency patch is typically far riskier than going back to last week's build. We always favor rollbacks, unless there were critical fixes in the current release we absolutely cannot afford to lose.
- jll29 4y agoThe only thing worse that comes to mind than repurposing flags is re-using UUIDs, which I have seen some production DB do (no, not mine, thank you very much!). Reusing the RAM occupied by flags might be do-able in a clean way using guards like the structured enum in Rust, which permits unions (objects that occupy the same space) that always have the right type (i.e., the compiler has knowledge what is in there at each point in time). This mechanism could in theory be extended beyond type-safety to accommodate other contexts in systems programming use cases where memory usage is extremely important.
- hftthrowaway22 4y ago
- akhmatova 4y agoRushing out a release without proper software engineering hygiene. Sounds familiar. From everything listed above, it would sounds like this must have been yet another one of those "Just F-ing push the change the server now, dweeb, the traders are going crazy" environments. That is to say: the real problems were most likely cultural, and not about the sum of a certain set of bad practices.
- randomhodler84 4y agoBack in the day $440M loss due to coding error was a landmark warning case. How could this happen?? In 2021 alone something like $10B was lost due to bugs in defi land. Something about the worst possible thing could happen tends to happen eventually and it gets worse every passing year.
- nikanj 4y agoThose $440M were lost by rich people who had invested in a hedge fund, not poor people who bought crypto lottery tickets in the hopes of getting rich quick
- quickthrower2 4y agoAlso a fuck load of money had been printed in the mean time. https://fred.stlouisfed.org/series/M2SL https://fred.stlouisfed.org/series/M2SL
- pingeroo 4y agoWas just about to comment along these lines. If I read about this a few years ago I would be shocked. Now after seeing so many flubs in the crypto space, my reaction is just 'meh'
- vmception 4y agoI actually always think of the Knight case and similar ones when people see a DeFi organization have an issue and extrapolate that to an issue with the entire DeFi concept. Its so obvious that those people have no clue whats going on in the markets they respect. Truth be told, many of them dont like markets at all. So its just a lack of exposure and compounded ignorance.
- colechristensen 4y agoMany traditional finance issues are fixable though, there are many more errors which don’t become big stories because they are reasonably reversed as only minor inconveniences.
- gzer0 4y agoThe incident happened after a technician forgot to copy the new Retail Liquidity Program (RLP) code to one of the eight SMARS computer servers, which was Knight's automated routing system for equity orders. RLP code repurposed a flag that was formerly used to activate an old function known as 'Power Peg'. Power Peg was designed to move stock prices higher and lower in order to verify the behavior of trading algorithms in a controlled environment. Therefore, orders sent with the repurposed flag to the eighth server triggered the defective Power Peg code still present on that server [1] > Power Peg was designed to move stock prices higher and lower in order to verify the behavior of trading algorithms in a controlled environment. This is insane. Make one wonder, what is or isn't actually being deployed in prod in 2022. [1] https://en.wikipedia.org/wiki/Knight_Capital_Group#2012_stock_trading_disruption https://en.wikipedia.org/wiki/Knight_Capital_Group#2012_stoc...
- codeulike 4y agocoder running down corridor to the trading room, bumping past people and sending sheaves of papers flying "Power Peg has triggered! Tell them Power Peg has triggered!"
- civilized 4y agoSo... you're saying the company activated a tool called "Power Peg"... and.... they got f***ed? Well, what did they expect?
- bananaquant 4y agoYep. They got... power pegged.
- bovermyer 4y agoInteresting. Five years prior, this story was posted on this blog: https://dougseven.com/2014/04/17/knightmare-a-devops-cautionary-tale/ https://dougseven.com/2014/04/17/knightmare-a-devops-caution...
- NovemberWhiskey 4y ago(2019)
- bfm 4y agoUpdated the title
- user3939382 4y agoHere’s a 225 million dollar oopsie from 2005 https://www.foxnews.com/story/typing-error-causes-225m-loss-at-tokyo-stock-exchange https://www.foxnews.com/story/typing-error-causes-225m-loss-...
- bfm 4y agoToday there was a 300B oopsie in Europe caused by a Citibank "glitch" bloomberg.com/news/articles/2022-05-02/citi-s-london-trading-desk-behind-rare-european-flash-crash
- chmod775 4y agoIt only was a sudden drop in share prices, which quickly rebounded. The amount of money that actually changed hands due to that mistake will be tiny in comparison.
- nuclearnice1 4y agoHere’s the same oops in 2001 https://www.wsj.com/articles/SB1007117680496415760 https://www.wsj.com/articles/SB1007117680496415760
- inter_netuser 4y agoPeanuts, just a regular day in DeFi. https://rekt.news https://rekt.news
- robofanatic 4y agoat the end .. its just money going from one account to another right? Its not like some physical thing that has perished and cant be brought back. Why is it difficult to reverse the transactions?
- ceejayoz 4y agoBecause those transactions cause other transactions, which cause others, and so on and so forth. You'd have to reset the market for the day. Imagine how pissed you'd be if you made money off Knight's mistake and it all just disappeared the next day.
- strgcmc 4y agoExcept, well, cancelling transactions obviously does happen, sometimes: https://www.reuters.com/business/lme-suspends-nickel-trading-day-after-prices-see-record-run-2022-03-08/ https://www.reuters.com/business/lme-suspends-nickel-trading... Knight was probably too messy to rollback cleanly, but that just means it's a matter of cost/complexity/politics... if you're a big enough player, then the exchange will do you favors, like in the LME case. Free markets, lol
- rubyskills 4y agoThis is much easier to do in a centralized futures market. I can't imagine a rollback in stocks being easy or possible.
- bfm 4y agoFrom the OP Rules were established after the “flash crash” of May 2010 to govern when trades should be canceled. Knight’s buying binge did not drive up the price of the purchased stocks by more than 30 percent, the cancellation threshold, except for six stocks. Those transactions were reversed. In the other cases, the trades stood.
- ceejayoz 4y ago> The LME announced that all trades will be voided from midnight until 8:15 a.m. on Tuesday when trading stopped and added that it was considering a closure of several days. > "People will be asking if this really a functioning market... This is meant to be a market of last resort and people can't get inventories to deliver against positions," said Colin Hamilton, managing director of commodities research at BMO Capital Markets. There’s gonna be a pretty high threshold for this sort of thing. Higher than “one company fucked up and wants a do-over”.
- dang 4y agoRelated: Knight Capital Says Trading Glitch Cost It $440 Million - https://news.ycombinator.com/item?id=4329101 https://news.ycombinator.com/item?id=4329101 - Aug 2012 (90 comments)
- Terry_Roll 4y agoI read this >Under stock exchange rules, Knight would have been required to pay for those shares three days later. However, there was no way it could pay, since the trades were unintentional and had no source of funds behind them. The only alternatives were to try to have the trades canceled, or to sell the newly acquired shares the same day. And then I understand why /r/WallStreeBets and /r/Antiwork is gaining traction. All it takes is a bit of organisation and the adoption of Govt tactics and practices which is ultimately violence and then just maybe you might see a Govt that works for the people and not the criminals, but I cant picture Bernie Sanders wielding a pitchfork! Still I see Musk was market making with his tweet. I dont think you can be any more blatant! LOL https://twitter.com/elonmusk/status/1520650036865949696?cxt=HHwWgICjsdvJt5oqAAAA https://twitter.com/elonmusk/status/1520650036865949696?cxt=...
- throwyawayyyy 4y agoRandom, but I interviewed at Knight Capital for a software engineering position a few weeks before this all went down. I was in London, so the interview was done over the phone. Picture me in the evening, handwriting C to solve some problem (the fog of time too thick to remember what that problem was), then reading out what I'd written, semicolons and all, to the interviewer. Because of course there was no shared doc. I did very badly. But then, so did they.
- nogridbag 4y agoSame! Although I interviewed in person in their NYC office. I was very junior at the time and the team I interviewed with was awesome. I (luckily) didn't get the job. I did a few more interviews and accepted an offer from another company where I met my wife!
- Bilal_io 4y agoDo you return the wife when you leave the job? In all seriousness, you dodged a bullet. Do you remember what problem or set of problems they asked you to solve? Also, what made the interviewing team awesome in your opinion?
- nogridbag 4y agoThe interview was incredibly basic compared to modern FANG style full day interviews. The coding portion was basically a fizzbuzz question. They had a one page printout of a Java method and I needed to give the output. I answered it correctly nearly immediately as it was simply testing pass by reference vs pass by value. The main interview was all verbal q&a core java questions.
- MikeDelta 4y ago> and accepted an offer from another company where I met my wife! That really cheered me up this morning, I met my wife at work as well.
- coolhoody 4y ago
- rewsiffer 4y agoWow, I just started a new position at a tech company in the NYC financial sector, and this was part of today’s training!
- paulpauper 4y agosorta like a precursor to a smart contract hack
- photochemsyn 4y agoReally good write-up. Perhaps there's a dawning realization that the model of 'move fast and break things' is fatally flawed? > "Knight’s IT project managers and CIO should have pushed back on the hyper-aggressive delivery schedule... Thirty days to implement, test, and deploy major changes to an algorithmic trading system that is used to make markets daily worth billions of dollars is impulsive, naive, and reckless." The fact that since 2008, the portion of all stock trades in the U.S. taking place away from public markets has risen from 15 percent to more than 40 percent is also kind of astonishing. It's long past time to re-erect the walls between commercial and investment banking.
- jll29 4y agoExcellent report of a fault case study. Yes, check lists help, not just in medicine, but also in aviation. But automating things is of course the best option.
- jmyeet 4y agoWhat's amazing to me is that there are many software engineers here who can recognize how errors like this can so easily creep into software (some would say they're inevitable in any sufficiently complex software) but they still somehow think immutable smart contracts on blockchains are somehow still the future. Crazy.
- killingtime74 4y agoThey know but they don’t care. Crypto is get rich quick
- sonicggg 4y agoMy brother worked in a medium-sized fintech not long ago. This example is shown in their onboarding process. I guess the same goes for other companies out there.
- toddm 4y agoImportant to add that this occurred in August 2012 and not in 2019 as the title implies. I had interviewed there in my garden year (2010-2011) and was ultimately not considered for a role as a high-frequency quant.
- bmitc 4y agoWhat is a garden year?
- toddm 4y agoIt's paid time off from a financial entity (hedge fund, prop shop, etc.) where you are paid not to compete in their space. In my case, I had a year to sit-out as a HFT quant from 2010-2011. Note that it sounds great - you're getting paid not to work - but nobody likes to make you an offer, as it's non-binding and your perceived/real value is declining as the time goes by.
- bibliographie 4y agoHey on that topic of a non compete, I've recently been offered a different internal role that would force me to sign a non compete. What was your experience like interviewing/searching for new roles? I would also like to hear what your general opinion on what it make it worth it to sign a non compete, although it sounds like you a generally negative view of them from this comment.
- anonu 4y agoMy small footnote to this story: we had the look on the "blind bid" to purchase the portfolio of erroneous trades on the program trading desk I worked on. Ultimately we didn't bid (thank you risk department) and it traded away to Goldman as the article correctly reports. Would've been a great trade though. I estimated it netted Goldman $2m+
- faangiq 4y agoNo lessons have been learned. There is no catharsis. This blog post has meant nothing.
- anonu 4y agoOriginal Post-Mortem: https://www.sec.gov/litigation/admin/2013/34-70694.pdf https://www.sec.gov/litigation/admin/2013/34-70694.pdf
- alchemist1e9 4y agoThere are several more of these of approximately the same order of magnitude but those firms are able to keep it successfully out of the media.
- geranim0 4y agoWho gets to deposit the bribes?
- eek2121 4y agoFascinating read. I love code horror stories such as this.
- blobbers 4y agoMy experience in quant finance is that strong engineering teams are viewed as cost centers and are not well compensated the way they are at “real” tech companies, especially when the firm isn’t doing well. Management foolishly tends to only value seeking alpha, rather than reliability of existing alpha. Unsurprisingly poor management tends not to realize when their software is built on a house of cards. Definitely saw bugs cause large losses (8 digit numbers).
- smabie 4y agoFeel like this way more of the case in the past than it is now. Engineers at top prop shops today are compensated more much generously than at all but the most senior engineers at a FAANG. For example, word is that new grad hires at Jane Street are getting first year guarantees at around the 450k range, which is much higher than an entry level engineer at a FAANG would make. Culturally, there's some truth in what you say though I think: alpha generators are probably more respected and are pulling more TC in good years.
- linuxhansl 4y ago> However, a dark cloud remains: market data suggests that 70 percent of U.S. equity trading is now executed by high-frequency trading firms Is it just me, or was this one of the scarier statements in the article? High frequency trading does not seem to add anything, it's not a market instrument to balance anything. It's more like gambling with the odds slightly in your favor. Oh course, I'm a laymen when it comes to high frequency trading and might totally wrong.
- smabie 4y agoMarket makers provide cheap liquidity. Or in other words, compare the total transaction costs (bid-ask spread + commission) you would to pay today to trade 100 shares of stock today vs 30 years ago. The difference is staggering.
- jspaw 4y agoHindsight is a helluva drug. The SEC report cannot be viewed as a “postmortem.” https://www.kitchensoap.com/2013/10/29/counterfactuals-knight-capital/ https://www.kitchensoap.com/2013/10/29/counterfactuals-knigh...