10 ms·
Tokyo Stock Exchange Blackout: One Piece of Hardware Took Down a Market
- deleted 6y ago[deleted]
- gautamcgoel 6y agoIt's amazing to me a single hardware failure can lead to this kind of chaos. Wasn't there any redundancy? Didn't anyone forsee this possibility?
- bunnie 6y agoFrom the article: "When the error happened, the system should have carried out what’s called a failover -- an automatic switching to the No. 2 device. But for reasons the exchange’s executives couldn’t explain, that process also failed. That had a knock-on effect on servers called information distribution gateways that are meant to send market information to traders."
- seppel 6y agocouldn't explain = didn't want to :) But anyway, it also shows that the failover was not properly tested.
- rightbyte 6y agoOr that the fault killing the first on also killed the failover? Like a rampage JIS encoding control character running over the wires.
- marcan_42 6y agoIt was a memory error. Hardware failure. The failover process is what went majorly wrong here. It seems they had a press conference with more details in Japanese.
- karmakaze 6y agoAutomatic failover is like data backups, if not regularly tested it's like it doesn't exist.
- vmception 6y agoThats a bit extreme in my experience, I would say its just closer to luck
- rwbhn 6y agoI would be very curious to hear about your favorable experience with untested failover systems. My last job we tested our active/standby failover like clockwork every 3 months. It served us very well.
- vmception 6y agoMy experience is that they worked when needed and it was a sigh of relief, and a little pat on the back that we perceived reality correctly when we had the foresight to set them up.
- ncmncm 6y agoAn untested failover system is indistinguishable from a box burning power and doing nothing. On a failure, it might work. But there might not be a failure. Same difference.
- karmakaze 6y agoIt's sort-of a saying in the industry that can be applied to many things. The biggest problem is the people who might know how to make it work are no longer there or have forgotten. If it's automated there's a chance it might work, but if it's slightly off then it's even harder to comprehend and adjust. I've seen backups and failovers not work so many times that it's an amazing surprise when one actually works--usually after laborious manual off-script intervention and invention. I was being somewhat kind to failovers, I have seen smallish backups work on several occasions. An untested disaster recovery is the worst. Edit: think of an unexercised procedure as not "It works for me" but rather "it worked for me once."
- numpad0 6y agoSomeone speculated that since this is a RAM failure, the storage in question might kept crawling with reduced RAM and full workload so failover didn’t kick in(either failed from high load, or didn’t trigger from the fact that faulty RAM is successfully isolated). Sounds plausible.
- toast0 6y agoRAM failures are fun. Maybe 15 years ago, I had one box reboot itself and come back up with only ~ 16 MB out of an expected 1 GB or so; it was pretty overprovisioned and was able to sort of keep up working from swap but was sending alarms on a system that was always quiet (near idle cpu normally, important but low throughput system). More recently, I've had a couple systems with so high a rate of correctable ECC errors that machine check processing ate nearly all the cpu; but it still had enough to pass healthchecks, but not handle requests in a timely fashion, but didn't manage to get an uncorrectable error that would have paniced the machine so failover could work.
- cbhl 6y agoIf you don't regularly test your failover, chances are it will not kick in when the primary fails. Especially if the primary is very reliable. Very common pattern. Ideally, you periodically test your ability to failover. But if it doesn't work, well, there's a chance that you just caused a user-facing outage with your test.
- StillBored 6y agoActive->standby failover is usually that way because it doesn't support active->active. Which likely means that the active->standby system is bolted on with some 3rd party technology that isn't integrated with the application. I've been involved with a number of failover systems where even when it worked there was the possibility that you might hit a _known_ condition that causes the fail-over to fail. Pretty scary knowing the product your working on has a couple critical holes in the fail-over that management papered over, which while rare could happen. A lot of these solutions are the equivalent of pull the power on one machine move the disk to the other and power it on. The assumption being that the storage mirror/replication/etc being used to maintain transnational consistency for the "move the disk" part is actually going to be consistent when that happens.
- AnimalMuppet 6y agoThis happened (or at least, was detected) an hour before trading opened. It should have failed over then. To me, that means that you could validly test your failover an hour before trading opened (or, perhaps wiser, an hour after trading ended). If it doesn't work, you learned without causing a user-facing outage.
- ycombobreaker 6y agoThe system is online, publishing market data and and accepting orders before trading starts. The exchange initially announced a delay in opening, and only later announced staying down for the day. It doesn't sound like they had the option to do what you described. The staff trying to resolve the problem was presumably doing everything they could before giving up for the day.
- 6y ago
- jbm 6y agoI have worked with Fujitsu in Japan before. I could see the following as one way it could happen. There is a large disconnect with large system integrators there and the actual developers / architects; usually 2 or 3 levels of subcontracting. This means there are 2 to 3 levels of intermediaries taking a margin too, and as such, the actual developer doesn't have the fiscal wherewithal to push back on requirements or deadlines too strongly. (How do you save money on a 300k yen salary in Tokyo?) If I was to guess, the requirement was there, but by the time someone technical got into the real nitty gritty, they discovered that the timeline was too tight to effectively do the testing. Instead of pushing back, they just rushed it through with a lot of overtime work.
- smabie 6y agoAlso, though maybe this is unfair, but Japanese culture doesn't work very well for software development. I worked with a couple on a project and they are absolutely unwilling to ask for help, ask questions, or proactively fix problems. Of course this is a very small sample, but I've heard very similar things from a friend who worked at a Japanese conglomerate for a couple years in Japan.
- xvilka 6y agoThey treat mistake/bug/problem as a failure of a developer or engineer, not something to be proactively looking for and taking as a lesson for both engineer and code improvement.
- deleted 6y ago[deleted]
- stanrivers 6y ago"A data device critical to the Tokyo Stock Exchange’s trading system had malfunctioned, and the automatic backup had failed to kick in. It was less than an hour before the system, called Arrowhead, was due to start processing orders in the $6 trillion equity market. Exchange officials could see no solution." You know, just from a human perspective, talk about a bad day. It's like seeing the cruise liner heading for the port too fast, knowing it is going to crash and cause immense damage, and realizing there is absolutely nothing that can now be done to prevent the damage. Except, in this case, it's just way worse.
- tuatoru 6y agoHow many people died? How many were hospitalized?
- gruez 6y agoAre you talking about the cruise ship crashing? It's a hypothetical. Nobody was hurt.
- stanrivers 6y agoFair point! I was talking about monetary damages - excluding any kind of loss of life or personal injury. I could have picked a better metaphor!
- Scoundreller 6y agoI question the monetary losses too. How much was lost by being unable to sell something today that you weren’t willing to sell yesterday?
- outworlder 6y ago> How much was lost by being unable to sell something today that you weren’t willing to sell yesterday? Could be a lot. Markets are not disconnected. If only Japan existed, sure. Some orders would no longer exist, but nothing major, because noone else would be trading anyway. But suppose this was during an economic downturn. You are trying to SELL SELL SELL because all countries worldwide are feeling some economic pressure and you can't because the exchange is down. The next day, whatever assets you had are now worth a fraction of what they were one day before. Oh, and there are some financial instruments that expire.
- Stierlitz 6y ago‘Upgraded Version of Tokyo Stock Exchange's "arrowhead" Trading System’ https://archive.is/kkuxu https://archive.is/kkuxu
- bob1029 6y agoI've seen the guts of a few major financial organizations, and there are some common themes regarding their infrastructures. The one that really sticks out to me as an engineer is the fact that the whole system in most cases seems to be tied together by a fragile arrangement of 100+ different vendors' systems & middleware that were each tailored to fit specific audit items that cropped up over the years. Individually, all of these components have highly-available assurances up and down their contracts, but combine all these durable components together haphazardly and you get emergent properties that no one person or vendor can account for comprehensively. When the article says a full reset entails killing the power and restarting, this is my actual experience. These complex leviathans have to be brought up in a special snowflake sequence or your infra state machine gets fucked up and you have to start all over. When dependency chains are 10+ systems long and part of a complex web of other dependency chains it starts to get hopeless pretty quickly.
- outworlder 6y ago> that were each tailored to fit specific audit items that cropped up over the years. Or worse, "compliance" line items, that some tool or some company identified in their cookie-cutter processes. As long as that line item goes away, noone really cares what the long term implications are.
- Aeolun 6y agoThat’s closer to the truth in my experience.
- x87678r 6y agoExternal Audit are the most important people in financial co's Technology. Everything else is a nice to have.
- lmilcin 6y agoExactly. If you are manager, a disaster is something that is only potential. Audit and its consequences are unavoidable.
- fovc 6y agoDepending on your POV, financial exchanges are a great/awful example of the "behind the surface" complexity of modern life. You'd think with only a few order types and not that many tickers you could stand up an exchange using rust in a few nights and weekends, no? fsync liberally, pay colin his tarsnap dues, and off you go! /s
- smabie 6y agoYeah you probably could. The problem is that major exchanges like nyse or cme have hundreds of order types and absolutely staggering volume. Not to mention the billions of dollars on the line as well.
- paxys 6y agoTrading systems are generally regarded to be some of the most challenging and simultaneously unrewarding type of engineering out there. I don't think there is anyone, even on HN, who will look at the NYSE platform and go "I could build that in a weekend".
- notacoward 6y agoSounds very much like a fault plus a bad HA implementation took down the market. Couldn't care less about the fault. I'd really like to hear about the bug(s) in the HA software.
- ve55 6y agoI wonder if a solution such as "Cancel all orders and publicly announce this as well as trading commencement time such as +1 hours ahead, then reboot the server just before that and resume the day as normal" was considered, and if so why it wouldn't have worked, given it seems to satisfy the constraints mentioned in the article
- ycombobreaker 6y agoIf a system supports "Good 'Til Canceled" orders, casual application of "cancel all orders" will wipe out orders that are months old. Maybe that is called for in an extreme case (data center destroyed, offsite backups cannot be recovered), but at a minimum it is extremely rude. GTC orders even live through corporate actions, getting repriced as appropriate. They treat that order lifetime seriously.
- ncmncm 6y agoIt is not actually a very bad thing for the whole market to go down for everybody at the same time. What is bad is for part of the market to go down, or for it to go down for some people but not others. This failure highlights how frequent, scheduled testing of your failover system is needed in order to be able to say, honestly, that you even have a failover system, and not just another box burning power and doing nothing. When you have a choice, it is often better to have both systems running all the time, at less than half capacity and sharing the load; or each doing all the work, and throwing away half. If you choose the former, traffic can sometimes peak over half capacity, without loss. If the latter, you can check that they are producing the same answers, too.
- Rochus 6y agoApart from the enormous loss of sales and the lost profits for all involved.
- ncmncm 6y agoEvery person's profit would have been somebody else's loss. They match exactly, less transaction costs which were saved by both sides. Of course somebody else didn't collect on the transaction costs.
- neximo64 6y agoExcept if new shares are issued at no cost. Or new currency is minted/credit is created which happens literally every day.
- freddie_mercury 6y agoWhat new currency is minted or traded on the Tokyo Stock Exchange?
- neximo64 6y agoJapanese yen of which the shares traded on the TSE are denominated
- fomine3 6y agoI really impressed how the CIO replies explains the problem on press conference. He knows the system and explains the issue perfectly.
- xwolfi 6y agoI'm trying to find these details and I cant watch videos: would you be able to summarize ?
- xvilka 6y agoShouldn't something like Lasp[1] help to build better distributed and fault-tolerant systems of this scale? Or using the purer programming languages and frameworks, like Jane Street with OCaml, Standard Chartered with Haskell. [1] https://lasp-lang.readme.io/docs https://lasp-lang.readme.io/docs
- neurotech1 6y agoLeaders frankly accepting responsibility, known as the Asoh Defense[0] is named after Capt. Kohei Asoh after accidentally landing a DC-8, JAL Flight 2 in San Francisco bay. [0] https://en.wikipedia.org/wiki/Japan_Airlines_Flight_2#The_%22Asoh_defense%22 https://en.wikipedia.org/wiki/Japan_Airlines_Flight_2#The_%2...
- ezoe 6y agoThe interesting aftermath from this situation is the Top of Tokyo Stock Exchange(東証) as a company is recognizing the technical situation well. That's not the case for most of the companies in Japan as most of the companies simply outsource their system, but 東証 was not. While Japanese so-called journalists are completely blind at technology, even worse, they don't even have a basic literacy or listening skill, questioning things they already told, or fart out a question like "But computers won't break, isn't it?" while the cause is likely be soft memory error.