10 ms·
My curiosity is killing me as to what exactly went wrong. And as someone living in The Netherlands I'm also kind of mad at the apparent fragility of this huge c
by miggol 5y ago
My curiosity is killing me as to what exactly went wrong. And as someone living in The Netherlands I'm also kind of mad at the apparent fragility of this huge chunk of our infrastructure. If this had happened on a work day, it would have been a real national emergency, rather than just a huge national inconvenience.
I would appreciate one of those post-outage "what happened" reports like with the AWS and Facebook outages last year. But outside of IT I don't think anyone really expects those over here. And there might be national security considerations preventing such disclosure until any chance of a repeat has been engineered away, at which point everyone will have forgotten.
- tjr225 5y agoPhysical infrastructure can be fragile/fail catastrophically, as well. I recall starting a job near Seattle, 40 miles from my home in Olympia, WA and a passenger train derailed onto the highway killing a few people and backing up traffic for days. That was not a pleasant first day! https://en.wikipedia.org/wiki/2017_Washington_train_derailment https://en.wikipedia.org/wiki/2017_Washington_train_derailme...
- Wowfunhappy 5y agoI assume, however, that this didn't ground every single train in the state!
- eepp 5y agoDamn this made me read a lot about this specific derailment, but also about Puget Sound, Salish Sea, Positive train control and all this 220-MHz stuff. Thanks!
- gonzo41 5y agoI bet their schedule once ran on a commodore 64, and then they upgraded to a 386, and then ignored the IT department for 30 years. Hence today, chickens are coming home to roost.
- postingposts 5y agoMy very good guess is that you had a database failure caused by a crypto virus. They will not wish to announce that this was the root cause until A. A plan of action to avoid a future infection can be enumerated. B. They can be certain they will restore the functionality of the trains. The other part of my guess is that they have a SQL db which is stored on /likely/ a windows server instance. I’d even surmise that this instance may be hosted on Azure but that’s speculation, not a good guess.
- lucb1e 5y agoThis is one of my leading theories as well, although I thought they usually hit on Friday night in order to catch admins off-guard and encrypt as much as possible before being stopped? "The IT failure occurred at the end of [Sunday] morning". On Sunday there is also not as much pressure to get things running as there would be on Monday. Or perhaps this is the incentive to pay now and get things fixed on time for the work week? In that case Friday night would again have made more sense unless the attackers have some very specific insight into how much slower restoring without paying is. My alternative theory is an expired certificate that makes some core systems just not talk to each other anymore. The announcements on the stations, for example, were also out, and they lost control of the mobile apps (couldn't make the apps show that trains didn't run, the in-app scheduler showed all was A-OK), and that sounds quite dissimilar from the core train operation service, making me think it's more of an infrastructure than a specific system's problem. On the other hand, once the app could be controlled again (assuming this singular underlying cause), you'd think they could then also start putting train service back in place and that didn't happen for hours still. I can't really make the pieces fit together for any theory, so then presumably something multi-faceted (one thing tripping one or two other things so it escalated from restarting one component to not being able to restart the trains anymore the whole day).
- withinboredom 5y agoCertificates make the most sense to me. The delay can possibly be due to cached certificates, having to find the person to sign the CSR (assuming in-house PKI or even just someone with access to the right email address for external PKI), and/or a CA that isn’t valid on the machines.
- quartz 5y agoFrom the notice: > It affected the system that generates up-to-date schedules for trains and staff. ...boy oh boy did trying to look into what NS uses for crew scheduling ever send me down a rabbit hole. I don't know if the systems have changed but while poking around online I found this[1] doc about the Netherlands' timetable revamp around 2006 and it talks about the complexity of TURNI-- their on-the-fly crew scheduling system. > A typical workday at NS includes approximately 15,000 trips for drivers and 18,000 for conductors. The resulting number of duties is approximately 1,000 for drivers and 1,300 for conductors. This leads to extremely difficult crew scheduling instances. Nevertheless, because of the highly sophisticated applied algorithms, TURNI solves these cases in 24hours of computing time on a personal computer. Therefore, we can construct all crew schedules for all days of the week within just a few days. Then I found more detail about TURNI's implementation in this[2] paper about optimizing crew scheduling for timetables. > In the railway industry the sizes of the crew scheduling instances are, in general, a magnitude larger than in the airline industry. Moreover, crew can be relieved during the drive of a train resulting in much more trips per duty than typical in airlines. In other words, the combinatorial explosion is much higher. The latter has made the application of these models in the railway industry prohibitive until recently. Cool stuff. Finally, gleaning from ns.nl's careers page[3] everything else in their IT land outside this system runs off SAP (likely including the actual distribution of the output of crew scheduling) so if I had to gamble I'd say the failure happened somewhere in the integration between them. [1] https://homepages.cwi.nl/~lex/files/Interfaces.pdf https://homepages.cwi.nl/~lex/files/Interfaces.pdf [2] https://repub.eur.nl/pub/11701/ei200803.pdf https://repub.eur.nl/pub/11701/ei200803.pdf [3] https://werkenbijns.nl/werkgebieden/it/sap-specialist-bij-ns/ https://werkenbijns.nl/werkgebieden/it/sap-specialist-bij-ns... sidenote: If anyone out there is an SAP specialist ns.nl looks like a pretty great place to work: 36hr week, 5 weeks vacation, pension, and free unlimited 2nd class + low cost 1st class train travel.
- deleted 5y ago[deleted]
- incompatible 5y agoIf they are scheduling days in advance, then why would a failure of the scheduling system cause an instant outage?
- totetsu 5y agoIf it happed on a work day I wonder if peoples pandemic logic would kick in and you would start hearing, "It's not worth closing down the whole economy just because of the risk of a few people dying in a train crash"
- jokoon 5y ago> And there might be national security considerations That especially. I wonder how expensive that thing is going to cost.