28 ms·
UK air traffic control meltdown
- cratermoon 3y ago"the description sounds like the procedure is working directly on the textual representation of the flight plan, rather than a data structure parsed from the text file. This would be quite worrying, but it might also just be how it is explained." Oh, this is typical in airline industry work. Ask programmers about a domain model or parsing, they give you blank stares. They love their validation code, and they love just giving up if something doesn't validate. It's all dumb data pipelines At no point is there code models the activities happening in the real world. In no system is there a "flight plan" type that has any behavior associated with it or anything like a set of waypoint types. Any type found would be a struct of strings in C terms, passed around and parsed not once, but every time the struct member is accessed. As the article notes, "The programming style seems very imperative.".
- jameshh 3y agoThat's super interesting (and a little terrifying). It's funny how different industries have developped different "cultures" for seemingly random reasons.
- cratermoon 3y agoIt was terrifying enough for me in the gig I worked on that dealt with reservations and check-in, where a catastrophic failure would be someone boarding a flight when they shouldn't have. To avoid that sort of failure, the system mostly just gave up and issued the passenger what's called an "Airport Service Document": effectively a record that shows the passenger as having a seat on the flight, but unable to check-in. This allows the passenger to go to the airport and talk to an agent at the check-in desk. At that point, yes, a person gets involved, and a good agent can usually work out the problem and get the passenger on their flight, but of course that takes time. If you've ever been a the airline desk waiting to check-in and an agent spends 10 minutes working with a passenger (passengers), it's because they got an ASD and the agent has to screw around directly in the the user-hostile SABRE interface to fix the reservation.
- 3pac 3y agoSABRE is pretty good compared to the card file it replaced.
- cratermoon 3y agoIt's better to say SABRE replicated, in digital form, that card file. And even today the legacy of that card form defines SABRE and all the wrappers and gateways to it.
- touisteur 3y agoGiving up if something doesn't validate is indeed standard to avoid propagating badly interpreted data, causing far more complex bugs down the line. Validate soon, validate strongly, report errors and don't try to interpret whatever the hell is wrong with the input, don't try to be 'clever', because there lie the safety holes. Crashing on bad input is wrong, but trying to interpret data that doesn't validate, without specs (of course) is fraught with incomprehension and incompatibilities down the line, or unexpected corner cases (or untested, but no one wants to pay for a fully tested all-goes system, or just for the tools to simulate 'wrong inputs' or for formal validation of the parser and all the code using the parser's results). There are already too many problems with non-compliant or legacy (or just buggy) data emitters, with the complexity in semantics or timing of the interfaces, to try and be clever with badly formatted/encoded data. It's already difficult (and costly) to make a system work as specified, so subtle variations to make it more tolerant to unspecificied behaviour is just asking for bugs (or for more expensive systems that don't clear the purchasing price bar).
- cratermoon 3y agoThere's a difference between parsing and validating. https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-validate/ https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va... You're right about all the buggy stuff out there, and that nobody wants to pay to make it better, though.
- touisteur 3y agoFrom a safety-critical standpoint, I've always found this article interesting but strange. You want both, before taking into account any data from anything outside of the system. Do both. As soon as possible. Don't propagate data you haven't validated in any way your spec says so. If you have more stringent specs than any standard you're using, be explicit about it, reject the data with a clear failure report. Check for anything that could be corrupted, misformated, something that you're not expecting and could cause unexpected behaviour. I feel the lack of investment in destroying the parsing- (and validation-) related classes of bugs is the worst oversight in the history of computing. We have the tools to build crash-proof parsers (spark, Frama-C, and custom model checked code generators such as recordflux) that - not being perfect in any way - if they had a tiny bit of the effort the security industry put in mending all the 'Postel's law' junk out there, we'd be working on other stuff. I built, with an intern, an in-house bit-precise code generator for deserializers that can be proved absent of runtime errors, and am moving to semantics checks ('field X and field Y can only present together', or 'field Y must be greater or equal to the previous time field Y was present'). It's not that hard, compared to many other proof and safety/security endeavours.
- asimpleusecase 3y agoAnd why could the system not put the failed flight plan in a queue for human review and just keep on working for the rest of the flights? I think the lack of that “feature” is what I find so boggling.
- cratermoon 3y ago> why could the system not put the failed flight plan in a queue Because it doesn't look at the data as a "flight plan" consisting of "way points" with "segments" along a "route" that has any internal self-consistency. It's a bag of strings and numbers that's parsed and the result passed along, if parsing is successful. If not, give up. In this case fail the entire systemand take it out of production. Airline industry code is a pile of badly-written legacy wrappers on top of legacy wrappers. (Mostly not including actual flight software on the aircraft. Mostly). The FPRSA-R system mentioned here is not a flight plan system, it's an ETL system. It's not coded to model or work with flight plans, it's just parsing data from system A, re-encoding it for system B, and failing hard if it it can't.
- slt2021 3y agogood ETLs are usually designed to separate good records from bad records, so even if one or two rows in the stream do not conform to schema - you can put them aside and process the rest. seems like poor engineering
- jandrese 3y agoThe problem is that it means you have a plane entering the airspace at some point in the near future and the system doesn't know it is going to be there. The whole point of this is to make sure no two planes are attempting to occupy the same space at the same time. If you don't know where one of the planes will be you can't plan all of the rest to avoid it. The thing that blows my mind is that this was apparently the first time this situation had happened after 15 million records processed. I would have expected it to trigger much more often. It makes me wonder if there wasn't someone who was fixing these as they came up in the 4 hour window, and he just happened to be off that day.
- darkclouds 3y agoInteresting to see that flight plans over the UK have to be filed 4 hours in advance. No mention of plane, pilot, passenger and cargo manifests. So why the 4 hour lead time, is this the time it takes UK Authorities to look people up or workout if the cargo could be dangerous in an airborne Anthrax (Gruinard) Island [1] or Japanese subway Sarin [2], or an IRA favourite, fertilizer bomb thats bypassed the usual purchase reporting regulations used by people like Jeremy Clarkson and Harry Metcalfe as their store of wealth[3]? It makes me wonder just how much more surveillance of the population exists, knowing I cant even step out of the front door without attracting surveillance of the type that followed Dr David Kelly. Sure its not a cyber attack per se, carried out over the internet like a DDOS attack or a brute force password guessing attack with port knocking mitigation, but how would one carry out a cyber attack on this system if the only attack vector is from people submitting flight plans? There sure is a constant playing down of the cyber attack angle to this which makes me think someone wants to Blurred Lines! One point on the lack of uniquely named global way points, which is the main crux of the problem falling over if some are to be believed. The USA demonstrates a disproportionate number of similar names, by virtue of Europeans migrating to the US [4]. So has this situation arisen with this system in other parts of the world like in the US? How can a country that created the globe spanning British Empire become so insular with regards to air travel in this way? I'd agree with the initial assessment that there appears to be a lack of testing, but are the specifications simply not fit for purpose? I'm sure various pilots could speak out here, because some of the regulations require planes to be minimally distanced from each other when transiting across the UK. On the point of ICAO and other bodies to eradicate non-unique waypoint names, its clear there is some legacy constraint still impeding the safety of air travellers, perhaps caused by poor audio quality analogue radio, so perhaps its time for the unambiguous and globally recognised What 3 Words form of location identifier, to come into effect? The UK police already prefer it to speed up response times [4]. And although the same location can create 3 different words, suggesting drift with GPS [5], even if What 3 Words could not be used for a global system, having something a bit longer to create an easily recognisable human globally unique identifier is needed for these flight plans and perhaps maritime situations. Obviously global coordination will be like herding cats, and if such a fixed size global network of cells were introduced, some area's like transiting over the Atlantic or Pacific could command bigger cells, but transiting over built up areas like London would require smaller sized identifiable cells. But IF ever there was a time for the New World Order to step up to the plate and assert itself, to create a Globally Unique Place ID (GUPID) for the whole planet, now is the time. On the point of humans were kept safe, only by the sheer common sense of the pilots and traffic control tower staff, its not something NATS did or should claim, their systems were down, so everyone had to resort back to pen and paper and blocks in queues, and apart from Silverstone when the F1 British Grand Prix is on, is air space ever that densely populated. NATS were caught with their pants down at so many levels of altitude, is this laissez faire UK management style that saw the Govt having to step in to bail out the banks during the financial crisis, still infecting other parts of UK life and still coming to light? It's beginning to look a lot like Christmas! [1] https://www.youtube.com/watch?v=_8Zr0IPtx80 https://www.youtube.com/watch?v=_8Zr0IPtx80 [2] https://www.youtube.com/watch?v=RTr1lquCQMg https://www.youtube.com/watch?v=RTr1lquCQMg [3] https://youtu.be/LS54AJSadT4?t=279 https://youtu.be/LS54AJSadT4?t=279 [4] https://en.wikipedia.org/wiki/List_of_U.S._places_named_after_non-U.S._places https://en.wikipedia.org/wiki/List_of_U.S._places_named_afte... [5] https://www.bloomberg.com/news/articles/2019-03-21/u-k-police-are-using-three-words-to-speed-up-response-times?leadSource=uverify%20wall https://www.bloomberg.com/news/articles/2019-03-21/u-k-polic... [6] https://support.what3words.com/en/articles/2212837-why-do-i-see-different-words-every-time-i-press-the-location-button https://support.what3words.com/en/articles/2212837-why-do-i-...
- NovemberWhiskey 3y agoThis is not the first time this has happened; the phenomenon has even got a name - "poison flight plan".
- krisoft 3y ago> the phenomenon has even got a name - "poison flight plan". Maybe, but it must not be a common phrase because your comment is the first result when I search for it. And it is also mentioned in this article: http://www.aero-news.net/subsite.cfm?do=main.textpost&id=ce27c0d8-dbb5-46fa-be6e-e47af692dba1 http://www.aero-news.net/subsite.cfm?do=main.textpost&id=ce2... And that's about it? Do you have any other sources?
- tpmx 3y agoI think that term was invented four days ago by that article writer. There are four other occurrances before then and they're about PS2 games.
- NovemberWhiskey 3y agoThis term was in wide circulation when I was consulting at NATS in the 2000-2005 time frame.
- krisoft 3y agoThat is indeed very currious. So you say NATS was amaware of this vulnerability?
- quickthrower2 3y agoIronically the term about name clashes has a name clash!
- dboreham 3y agoThe generic term I'm familiar with is "ping of death".
- omginternets 3y agoOf course they blamed the French ^^
- gumballindie 3y ago[flagged]
- sergers 3y agoWell that's DailyMail for you, where they tag anything parenting or healthy as "femail" section... cause you know only women are looking at that stuff. Lol. Anyways I actually think that's just reasonable response, system goes down/related system goes down , and in reviewing they are making frivolous updates to names that aren't needed. I would question these updates (while they may be minor part of overall updates occuring).
- cratermoon 3y agoAt least until the 70s most newspapers had a section called "Women" or something similar. Even the news about the 60s/70s women's movement appeared there, not in the main "news" sections. Those sections were mostly renamed around that time to "Lifestyle", "Home", or just "Features".
- vixen99 3y agoIs this the UK or US edition? It's always easy fun to have a go at the Daily Mail which presumably you read regularly else you wouldn't be commenting. Its sin seems to be that it's not a serious broadsheet. It's a tabloid with very broad appeal that has to be profitable and therefore tries to reflect the requirements of the British public for such a publication. Perhaps you should lower your expectations. 'Tag anything parenting or healthy ...'? No, that's not correct. Here are a few health & food related items back to mid-September that did not appear in 'female'. You are right about parenting; most parenting in the UK is still undertaken primarily (in terms of executive action) by females so items on this topic are reasonably included in 'female'. The growing number of people who don't have children probably appreciate this sub-grouping by the Mail. You may not approve but this is what happens. Single males with dependent children are not known for objecting to checking out that section. It's not forbidden. https://www.dailymail.co.uk/wires/pa/article-12505173/Healthy-lifestyle-key-preventing-depression--regardless-genetic-risk.html https://www.dailymail.co.uk/wires/pa/article-12505173/Health... https://www.dailymail.co.uk/wires/ap/article-12504751/Eggplant-stuffed-pita-sandwiches-power-quick-pickle.html https://www.dailymail.co.uk/wires/ap/article-12504751/Eggpla... https://www.dailymail.co.uk/health/article-12504649/Suicide-mans-family-Ozempic-warning-label.html https://www.dailymail.co.uk/health/article-12504649/Suicide-... https://www.dailymail.co.uk/health/article-12504813/Anthony-Fauci-no-mask-mandates-covid-deadly-wave.html https://www.dailymail.co.uk/health/article-12504813/Anthony-... https://www.dailymail.co.uk/health/article-12503801/Cancer-number-preventable-cases-UK.html https://www.dailymail.co.uk/health/article-12503801/Cancer-n... https://www.dailymail.co.uk/wires/reuters/article-12503815/Waitrose-Aldi-cut-prices-Britains-food-inflation-picture-improves.html https://www.dailymail.co.uk/wires/reuters/article-12503815/W... https://www.dailymail.co.uk/wires/reuters/article-12503299/Ripe-change-Activist-investors-eye-food-consumer-goods.html https://www.dailymail.co.uk/wires/reuters/article-12503299/R... https://www.dailymail.co.uk/news/article-12468365/One-woman-reveals-community-pantry-demand-Woolworths-donates-13-million-meals-help-fight-hunger.html https://www.dailymail.co.uk/news/article-12468365/One-woman-... https://www.dailymail.co.uk/wires/reuters/article-12502685/Waitrose-cuts-prices-Britains-food-inflation-picture-improves.html https://www.dailymail.co.uk/wires/reuters/article-12502685/W... https://www.dailymail.co.uk/wires/ap/article-12501533/Food-recalls-pretty-common-things-like-rocks-insects-plastic.html https://www.dailymail.co.uk/wires/ap/article-12501533/Food-r... https://www.dailymail.co.uk/news/article-12490747/How-safe-childrens-school-dinners.html https://www.dailymail.co.uk/news/article-12490747/How-safe-c...
- tantalor 3y ago> The programming style is very imperative Is that supposed to be a meaningful statement?
- deleted 3y ago[deleted]
- amiga386 3y ago[flagged]
- tome 3y agoYes, typically it would be used to mean things like the code mutates data in place rather than using persistent data structures, explicitly loops over data rather than using higher-order map, fold etc. operations, and explicitly checks tag bits rather than using sum types.
- tantalor 3y agoFine, I'll give you that (sounds like a generic description) but there's nothing like that from the description given in the article and the paragraph immediately before that statement. It's almost as if the author completely made that up.
- rcostin2k2 3y agoThe fact that they blamed the French flight plan already accepted by Eurocontrol proves that they didn't really know how the software works. And here the Austrian company should take part of the blame for the lack of intensive testing.
- littlestymaar 3y agoThey blamed the French because they are British, that's it. It's hard to get rid of bad habits.
- Diggsey 3y agoSoftware has bugs, that's not really the damning part... The damning part is that in four hours and two levels of support teams, there was noone who actually knew anything about how the system worked who could remove the problematic flight plan so that the rest of the system could continue operating! What exactly is the point of these support teams when they can't fix the most basic failure mode (a single bad input...)
- deleted 3y ago[deleted]
- jahewson 3y ago> What exactly is the point of these support teams when they can't fix the most basic failure mode (a single bad input...) To collect money on support contracts, I suspect.
- Maxion 3y agoTry to get developers who love to code and create to stay on a support team and be on an on-call roster. I betcha at least half will say no, and the other half will either leave or you'll run out of money paying them.
- gonzo41 3y agoAnd when did you last test your monthly backups? But seriously. If you fill out all the positions in an org chart it's easy to think you're delivering, and for a lot of situations it usually works. Anointing someone a manager usually works out because people can muddle through. It doesn't work in medicine, or as it turns out, air traffic control. Lesson learned for about the next ~5 years.
- blibble 3y agoI wouldn't expect level 1 and level 2 to be able to diagnose a problem like this level 3 (devs) should have been brought in much quicker though
- 3y ago
- thrdbndndn 3y ago> Flight Plan Reception Suite Automated (FPRSA-R) Where does the "-R" come from?
- reactordev 3y agoSo they forgot to "geographically disparate" fence their queries. Having built a flight navigation system before, I know this bug. I've seen this bug. I've followed the spec to include a geofence to avoid this bug.
- deleted 3y ago[deleted]
- sam0x17 3y agoWhy on earth do they not have GUIDs for these navigation points if the names are not globally unique and inter-region routes are commonplace?
- amoerie 3y agoLong story: because changing identifiers is a considerable refactoring, and it takes coordination with multiple worldwide distributed partners to transition safely from the old to the new system, all to avoid a hypothetical issue some software engineer came up with Short story: money. It costs money to do things well.
- ftxbro 3y ago> Long story: because changing identifiers is a considerable refactoring is this what refactoring means
- NBJack 3y agoYes. It would cascade into: Changes in how ATCs operate Changes in how pilots operate Changes in how airplanes receive these instructions (including the flight software itself, safety systems, etc.) Changes in how airplanes are tested Changes in how pilots are trained Etc. In this case, the refactoring requires changes to hardware, software, training, manufacturing, and humans.
- 3y ago
- rglover 3y agoThis is one of the many reasons there should be a universal data standard using a format like JSON. Heavily structured, easy to parse, easy to debug. What you lose in footprint (i.e., more disk space), you gain in system stability. Imagine a world where everybody uses JSON and if they offer an API, you can just consume the data without a bunch of hoop jumping. Failures like this would vanish overnight.
- 0xffff2 3y agoBroadly speaking I think this is done for new systems. What you need to identify here is how and when you transition legacy systems to this new better standard of practice.
- rglover 3y agoI'd argue in favor of at least an annual review process. Have a dedicated "feature freeze, emergencies only" period where you evaluate your existing data structures and queue up any necessary work. The only real hang up here is one of bad management. In terms of how, it's really just a question of Schema A to Schema B mapping. Have a small team responsible for collection/organization of all the possible schemas and then another small team responsible for writing the mapping functions to transition existing data. It would require will/force. Ideally, too, jobs of those responsible would be dependent on completion of the task so you couldn't just kick the can. You either do it and do it correctly or you're shopping your resume around.
- tristor 3y agoThe problem is systems written in the 1970s in FORTRAN to run on Mainframes don't speak JSON.
- rglover 3y agoGreat. It should be fixed by replacing the FORTRAN systems with a modern solution. It's not that it can't be done, it's that the engineers don't bother to start the process (which is a side-effect of bad incentive structure at the employment level).
- sp0ck 3y agoSmall suggestion. Don't choose obscure language (in terms of popularity, 28th on TIOBE index with 0.65% rating) to visualize structure and algorithms. Otherwise you risk average viewer will stop reading the moment he encounter code samples. There are 27 more popular languages, some of them orders of magnitute more.
- louthy 3y agoMaybe he doesn’t care if people stop reading and he’d prefer to use the language he’s most comfortable with? It’s his blog after all, not yours. Additionally, perhaps he’s making the point that a language with an expressive type system makes solving problems like this trivial.
- sp0ck 3y agoIf you don't care about readers reading it or not then what is the point to publish an article ?
- rjh29 3y agoThe code is a relatively small part of the article, and quite far into it I might add.
- daaaaaaan 3y agoI appreciated the Haskell examples, they aren't particularly hard to follow. How do you think those more popular languages got more popular?
- redleader55 3y agoI imagine, for this kind of system, there is only one supplier. Why not force that supplier, as part of their 10-15 yr contract, to publish the source code for everything, not necessarily as FOSS. This way if there are bugs they can be reported and fixed.
- passwordoops 3y agoI agree. But this would assume that: 1- the people writing and approving the specs even understand why this might be a good suggestion 2- the people ultimately approving the contract aren't in bed with the supplier
- FeepingCreature 3y ago3- the people operating the system are capable of maintaining its source code
- dboreham 3y agoThere's always prison for those people.
- dang 3y agoRelated. Others? Coincidentally-identical waypoint names foxed UK air traffic control system - https://news.ycombinator.com/item?id=37430384 https://news.ycombinator.com/item?id=37430384 - Sept 2023 (64 comments) UK air traffic control outage caused by bad data in flight plan - https://news.ycombinator.com/item?id=37402766 https://news.ycombinator.com/item?id=37402766 - Sept 2023 (20 comments) NATS report into air traffic control incident details root cause and solution - https://news.ycombinator.com/item?id=37401864 https://news.ycombinator.com/item?id=37401864 - Sept 2023 (19 comments) UK Air traffic control network crash - https://news.ycombinator.com/item?id=37292406 https://news.ycombinator.com/item?id=37292406 - Aug 2023 (23 comments)
- a_wild_dandan 3y agoThe recent episode of The Daily about the (US) aviation industry has convinced me that we’ll see a catastrophic headline soon. Things can’t go on like this.
- switch007 3y agoThe title of this post made me think there was a new, current meltdown !
- c7DJTLrn 3y agoThis is an interesting engineering problem and I'm not sure what the best approach is. Fail safe and stop the world, or keep running and risk danger? I imagine critical systems like trading/aerospace have this worked out to some degree.
- fbdab103 3y agoAbsolutely no idea on what is correct, but I love to reference this article on software practices at NASA[0], They Write the Right Stuff. [0] https://www.fastcompany.com/28121/they-write-right-stuff https://www.fastcompany.com/28121/they-write-right-stuff
- crabbone 3y agoThere isn't and cannot be a preference to either one. It always depends on what the system is doing and what the consequences would be... Pacemaker cannot "fail safe" for example, under no circumstances. It's meaningless to consider such cases. But if escalation to a human operator is possible, then it will also depend on how the system is meant to be used. In some cases it's absolutely necessary that the system doesn't try to handle errors (eg. if say a patient is in a CT machine -- you always want to stop to, at least, prevent more radiation), but in the situation like the one with the flight control -- my guess is that you want the system to keep trying while alerting the human operator. But then it can also depend on what's in the contract and who will get the blame for the system functioning incorrectly. My guess here is that failing w/o attempting to recover was, while an overkill, a safer strategy than to let eg. two airplanes be scheduled for the same path (and potentially collide).
- cellularmitosis 3y agoThe best approach is to simply print the error to the screen, rather than burying it in a “low level log” which only the software vendor has access to. They had a four hour buffer until the world stopped, but most of that was pissed away because no one knew what the problem was.
- throw7 3y agoWell, I certainly hope they've at least stopped issuing waypoints with identical names... although it wouldn't surprise me if geographically-distant is the best we can do as a species.
- SoftTalker 3y agoThey appear to be sequences of 5 upper-case letters. Assuming the 26-character alphabet, that should allow for nearly 12 million unique waypoint IDs. The world is a big place but that seems like it should be enough. The more likely problem is that there is (or was) no internationally-recognized authority in charge of handing out waypoint IDs, so we have at least legacy duplicates if not potential new ones.
- seabass-labrax 3y agoYou have to reduce that to the (still massive) set of IDs that are somewhat pronounceable in languages that use the Latin script. You don't want to be the air traffic controller trying to work out how to say 'Lufthansa 451, fly direct QXKCD'. Nonetheless, I think the there is little cause for concern about changing existing IDs. There might be sentimental attachment, but it takes barely a few flights before the new IDs start sticking, and it's not like pilots never fly new routes.
- SoftTalker 3y agoI thought that is what the "ICAO pronunciation" was for? "Fly direct Quebec Xray Kilo Charlie Delta"
- drachir91 3y agoNo, waypoints aren't spelled out with the ICAO alphabet. They are mnemonics that are pronounced as a word and only spelled out if the person on the receiving end requests it because of bad radio reception, or unfamiliarity with the area/waypoint. For example, Hungarian waypoints, at least the more important ones are normally named after cities, towns or other geographical locations near them, and use the locations name or abbreviated name, being careful that they can be pronounced reasonably easily for English speakers. Like: ERGOM (for the city Esztergom), ABONY (for the town Füzesabony), SOPRO (for Sopron), etc.
- Rochus 3y agoThis is apparently just an opinion, no additional inside information than we had from the report (https://news.ycombinator.com/item?id=37401981 https://news.ycombinator.com/item?id=37401981), isn't it? EDIT: downvoting this question instead of responding is a pretty strange reaction.
- seabass-labrax 3y agoYou are correct, but it's an opinion that bridges the gap editorially between those knowledgable about ATC but not data, and those knowledgable about data but not ATC. This is a valuable service to provide, as both fields are rather complex.
- Rochus 3y agoThanks. I didn't have the patience to read it all. I initially hoped that the author was a field expert or even someone with inside knowledge, but he is apparently from a completely different domain and not in the UK, and there were assumptions about things the report was rather specific about (as specific as such reports usually are). It would be more useful if people would take a closer look at the report and draw the right conclusions about organizational failures and how to avoid them. All the great software technologies to achieve memory safety, etc. are of little use if the analyses and specifications are flawed or the assumptions of the various parties in a system of systems do not match. But people seem to prefer to speculate and argue about secondary issues.
- lagt_t 3y agoDude this isn't reddit dont worry about the votes.
- Rochus 3y agoSince it's not Reddit but HN, it's all the stranger to dismiss a perfectly legitimate question. But times and mores seem to change much faster than I realize.
- 3y ago
- lbriner 3y agoI seem to remember another problem at NATS which had the same effect. Primary fell over so they switched over to a secondary that fell over for the exact same reason. It seems like you should only failover if you know the problem is with the primary and not with the software itself. Failing over "just because" just reinforces the idea that they didn't have enough information exposed to really know what to do. The bit that makes me feel a bit sick though is that they didn't have a method called "ValidateFlightPlan" that throws an error if for any reason it couldn't be parsed and that error could be handled in a really simple way. What programmer would look at a processor of external input and not think, "what do we do with bad input that makes it fall over?". I did something today for a simple message prompt since I can't guarantee that in all scenarios the data I need will be present/correct. Try/catch and a simple message to the user "Data could not be processed".
- d1sxeyes 3y agoWell, if the primary is known not to be in a good state, you might as well fail over and hope that the issue was a fried disk or a cosmic bit flip or something. The real safety feature is the 4 hour lead time before manual processing becomes necessary. One of the key safety controls in aviation is “if this breaks for any reason, what do we do”, not so much “how do we stop this breaking in the first place”.
- zaphar 3y agoI'm no aviation safety controls expert but it seems to me that there are two types of controls that should be in place: 1. Process controls: What do we do when this breaks for any reason. 2. Engineering controls: What can we do to keep this from breaking in the first place? Both of them seem to be somewhat essential for a truly safe system.
- mixdup 3y agoIt's very hard to ensure you capture every single possible failure mode. Yes, the engineering control is important but it's not the most critical. What to do if it does fail (for any reason) is the truly critical control, because it solves for the possibility of not knowing every possible way something might fail and therefore missing some way to prevent a failure
- anupj 3y agoGreat writeup
- SillyUsername 3y agoBugs happen. Fact of being written by fleshy meatballs. What should also have been highlighted is that they clearly had no easy way of finding the specific buggy input in the logs nor simulating it without contacting the manufacturer.
- FeepingCreature 3y agoNo way or no procedure.
- clnq 3y agoIt sounds like a simple functional smoke test throwing random flight plans at the system would have eventually (and probably pretty soon) triggered this. I hope they at least do it now. This reminds me of: https://danluu.com/wat/ https://danluu.com/wat/
- CodeL 3y ago[flagged]
- dundarious 3y agoI wish the article contained some explanation of why the processing for NATS requires looking at both the ADEXP waypoints and the ICAO4444 waypoints (not a criticism per se, it may not have been addressed in the underlying report). Just looking at the ADEXP seems sufficient for the UK segment logic. I'm guessing it has something to do with how ICAO4444 is technically human readable, and how in some meaningful sense, pilots and ATC staff "prefer" it. e.g., maybe all ICAO4444 waypoints are "significant" to humans (like international airports), whereas ADEXP waypoints are often "insignificant" (local airports, or even locations without any runway at all). Of course with 20/20 hindsight, it seems obviously incorrect to loop through the ICAO4444 waypoints in their entirety, instead of "resuming" from an advanced position. But why look at them at all?
- masklinn 3y agoPossibly it needs the ICAO information to communicate with some systems, but has to work in ADEXP to have sufficient granularity (the essay mentions the possibility of “clipping”, a flight going through the UK between two ICAO waypoints).
- dundarious 3y agoYes, I'm essentially wanting to know more about those existing ICAO-based systems, be they machine or not.
- t0mas88 3y agoThey use the ADEXP to determine which part of the route is in the UK. Because the auto generated points are ATC area handover points. So this data is the best way so see which part of the route is within the UK airspace. Then it needs to find the ICAO part that corresponds, because the controller needs to use the ICAO plan that the pilot has. If the controller sees other (auto generated) waypoints that the pilots don't have you get problems during operation. A simple example is that controllers can tell pilots to fly in a straight line to a specific point on their filed route (and do so quite often). The pilot is expected to continue the filed route from that point onwards. They can also tell a pilot to fly direct to some random other point (this also happens but less often). The pilot is then not expected to pick up a route after that point. The radio instruction for both is exactly the same, the only difference is whether the point is part of the planned route or not. So the controller needs to see the exact same route as the pilots have, not one with additional waypoints added by the IFPS system.
- Gud 3y agoA day I don't want to remember. Took me 15 hours to reach my destination instead of 2. Had to take train, bus, then train again. 30 minutes after I had booked my tickets, everything was fully booked for two days.
- algas 3y agoI waited in the airport for 6 hours before learning that my flight was cancelled, and had to rebook... I was flying to New York to see my family, so I didn't really have any alternate transportation options!
- bbx 3y agoThat's a shame, sorry to hear that. I got more lucky: I had to wait for 6 hours too but my flight suddenly resumed (must have been one of the first few). I didn't have any alternatives to go home either so I feel for all of those stuck in a foreign country.
- conradfr 3y agoDid you meet John Candy along the way?
- supernova87a 3y agoDid the creator of the flight plan software engage in adversarial testing to see if they could break the system with badly formed flight plans? Or was / is the typical practice to mostly just see if the system meets just the "well-behaved" flight plan processing requirements? (with unit tests, etc)
- bombcar 3y agoI think we all know the answer to this. A huge portion of "exploits" in the last 20 years have been "internal business APIs" if you will being exposed to malicious actors.
- jacquesm 3y agoTrusted input rarely should be trusted. It's input. You need to validate it as if it is hostile and have a process for dealing with malformed input. Now of course, standing by the sidelines it is easy to criticize and I'm sure whoever worked on this wasn't stupid. But I've seen this error often enough now in practice that I think that it needs to be drilled into programmers heads more forcefully: stuff is only valid if you have just validated it. If you send it to someone else, if someone you trust sends it to you, if you store in a database and then retrieve it and so on then it is just input all over again and you probably should validate it for being well-formed. If you don't do that then you're a bitflip, migration or an update away from an error that will cause your system to go into an unstable state and the real problem is that you might just propagate the error downstream because you didn't identify it. Input is hard. Judging what constitutes 'input' in the first place can be harder.
- lgeorget 3y agoFrom what I gathered from the article, the input WAS valid. It's the software that was unable to handle a specific case of valid input.
- jacquesm 3y agoThat's fine, and is exactly the kind of case that I was thinking of: your software has a different idea of what is valid than an upstream piece of software, so from your perspective it is invalid. So you need to pull this message out of the stream, sideline it so it can be looked at by someone qualified enough to make the call of what's the case (because it could well be either way) and processing for all other messages should continue as normal. After all the only reason you can say with confidence that it in fact was valid is because someone looked at it! You can only do that well after the fact. A message switch [1] that I worked on had to deal with messages sources from 100's of different parties and while in principle everybody was working from the same spec (CCITT [2]) every day some malformed messages would land in the 'error' queue. Usually the problem was on the side of the sender, but sometimes (fortunately rarely) it wasn't and then the software would be improved to be able to handle that case correctly as well. Given the size of the specs and the many variations on the protocols it wasn't weird at all to see parties get confused. What's surprising is that it happens as rarely as it does. The big takeaway here should be that even if something happens very rarely it should still not result in a massive cascade, the system should handle this gracefully. [1] https://www.kvsa.nl/en/ https://www.kvsa.nl/en/ [2] https://en.wikipedia.org/wiki/Group_4_compression https://en.wikipedia.org/wiki/Group_4_compression
- failbuffer 3y ago> The manufacturer was able to offer further expertise including analysis of lower-level software logs which led to identification of the likely flight plan that had caused the software exception. This part stood out to me. I've found it super helpful to include a reference to which piece of days in working with in log messages and exceptions. It helps isolated problems so much faster.
- codeulike 3y agoThis is a great post. My reading of it: - waypoint names used around the world are not unique - as a sortof cludge, "In order to avoid confusion latest standards state that such identical designators should be geographically widely spaced." - but still you might get the same waypoint name used twice in a route to mean different places - the software was not written with that possibilty in mind - route did not compute - threw 'critical exception' and entered 'maintenance mode' - i.e. crashed - backup system took over, hit the same bug with the same bit of data, also crashed - support people have a crap time - it wasnt until they called the software supplier that they found the low level logs that revealed the cause of the problem
- noman-land 3y agoMy jaw kept dropping with each new bullet point.
- xvector 3y agoSame, is aviation technology really this primitive?
- H8crilA 3y agoIt is mostly quite primitive, but it also works amazingly well. For example ILS or VOR or ATC audio comms can all be received and read correctly using hardware built from entry level ham radio knowledge. Altimeters still require a manual input of pressure. Fuel levels can be checked with sticks. Kinda the opposite of a modern web/mobile app, complicated, massively bloated and breaks rather often :).
- rozap 3y agoshhh, nobody tell xvector that unleaded avgas finally happened in 2022 :)
- brianpan 3y agoYou might find it interesting that the SF subway runs on floppy disks. Not the fancy new 3.5" ones, either. https://sfstandard.com/2023/02/02/sfs-market-street-subway-runs-on-reagan-era-floppy-disks/ https://sfstandard.com/2023/02/02/sfs-market-street-subway-r...
- throw74848 3y ago[flagged]
- crabbone 3y agoI want to comment specifically on: > The software and system are not properly tested. Followed by suggesting to do fuzzing tests. * Automatically generating valid flight paths is somewhat hard (and you'd have to know which ones are valid because the system, apparently, is designed to also reject some paths). It's also possible that such a generator would generate valid but improbable flight paths. There's probably an astronomic number of possible flight paths, which makes exhaustive testing impossible, thus no guarantee that a "weird" path would've been found. The points through which the paths go seem to be somewhat dynamic (i.e. new airports aren't added every day, but in a life-span of such a system there will be probably a few added). More realistically some points on flight paths may be removed. Does the fuzzing have to account for possibilities of new / removed points? * This particular functionality is probably buried deep inside other code with no direct or easy way to extricate it from its surrounding, and so would be very difficult to feed into a fuzzer. Which leads to the question of how much fuzzing should be done and at what level. Add to this that some testing methodologies insist on divorcing the testing from development as not to create an incentive for testers to automatically okay the output of development (as they would be sort of okaying their own work). This is not very common in places like Web, but is common in eg. medical equipment (is actually in the guidelines). So, if the developer simply didn't understand what the specification told them to do, then it's possible that external testing wasn't capable of reaching the problematic code-path, or was severely limited in its ability to hit it. * In my experience with formats and standards like these it's often the case that the standard captures a lot of impossible or unrealistic cases, hopefully a superset of what's actually needed in practice. Flagging every way in which a program doesn't match the specification becomes useless or even counter-productive because developers become overloaded with bug reports most of which aren't really relevant. It's hard to identify the cases that are rare but plausible. The fact that the testers didn't find this defect on time is really just a function of how much time they have. And, really, the time we have to test any program can cover a tiny fraction of what's required to test a program exhaustively. So, you need to rely on heuristics and gut feeling.
- theptip 3y agoNone of this really argues against fuzz testing; even with completely bogus/malformed flight plans, it shouldn't be possible for a dead letter to take down the entire system. And, since it's translating between an upstream and downstream format (and all the validation is done when ingesting the upstream), you probably want to be sure anything that is valid upstream is also valid downstream. It's true that fuzz testing is easiest when you can do it more at the unit level (fuzz this function implementing a core algorithm, say) but doing whole-system fuzz tests is perfectly fine too.
- cjbprime 3y agoGreat post. This part goes too far, I think: > Human lives were kept safe at all times > The consequence of all this was not that any human lives were put in danger, .. When you're arguing that cancelling 2000 flights cost £100M and that no human danger was incurred, something should feel off. That might be around 600k humans who weren't able to be where they felt they needed to be. Did they have somewhere safe to sleep? Did they have all the medications they needed with them? Did they have to miss a scheduled surgery? Could we try to measure the effect on their well-being in aggregate, using a metric other than the binary state of alive or facing imminent death? You get the idea. Of course I agree with the version of the claim that says that no direct danger was caused from the point of view of the failing-safe system. But when you're designing a system, it ought to be part of your role to wonder where risk is going as you more stringently displace it from the singular system and source of risk that you maintain.
- deleted 3y ago[deleted]
- kodt 3y agoI mean it could have also saved lives by that logic. Did someone missing their flight mean they also missed a terrible pileup on the roadways after landing? We can imagine pretty much any scenario here.
- cjbprime 3y agoI agree with you that we don't know! But my thesis is that we should still do our best, when considering how much risk the systems we maintain should be willing to keep operating through.
- buildsjets 3y agoBut how many lives were saved by the reduced carbon emissions that were not produced by the cancelled flights?
- hermitcrab 3y ago>"in typical Mail Online reporting style: "Did blunder by French airline spark air traffic control issues?" The Daily Mail is a horrible, right-wing paper in the UK that blames 'foreigners' for everything. Particularly the French. Out of curiosity, is there a corresponding French paper that blames the English or the British for everything?
- dopidopHN 3y agoFrench here, as much as I wish It was the case for comical effect… I don’t think so. Our right wing press is also desperately economically liberal so anything privately run is inherently better. Maybe radio stations? Honestly, major respect to the daily mail for those snarky attacks that keep up the good spirits between our two countries. It’s maybe the food or the weather that make them aggro ? Idk, but don’t worry, we love to hate the perfide Albion. Too. Fellow French: am I wrong ? Maybe “valeur actuelle” could pull up that type of bullshit, but I think they are too busy blaming Islam to start thinking about our former colony across the channel.
- hermitcrab 3y ago>major respect to the daily mail for those snarky attacks There is really nothing to like or respect about the Daily Mail. https://www.globaljustice.org.uk/blog/2017/10/horrible-history-daily-mail/ https://www.globaljustice.org.uk/blog/2017/10/horrible-histo... >our former colony across the channel Touché! ;0)
- dghughes 3y agoIsn't that why France has a President the position was created just to blame them? lol
- dopidopHN 3y agoOur last iteration of the constitution grant them large powers and leeway But indeed, the unspoken rule is also that we hate them with a passion no matter what.
- 3y ago
- worik 3y agoI heard in the news that this was caused by a "bad flight plan". It is clear, even without any more information than that, it was a software failure (bad flight plan?) It will be interesting to see if Frequentis has to pay a price for causing this
- vachina 3y agoI think Frequentis will actually be paid more to “add feature to make the system more robust” and a bump in support schedules.
- anentropic 3y agoYeah it seems clear from the report that NATS published it wasn't a bad flight plan at all... the plan was valid to the relevant specifications But the specifications allow ambiguities (non-unique waypoint ids) and the software did not handle this particular ambiguity correctly
- cja 3y agoEvery system I've ever made has better error reporting that that one. Even those that only I use. First thing I get working in a new project is the system to tell me when something fails and to help me understand and fix the problem quickly. I then use that system throughout development such that it works very well in production. I'd love to talk to the people who made the system discussed in the article. Is one of them reading this? Can you explain how come this problem reported itself so badly?
- anentropic 3y agoYes it seems incredibly lame error reporting that they had to spend hours contacting the original vendor (to "analyse low-level software logs") just to find out which flight plan had crashed the system
- m1n1 3y agoIf you want to hear about how bad air traffic control is in the United States, you can listen/read here https://www.nytimes.com/2023/09/05/podcasts/the-daily/plane-collision.html https://www.nytimes.com/2023/09/05/podcasts/the-daily/plane-... There was a time recently when only 3 out of the 300+ air traffic control centers in the U.S. were fully staffed. All the rest were short-handed. Not sure how it stands today
- drachir91 3y agoWhat ticked me is that when the primary system threw in the towel, an EXACT SAME system took over and ran the exact same code on the exact same data as the primary. I know that with code and algorithms it's not always the case but even then you know what doing the same thing over and over expecting different results defines... Yes, it can be argued that the software should've had more graceful failure modes and this shouldn't have thrown a critical exception. It can be argued that the programmers should've seen this possibility. We can argue a lot of things about this. But the reality is that this is a mission-critical system. And for such systems, there're ways to mitigate all of these mistakes and allow the system to continue functioning. The easiest (but least safe) one would be to have the secondary system loaded with code that does the same thing but written by a different team/vendor. It reduces the chance from 100% to much-much less that if any input provokes an unforseen, system-breaking bug in the primary, the same input will provoke the same bug in the secondary. An even better solution is to have a triumvirate system, where all 3 have code written by different teams, and they always compare results. If 3 agree, great, if 2 agree, not so great but safe to assume that the bug is in the 1 not the 2 (but should throw an alert for the supervisors that the whole system is in a degraded mode where any further node failure is a showstopper), and if all disagree, grind everything to a halt because the world is ending, and let the humans handle it. It can be refined even further. And it's not something new. So why wasn't this system implemented in such a way? (Aside from cost. I don't care about anyones cost-cutting incentives in mission-critical systems. Sorry capitalism...)
- throwaway894345 3y ago> Aside from cost. I don't care about anyones cost-cutting incentives in mission-critical systems. Sorry capitalism... Capitalism is happy to have redundancy in mission critical systems all the time. Why would it care here?
- drachir91 3y agoI don't know but in recent years I'm increasingly seeing mission critical systems having only token or "apparent" rendundancies instead of real ones, and couldn't find any other rationale than cost savings and shareholder bottom lines. I'm not saying that capitalism = bad, it's mostly better than the alternatives, but just like its most direct competitor, it suffers from bad implementations across the world and unbounded human greed. A recent and very "in the face" example, also from the air travel industry would be the B737 Max and its AoA sensors. There were two, for two flight computers, but MCAS only used 1 flight computer and 1 AoA sensor, despite the already existing crosslinks between the flight computers and the sensors... Pofit maxing first with the "no need for a new type rating for the pilots", then cost-cutting first in aeronautical engineering (solving an airframe design problem with software, plus designing a flight envelope protection system that can overpower the human pilots). Then cost-cutting in software engineering and QC, rushing out software made by (probably) inexperienced in the field engineers and failing to properly test it and ensure that it had the needed redundancy.
- johnklos 3y agoSo the "engineering teams" couldn't tail /var/log/FPRSA-R.log and see the cause of the halt? I've had servers and software that I had never, ever used before stop working, and it took a lot less than four hours to figure out what went wrong. I've even dealt with situations where bad data caused a primary and secondary to both stop working, and I've had to learn how to back out that data and restart things. Sure, hindsight is easy, but when you have two different systems halt while processing the same data, the list of possible causes shrinks tremendously. The lack of competence in the "engineering teams" tells us lots about how horribly these supposedly critical systems are managed.
- slingnow 3y agoDamn, if only you had been there to instantly save the day by just running that simple command!
- johnklos 3y agoNo. That's silly. The logs would've / should've just shown that the program halted because it was confused about data. The actual commands to fix would've been quite different.
- seabass-labrax 3y agoYou're assuming that there is in fact a /var/log/FPRSA-R.log to tail - it would not at all surprise me if a system this old is still writing its logs to a 5.25 inch floppy in Prestwick or Swanwick^1. ^1: they closed the West Drayton centre about twenty years ago; I don't imagine they moved their old IBM 9020D too, if they still had it by then. My comment is nonetheless only slightly exaggerated ;)
- wolfendin 3y agoMy question is: why was the algorithm searching any section before the UK entry point. You can’t exit at a waypoint before you enter so there is no reason to search that space.
- seabass-labrax 3y agoI had been considering becoming an air traffic controller myself, and it rather tickles me to think I might have missed my once-in-a-lifetime opportunity to direct aircraft with the original pen-and-paper flight strip mechanism in the 21st century! Completely safe, excruciatingly low-capacity, and sounds like awfully good fun as a novelty (for the willing ATC, not the passengers stuck on the ground, I hasten to add).
- epolanski 3y agoQuite few non major airports are still heavily pen and paper reliant methods to some degree. An example are islands that serve few flights per week and can't justify heavy update investments. Airplanes are generally spaced by hours and you need to do your math about where the airplanes are by hand. But again there's so little planes that risks are minimal.
- seabass-labrax 3y agoIndeed, but the set of aerodromes that are large enough to have a tower controller but not large enough to have their own radar surveillance is shrinking all the time. Radar is getting cheaper and what with ADS-C and TA/RA, a big reason to have ATC even without radar is vanishing (namely that of preventing collisions close to the airport). Oceanic control is probably the closest you can get nowadays to routine ATC without radar, even though they now have automatic position reports via satellite.
- epolanski 3y agoI think an island in the middle of the Atlantic that is mostly used for refueling is exactly that kind of airport. Can't remember the name but I'm quite sure it belongs to Portugal.
- seabass-labrax 3y agoWas it part of the Azores?
- jliptzin 3y agoWhat I don’t understand in situations like this when thousands of flights are cancelled is how do they catch up? It always seems like flights are at max capacity at all times, at least when I fly. If they cancel 1,000 flights in one day, how do they absorb that extra volume and get everyone where they need to be? Surely a lot of people have their plans permanently cancelled?
- CamelCaseName 3y agoThere's always some empty capacity, whether it's non-rev tickets for flight crew and their families which are lower priority than paying customers or people who miss their flights. I had a cancelled flight recently and they booked people two weeks out because every flight from that day onward was full or nearly full. I showed up the next morning and was able to board the next flight because exactly one person had scanned in their boarding pass (was present at the airport) but did not show up for whatever reason to the airplane. Beyond that, people just make alternate plans, whether it's taking a bus or taxi home, traveling elsewhere, picking another airline, anything is possible.
- thedrbrian 3y agoYou don't. I work in logistics for a FMCG company and sometimes our main producer goes down and we run out of certain types of stock. We send as much out as we can and cancel the rest. If they really want the stock the customers can rebook an order for tomorrow because they aren't getting it today. And we just start adding extra stock to each delivery. It's the best of a bad situation. We don't have the money to have extra trucks and very perishable stock laying about and I know the airlines don't pay 300 grand a month to lease a 737 just to have it sat about doing nothing. There's very little slack.
- deleted 3y ago[deleted]
- bdamm 3y agoExactly. People's plans get pushed out into the evenings or during the less busy times, absorbed, then forgotten as collateral damage.
- bubblydoops 3y ago[dead]
- WalterBright 3y ago> the backup system applied the same logic to the flight plan with the same result Oops. In software, the backup system should use different logic. When I worked at Boeing on the 757 stab trim system, there were two avionics computers attached to the wires to activate the trim. The attachment was through a comparator, that would shut off the authority of both boxes if they didn't agree. The boxes were designed with: 1. different algorithms 2. different programming languages 3. different CPUs 4. code written by different teams with a firewall between them The idea was that bugs from one box would not cause the other to fail in the same way.
- wavemode 3y agoFirst thought that came to my mind as well when I read it. This failover system seems to be more designed to mitigate hardware failures than software bugs.
- WalterBright 3y agoI also understand that it is impractical to implement the ATC system software twice using different algorithms. The software at least checked for an illogical state and exited, which was the right thing to do. A fix I would consider is to have the inputs more thoroughly checked for correctness before passing them on to the ATC system.
- nightpool 3y agonot stronger isolation between different flight plans? it seems "obvious" to me that if one flight plan is causing a bug in the handling logic, the system should be able to recover by continuing with the next flight plan and flagging the error to operators to impact that flight only
- WalterBright 3y ago"unexpected errors" are not necessarily problems with the flight plans. They could be anything.
- julienmarie 3y agoThey should have used Erlang OTP
- garyfirestorm 3y ago> Safety critical software systems are designed to always fail safely. This means that in the event they cannot proceed in a demonstrably safe manner, they will move into a state that requires manual intervention. unrelated - this instantly caused me to think about tesla autopilot crashes that have been reported with emergency vehicles
- cellularmitosis 3y agoIt made me think of the “put a traffic cone on it” denial of service attack
- tw1984 3y agothis is what happens when you de-industrialize your nation and focus on things like finance that brings quick and cheap $.
- sargenem 3y ago[flagged]
- sargenem 3y agoThis looks like the perfect definition of “a man with two watches is always confused of the time”
- pmarreck 3y agoIt must suck to be responsible for a system that everyone depends on and millions of dollars are riding on so you are very reluctant to change it, even if you know it needs technical improvements. Formal verification or fuzzing could have helped them over that mistrust, but are not panaceas
- onetokeoverthe 3y ago[dead]
- circular_logic 3y agoWhat I still don't understand is how flight plans get approved? In my mind they would only be approved once all involved countries review and process the plan. That way we don't need this ridiculous idea of failing safe on the whole uk airspace for a single error. That day a single flight plan could have been rejected, perhaps just resubmitted and the bug quietly fixed in the background
- mihaaly 3y ago"Jesus, what a clusterfuck!" - J.K. Simmons in Burn After Reading
- fennecfoxy 3y ago"UK air traffic control: inquiry into whether French error caused failure" Of course bloody not. How is it a French airline's fault when it's a UK system? Systems like this should be foolproof with redundancies. If one entry is bad reject it and carry on, even.
- Neil44 3y agoWell if you want to get all nationalistic, the software was Austrian.
- fennecfoxy 3y agoIt's not nationalism just because I'm defending one country in favour of another. I'm not French and nor am I British. I feel neutral about both of them though I do live in the UK. It's just logic, not nationalism. :/
- ric2b 3y agoBut built to UK spec
- spuz 3y agoHas the culprit flight-plan been disclosed? I'd be interested to know how easy it is to create a realistic looking flight-plan through UK airspace that reproduces the problem. I.e. how much truth is there when NATS say this was a 1 in 15m probability?
- javier_e06 3y agoI worked once with 4G BTS (Base Transceiver Stations) where one of the issues was preventing the errors in the running board to propagate to the backup systems. There was no clean way to do it given the fact the malformed input will eventually reach the backup system producing the same error. The post talks about the system delaying the process to prevent backup up. Perhaps a solution would be going in the other direction having a staging step to prevent compromising the pipeline. Very interesting article.
- J8K357R 3y agoPoison Pill! Why on earth would the best failure mode be to cease operating? Just don’t accept the new plan being ingested and tell the person uploading that their plan was rejected. Impact one flight not thousands!
- pasc1878 3y agoOk if the system finds something that it does not understand what should it do - and how does the programmer know it will work?
- benrutter 3y agoI wondered this- I have absolutely no understanding of what's involved in flight system development, but does anyone know why it doesn't do this? By contrast, its normal for an API to return 500 if something goes wrong and keep serving other requests. It would seem insane if it crashed out and completely stopped. Any idea why the parallel isn't true for a flight system?
- jameshh 3y agoFor those of you still following this story, the flight plan that triggered the chaos has been identified! https://chaos.social/@russss/111048524540643971 https://chaos.social/@russss/111048524540643971! > Tonight we were wondering why nobody had identified the flight which caused the UK air traffic control crash so we worked it out. It was FBU (French Bee) 731 from LAX/KLAX to ORY/LFPO. > It passed two waypoints called DVL on its expanded flight plan: Devil's Lake, Wisconsin, US, and Deauville, Normandy, FR (an intermediate on airway UN859). > https://www.flightaware.com/live/flight/FBU731/history/20230828/0255Z/KLAX/LFPO https://www.flightaware.com/live/flight/FBU731/history/20230... > Credit to @marksteward and @benelsen for doing much of the legwork here.