15 ms·
Air traffic failure caused by two locations 3600nm apart sharing 3-letter code
- Optimal_Persona 2y agoWell, 3600 billionths of a meter IS kinda close...just sayin'
- dh2022 2y agoI read it the same way….
- marky1991 2y agoWhat did they mean, if not 'nanometers'?
- abracadaniel 2y agoNautical miles
- bilekas 2y agoI was thinking the same and thinking that’s a super weird edge case to happen. I’m obviously tired.
- FateOfNations 2y agoGood news: the system successfully detected an error and didn't send bad data to air traffic controllers. Bad News: the system can't recover from an error in an individual flight plan, bringing the whole system down with it (along with the backup system since it was running the same code).
- deleted 2y ago[deleted]
- wyldfire 2y ago> he system can't recover from an error in an individual flight plan, bringing the whole system down with it From the system's POV maybe this is the right way to resolve the problem. Could masking the failure by obscuring this flight's waypoint problem have resulted in a potentially conflicting flight not being tracked among other flights? If so, maybe it's truly urgent enough to bring down the system and force the humans to resolve the discrepancy. The systems outside of the scope of this one failed to preserve a uniqueness guarantee that was depended on by this system. Was that dependency correctly identified as one that was the job of System X and not System Y?
- martinald 2y agoYes I agree. The reason the system crashed from what I understand wasn't because of the duplicate code, it was because it had the plane time travelling, which suggests very serious corruption.
- kevin_thibedeau 2y agoWaves hand... This is not the SQL injection you're looking for. It's just a serious corruption.
- aftbit 2y agoIt seems fundamentally unreasonable for the flight processing system to entirely shut itself down just because it detected that one flight plan had corrupt data. Some degree of robustness should be expected from this system IMO.
- HeyLaughingBoy 2y agoIt depends on what the potential outcomes are. I've worked on a (medical, not aviation) system where we tried as much as possible to recover from subsystem failures or at least gracefully reduce functionality until it was safe to shut everything down. However, there were certain classes of failure where the safest course of action was to shut the entire system down immediately. This was generally the case where continuing to run could have made matters worse, putting patient safety at risk. I suspect that the designers of this system ran into the same problem.
- hobs 2y agoPeople posting on this forum saying "ah well software's failure case isn't as bad" > This forced controllers to revert to manual processing, leading to more than 1,500 flight cancellations and delaying hundreds of services which did operate.
- egypturnash 2y agoZero fatalities though. You could do a lot worse for a massive air traffic control failure.
- hobs 2y agoIt's true, not saying they did a bad job here, just that even minor problems in your code can exacerbated into giant net effects without you even considering it.
- lxgr 2y agoUnfortunately shutting down air traffic generally does not result in zero excess deaths: https://pmc.ncbi.nlm.nih.gov/articles/PMC3233376/ https://pmc.ncbi.nlm.nih.gov/articles/PMC3233376/
- d1sxeyes 2y agoYour source says “the fatality rate did not change appreciably”.
- lxgr 2y agoInjuries did increase, though, and I can't think of a plausible mechanism that would somehow cap expected outcomes at "injury but not death".
- d1sxeyes 2y agoSo we were talking about excess deaths, which means that supporting your argument with a paper that argues that a previous finding of excessive deaths was flawed is probably not the strongest argument you could make. Increased number of injuries but not deaths could be, for example, (purely making things up off the top of my head here) due to higher levels of distractedness among average drivers due to fear of terrorism, which results in more low-speed, surface-street collisions, while there’s no change in high speed collisions because a short spell of distractedness on the highway is less likely to result in an accident.
- ipunchghosts 2y agoTitle should be nmi
- jordanb 2y agoI do a lot of navigation and have never seen nautical miles abbreviated as "nmi."
- lxgr 2y agoI bet not everybody on here does, so picking the unambiguous unit sign would definitely avoid some double-takes.
- barbazoo 2y agoThe unit of "nm" is common among pilots but yeah technically it should be "NM".
- yongjik 2y agoNGL, two locations 3600 non-maskable interrupts apart would have been a much more interesting story.
- astrange 2y agoIt's like doing the Kessel run in less than twelve parsecs.
- buildsjets 2y agoMaybe that is true in your industry. It is not true in my industry. NM is the legally accepted abbreviation for nautical miles when used in the context of aircraft operations.
- joemi 2y agoStill, either "nmi" or "NM" would be better than the current and less correct "nm", even if "nm" is what's used in the article.
- jp57 2y agoFYI: nm = nautical miles, not nanometers.
- barbazoo 2y agoGiven the context, I'd say NM actually https://en.wikipedia.org/wiki/Nautical_mile https://en.wikipedia.org/wiki/Nautical_mile
- jp57 2y agoI was clarifying the post title, which uses "nm".
- pvitz 2y agoYes, it looks like they should have written "NM" instead of "nm".
- andkenneth 2y agoNo one is using nanometers in aviation navigation. Quite a few aviation systems are case insensitive or all caps only so you can't always make a distinction. In fact, if you say "miles", you mean nautical miles. You have to use "sm" to mean statute miles if you're using that unit, which is often used for measuring visibility.
- ianferrel 2y agoSure but I could imagine some kind of software failure caused by trying to divide by a distance that rounded two zero because the same location was listed in two databases that were almost but not exactly the same location. In fact I did when I first read the headline, then realized that it was probably nautical miles. That would be roughly consistent with the title and not a totally absurd thing to happen in the world.
- 2y ago
- jmvoodoo 2y agoSo, essentially the system has a serious denial of service flaw. I wonder how many variations of flight plans can cause different but similar errors that also force a disconnect of primary and secondary systems. Seems "reject individual flight plan" might be a better system response than "down hard to prevent corruption" Bad assumption that a failure to interpret a plan is a serious coding error seems to be the root cause, but hard to say for sure.
- mjevans 2y agoReject the flight plan would be the last case scenario, but where it should have gone without other options rather than total shutdown. CORRECT the flight plan, by first promoting the exit/entry points for each autonomous region along the route, validating the entry/exit list only, and then the arcs within, would be the least errant method.
- mcfedr 2y agoReject the plan surely should have come many places before shutdown the whole system!
- d1sxeyes 2y agoYou can’t just reject or correct the flight plan, you’re a consumer of the data. The flight plan was valid, it was the interpretation applied by the UK system which was incorrect and led to the failure. There are a bunch of ways FPRSA-R can already interpret data like this correctly, but there were a combination of 6 specific criteria that hadn’t been foreseen (e.g. the duplicate waypoints, the waypoints both being outside UK airspace, the exit from UK airspace being implicit on the plan as filed, etc).
- sandos 2y agoIf this is the case, then every system along the flightplan should pre-validate it before it gets accepted at the source?
- 2y ago
- perihelions 2y agoOriginal (2023) thread with 446 comments, https://news.ycombinator.com/item?id=37461695 https://news.ycombinator.com/item?id=37461695 ("UK air traffic control meltdown (jameshaydon.github.io)")
- amarshall 2y agoAnd the article itself in the older thread is a far more interesting read than this OP.
- Jtsummers 2y agoThere's been some prior discussion on this over the past year, here are a few I found (selected based on comment count, haven't re-read the discussions yet): From the day of: https://news.ycombinator.com/item?id=37292406 https://news.ycombinator.com/item?id=37292406 - 33 points by woodylondon on Aug 28, 2023 (23 comments) Discussions after: https://news.ycombinator.com/item?id=37401864 https://news.ycombinator.com/item?id=37401864 - 22 points by bigjump on Sept 6, 2023 (19 comments) https://news.ycombinator.com/item?id=37402766 https://news.ycombinator.com/item?id=37402766 - 24 points by orobinson on Sept 6, 2023 (20 comments) https://news.ycombinator.com/item?id=37430384 https://news.ycombinator.com/item?id=37430384 - 34 points by simonjgreen on Sept 8, 2023 (68 comments)
- perihelions 2y agoThere's also a much larger one, https://news.ycombinator.com/item?id=37461695 https://news.ycombinator.com/item?id=37461695 ("UK air traffic control meltdown (jameshaydon.github.io)", 446 comments)
- mstngl 2y agoI remembered this extensive article immediately (only that I've read it, not what and where to find). Thanks for saving me from endlessly searching it.
- steeeeeve 2y agoYou know there's a software engineer somewhere that saw this as a potential problem, brought up a solution, and had that solution rejected because handling it would add 40 hours of work to a project.
- ryandrake 2y ago... or there's a software engineer somewhere who simply assumed that three letter navaid identifiers were globally unique, and baked that assumption into the code. I guess we now need a "Falsehoods Programmers Believe About Aviation Data" site :)
- MichaelZuo 2y agoOr even more straightforward, just don’t believe anyone 100% knows what they are doing until they exhaustively list every assumption they are making.
- gregmac 2y agoWhich also means never assume the exhaustive list is 100%.
- MichaelZuo 2y agoBingo, without some means of credible verification, then assume it’s incomplete.
- Filligree 2y agoI wouldn't be able to produce such a list, even for areas where I totally do know everything that would be on the list.
- madcaptenor 2y agoEven more straightforward, just don’t believe anyone 100% knows what they are doing.
- 2y ago
- _pete_ 2y agoThe DVL really is in the details.
- spatley 2y agoHar! should have seen that one coming :)
- jrochkind1 2y agoI don't know how long that failure mode has been in place or if this is relevant, but it makes me think of analogous times I've encountered similar: When automated systems are first put in place, for something high risk, "just shut down if you see something that may be an error" is a totally reasonable plan. After all, literally yesterday they were all functioning without the automated system, if it doesn't seem to be working right better switch back to the manual process we were all using yesterday, instead of risk a catastrophe. In that situation, switching back to yesterday's workflow is something that won't interrupt much. A couple decades -- or honestly even just a couple years -- later, that same fault system, left in place without much consideration because it rarely is triggered -- is itself catastrophic, switching back to a rarely used and much more inefficient manual process is extremely disruptive, and even itself raises the risk of catastrophic mistakes. The general engineering challenge, is how we deal with little-used little-seen functionality (definitely thinking of fault-handling, but there may be other cases) that is totally reasonable when put in place, but has not aged well, and nobody has noticed or realized it, and even if they did it might be hard to convince anyone it's a priority to improve, and the longer you wait the more expensive.
- telgareith 2y agoDig into the OpenZFS 2.2.0 data loss bug story. There was at least one ticket (in FreeBSD) where it cropped up almost a year prior and got labeled "look into layer," but it got closed. I'm aware closing tickets of "future investigation" tasks when it seems to not be an issue any longer is common. But, it shouldnt be.
- Arainach 2y ago>it shouldnt be Software can (maybe) be perfect, or it can be relevant to a large user base. It cannot be both. With an enormous budget and a strictly controlled scope (spacecraft) it may be possible to achieve defect-free software. In most cases it is not. There are always finite resources, and almost always more ideas than it takes time to implement. If you are trying to make money, is it worth chasing down issues that affect a miniscule fraction of users that take eng time which could be spent on architectural improvements, features, or bugs affecting more people? If you are an open source or passion project, is it worth your contributors' limited hours, and will trying to insist people chase down everything drive your contributors away? The reality in any sufficiently large project is that the bug database will only grow over time. If you leave open every old request and report at P3, users will grow just as disillusioned as if you were honest and closed them as "won't fix". Having thousands of open issues that will never be worked on pollutes the database and makes it harder to keep track of the issues which DO matter.
- sam0x17 2y agoI've posted this here before, but they really need globally unique codes for all the airports, waypoints, etc, it's crazy there are collisions. People always balk at this for some reason but look at the edge cases that can occur, it's crazy CRAZY
- crote 2y agoComing up with a globally unique waypoint system is trivial. Convincing the aviation industry to spend many hundreds of millions of dollars to change a core data type used in just about every single aviation-related system, in order to avoid triggering rare once-a-decade bugs? That's a lot harder.
- lostlogin 2y ago> That's a lot harder. I wonder what 1,500 cancelled flights and 700,000 disrupted passengers adds up to in cost? And that’s just this one incident.
- amiga386 2y ago...an incident where they didn't parse the data as other systems already parsed the data. It sounds like the solution is better validation and test suites for the existing scheme, not a new less-ambiguous scheme
- buildsjets 2y agoThat’s not CRAZY at all. CRAZY is at 14° 4' 50.87" N. 145° 38' 16.22" E https://opennav.com/waypoint/US/CRAZY https://opennav.com/waypoint/US/CRAZY
- gadders 2y agoIf you want to, you can read the final report from the UK Civil Aviation Authority here: https://www.caa.co.uk/publication/download/23340 https://www.caa.co.uk/publication/download/23340 It's pretty readable and quite interesting.
- tempodox 2y agoWhen there's no global clearing house for those identifiers, maybe namespaces would help? Related: The editorialized HN title uses nanometers (nm) when they possibly mean nautical miles (nmi). What would a flight control system make of that?
- bigfatkitten 2y agoThe reason idents for radio navaids (VOR/NDB) are only three characters is because they are broadcast via morse code. They need to be copyable by pilots who are otherwise somewhat busy and not particularly proficient in Morse. For this purpose, they only need to be unique to that frequency within plausible radio range. 'nm' and 'NM' are the accepted abbreviations for nautical miles in the aviation industry, whether official or not.
- buildsjets 2y agoEvery aircraft I’ve ever flown as either Pilot in Command or required crewmember, and also every marine navigation system I have used in my life has displayed distance information as nm, Nm, or NM, interchangeably. I have never been confused by this, and I have never seen any other crew be confused. I have not ever seen any version of nmi used, in any variation of capitalization. This includes Boeing flight decks, Airbus flight decks, general aviation Garmin equipment, and a few MIL aircraft. And some boats.
- deleted 2y ago[deleted]
- chefandy 2y agoAs an aside, that site's cookie policy sucks. You can opt out of some, but others, like "combine and link data from other sources", "identify devices based on information transmitted automatically", "link different devices" and others can't be disabled. I feel bad for people that don't have the technical sophistication to protect themselves against that kind of prying.
- amiga386 2y agoThis is old news, but what's new news is that last week, the UK Civil Aviation Authority openly published its Independent Review of NATS (En Route) Plc's Flight Planning System Failure on 28 August 2023 https://www.caa.co.uk/publication/download/23337 https://www.caa.co.uk/publication/download/23337 (PDF) Let's look at point 2.28: "Several factors made the identification and rectification of the failure more protracted than it might otherwise have been. These include: • The Level 2 engineer was rostered on-call and therefore was not available on site at the time of the failure. Having exhausted remote intervention options, it took 1.5 hours for the individual to arrive on-site to perform the necessary full system re-start which was not possible remotely. • The engineer team followed escalation protocols which resulted in the assistance of the Level 3 engineer not being sought for more than 3 hours after the initial event. • The Level 3 engineer was unfamiliar with the specific fault message recorded in the FPRSA-R fault log and required the assistance of Frequentis Comsoft to interpret it. • The assistance of Frequentis Comsoft, which had a unique level of knowledge of the AMS-UK and FPRSA-R interface, was not sought for more than 4 hours after the initial event. • The joint decision-making model used by NERL for incident management meant there was no single post-holder with accountability for overall management of the incident, such as a senior Incident Manager. • The status of the data within the AMS-UK during the period of the incident was not clearly understood. • There was a lack of clear documentation identifying system connectivity. • The password login details of the Level 2 engineer could not be readily verified due to the architecture of the system." WHAT DOES "PASSWORD LOGIN DETAILS ... COULD NOT BE READILY VERIFIED" MEAN? EDIT: Per NATS Major Incident Investigation Final Report - Flight Plan Reception Suite Automated (FPRSA-R) Sub-system Incident 28th August 2023 https://www.caa.co.uk/publication/download/23340 https://www.caa.co.uk/publication/download/23340 (PDF) ... "There was a 26-minute delay between the AMS-UK system being ready for use and FPRSA-R being enabled. This was in part caused by a password login issue for the Level 2 Engineer. At this point, the system was brought back up on one server, which did not contain the password database. When the engineer entered the correct password, it could not be verified by the server. "
- mcfedr 2y agoBut no mention of this insane failure mode? If the article is to be believed
- deleted 2y ago[deleted]
- fyt2024 2y agoIs nm the official abbreviation for nautical miles? I assume it is natural miles. For me it is nanometers.
- andkenneth 2y agoContextually no one is using nanometers in aviation nav applications. Many aviation systems are case insensitive or all caps only so capitalisation is rarely an important distinction.
- joemi 2y agoSimilarly, no pilot in the Devil’s Lake region is using DVL to mean Deauville, and vice versa. :)
- buildsjets 2y agoOfficially, NM is the abbreviation for nautical miles when used in the context of aircraft operations. It’s not just a good idea, it’s the Law. Specifically, 14 CFR Part 1.2 of the United States Code of Federal Regulations. https://www.ecfr.gov/current/title-14/chapter-I/subchapter-A/part-1 https://www.ecfr.gov/current/title-14/chapter-I/subchapter-A...
- deleted 2y ago[deleted]
- cbhl 2y agoHmm, is this the same incident which happened last year? Or is this a new incident? From Sept 2023 (flightglobal.com): - https://archive.is/uiDvy https://archive.is/uiDvy - Comments: https://news.ycombinator.com/item?id=37430384 https://news.ycombinator.com/item?id=37430384 Also some more detailed analysis: - https://jameshaydon.github.io/nats-fail/ https://jameshaydon.github.io/nats-fail/ - Comments: https://news.ycombinator.com/item?id=37461695 https://news.ycombinator.com/item?id=37461695
- javawizard 2y agoFirst sentence of the article: > Investigators probing the serious UK air traffic control system failure in August last year [...]
- Joel_Mckay 2y agoIn other news, goat carts are still getting 100 furlong–firkin–fortnight on dandelions. =3
- convivialdingo 2y agoI guarantee that piece of code has a comment like /* This should never happen */ if (waypoints.matchcount > 2) {
- crubier 2y agoPossibly even just waypoint = waypointsMatches[0] Without even mentioning that waypointsMatches might have multiple elements. This is why I always consider [0] to be a code smell. It doesn't have a name afaik, but it should.
- gopher_space 2y agoRace condition?
- deleted 2y ago[deleted]
- CaptainFever 2y agoSilently ignoring conditions where there are multiple or zero elements?
- gitaarik 2y agoDon't you mean > 1 ?
- dx034 2y agoFrom the text it sounds like it looked up if a code was in the flight plan and at which position it was in the plan. It never looked up two codes or assumed there code only be one, just comparing how the plan was filed. I'm sure there'd be a better way to handle this, but it sounds to me like the system failed in a graceful way and acted as specified.
- GnarfGnarf 2y agoFunny airport call letters story: I once headed to Salt Lake City, UT (SLC) for a conference. My luggage was processed by a dyslexic baggage handler, who sent it to... SCL (Santiago, Chile). I was three days in my jeans at business meetings. My bag came back through Lima, Peru and Houston. My bag was having more fun than me.
- watt 2y agoWhy not pop in to a shop, get another pair of pants?
- GnarfGnarf 2y agoToo cheap... :o)
- J05ephu5M13r 2y agoIt's like déjà vu all over again, Yogi. Aug 2023: “UK air traffic woes caused by 'invalid flight plan data'” https://www.theregister.com/2023/08/30/uk_air_traffic_woes_invalid_data/ https://www.theregister.com/2023/08/30/uk_air_traffic_woes_i... -- (-11 down votes and counting)
- abstractbeliefs 2y agoThe very first line of the article states that this is a retrospective of the August '23 incident, hence the downvotes.
- J05ephu5M13r 2y ago[dead]
- ggm 2y agoCould you front-end the software with a proxy which bounces code-collision requests and limit the damage to the specific route, and not the entire systems integrity? This is hack-on-hack stuff, but I am wondering if there is a low cost fix for a design behaviour which can't alter without every airline, every other airline system worldwide, accommodating the changes to remove 3-letter code collision. Gate the problem. Require routing for TLA collisions to be done by hand, or be fixed in post into two paths which avoid the collision. (intrude an intermediate waypoint)
- kccqzy 2y agoThe low cost fix is to fix the FPRSA-R software. The front-end proxy cannot easily detect code collisions because such code collisions are not readily apparent. The ICAO flight plan allows omitting waypoints and it is one of the omitted waypoint that collides with a non-omitted waypoint. If you have taken the trouble of introducing such a sophisticated code collision detection in the front-end proxy, you might as well apply the same fix to the real FPRSA-R software.
- ggm 2y agoMy primary fear was the compliance issues re-coding the real s/w would drive this into the ground, where the compliance for a front fix might be lower. You could hash-map the collisions seen before and drive them to a "don't do that" and reduce the risk to new ones. "we fall over far less often now" Thank you for neg votes kind strangers. Remember ATC is rife with historical kludges, including using whiteout on the giant green screens to make the phosphor get ignored by the light gun. This is an industry addicted to backwards compatibility to the point that you can buy dongle adapters for dot-matrix printers at every gate, such that they don't have to replace the printer but can back-end the faster network into it. C/F rebooting a 787 inside the maximum-days-without-a-reboot
- deleted 2y ago[deleted]
- NovemberWhiskey 2y agoSo, exactly the same airline (French Bee) and exactly the same route (LAX-ORY) and exactly the same waypoint (DVL) as last September, resulting in exactly the same failure mode: https://chaos.social/@russss/111048524540643971 https://chaos.social/@russss/111048524540643971 Time to tick that "repeat incident?" box in the incident management system, guys.
- riffraff 2y agoIt's an article about that same accident
- NovemberWhiskey 2y agoD'oh - the flight number was different and the lack of year on the post had me thinking it was just the same again.
- mmaunder 2y agoThere’s little to no authentication on filing flight plans which makes this a potentially bigger problem. I’m sure it’s fixed but the mechanism that caused the failure is an assertion that fails by disconnecting the critical systems entirely for “safety”. And the backup failed the same way. Bet there are similar bugs.
- deleted 2y ago[deleted]
- aeroevan 2y agoWhat's crazy is that this hasn't happened before, waypoints that share a name isn't uncommon
- polskibus 2y agoI would’ve thought that in flight industry they got the „business key” uniqueness right ages ago. If a key is multi-part then each check should check all parts not just one. Alternatively, force all airport codes to be globally unique.
- dboreham 2y agoHeadline still hasn't been fixed? (Correct abbreviation is NM).
- mjan22640 2y agoThe title sounds like an AMD cpu issue.
- craigds 2y agooh nautical miles ! not nanometres as you might assume from being used to normal units
- jojohohanon 2y agoIs it just me or was it basically impossible to decipher what those three letter codes were?
- entropyie 2y agoInitially read this as 3600 nanometres... :-)
- junon 2y agoFor the people skimming the comments and are confused: 3600nm here is nautical miles, not nanometers. My first thought was that this was some parasitic capacitance bug in a board design causing a failure in an aircraft.
- deleted 2y ago[deleted]
- IlliOnato 2y agoWhat brought me to read this article was a confusion: how can two locations related to air traffic be 3600 nanometers apart? Was it two points within some chip, or something? Only way into the article it dawned to me that "nm" could stand for something else, and guess it was "nautical miles". Live and learn... Still, it turned out to be an interesting read)
- mirages 2y ago"and it generated a critical exception error. This caused the FPRSA-R primary system to disconnect, as designed," as designed here sounds a big PR move to hide the fact they let an uncaught exception crash the entire software ... How about : don't trust your inputs guys ?
- mkj 2y agoSounds like the kind of thing fuzzing would find easily, if it was applied. Getting a spare system to try it on might be hard though.
- muffwiggler 2y ago3600 nanometers? That's cool.
- deleted 2y ago[deleted]
- whiteandmale 2y ago[dead]
- jll29 2y agoUnique IDs that are not really unique are the beginning of all evil, and there is a special place in hell for those that "recycle" GUIDs instead of generating new ones. Having ambiguous names can likewise lead to disaster, as seen here, even if this incident had only mild consequences. (Having worked on place name ambiguity academically, I met people who flew to the wrong country due to city name ambiguity and more.) At least artificial technical names/labels should be globally unambiguous.
- cryptonector 2y ago> Just 20s elapsed between the receipt of the flightplan and the shutdown of both FPRSA-R systems, causing all automatic processing of flightplan data to cease and forcing reversion to manual procedures. That's quite a DoS vulnerability...
- klysm 2y agoI’m curious what part of the code rejected the validity of the flight plan. Im also curious what keys are actually used for lookups when they aren’t unique??
- deleted 2y ago[deleted]
- QuiltOverture 2y ago[dead]