10 ms·
Airbus A320 – intense solar radiation may corrupt data critical for flight
- ChrisArchitect 10mo agoMore discussion: https://news.ycombinator.com/item?id=46082296 https://news.ycombinator.com/item?id=46082296
- jMyles 10mo agoThis is one of the rare cases where, IMO, it makes sense to use a modified title as you've done here.
- jb1991 10mo ago[flagged]
- deleted 10mo ago[deleted]
- op00to 10mo agoSolar radiation like solar wind, or sunlight? They don’t say.
- mr_toad 10mo ago“Analysis of a recent event” I presume they mean a Coronal Mass Ejection.
- fwip 10mo agoI feel like the event was something that happened to a plane. That said, I wouldn't think sunlight would be penetrating to the chips running the plane.
- dtagames 10mo agoGamma rays penetrate everything and have definitely been known to disrupt computer circuits.
- fwip 10mo agoYes, which is why the solar flare scenario makes more sense.
- awesome_dude 10mo ago> The grounding of Airbus A320neo aircraft around the world can be traced back to an incident on a JetBlue flight operating a Cancun to New Jersey service on 30 October. > At least 15 passengers were injured and taken to the hospital after a sudden drop in altitude on the flight from Mexico was forced to make an emergency landing in Florida, US aviation officials said at the time. > The Thursday flight from Cancun was headed to Newark, New Jersey, when the altitude dropped, leading to the diversion to Tampa International Airport, the US Federal Aviation Administration said in a statement. > Pilots reported “a flight control issue” and described injuries including a possible “laceration in the head,” according to air traffic audio recorded by LiveATC.net. > Medical personnel met the passengers and crew on the ground at the airport. Between 15 and 20 people were taken to hospitals with non-life-threatening injuries, said Vivian Shedd, a spokesperson for Tampa Fire Rescue. > Pablo Rojas, a Miami-based attorney who specialises in aviation law, said a “flight control issue” indicated that the aircraft wasn't responding to the pilots. https://www.stuff.co.nz/travel/360903363/what-happened-flight-sparked-grounding-airbus-a320neo-flights-around-world https://www.stuff.co.nz/travel/360903363/what-happened-fligh...
- lostlogin 10mo ago> At least 15 passengers were injured and taken to the hospital after a sudden drop in altitude on the flight from Mexico was forced to make an emergency landing in Florida, US aviation officials said at the time. I’m surprised passengers are allowed to unbuckle for so much of each flight. You can get injured while buckled it, but that seems less common.
- bparsons 10mo agoThere was a very large CME ten days ago. The NOAA scale had predicted a high likelihood of disruptions, and had specifically suggested that spacecraft and high altitude aircraft could be impacted. https://www.swpc.noaa.gov/noaa-scales-explanation https://www.swpc.noaa.gov/noaa-scales-explanation https://kauai.ccmc.gsfc.nasa.gov/CMEscoreboard/prediction/detail/5065 https://kauai.ccmc.gsfc.nasa.gov/CMEscoreboard/prediction/de...
- glaucon 10mo agoFWIW the "industry sources say" line on the incident is that it occurred on 30 October[1], so further back than ten days ago but of course there may have been other CME incidents at that time. The European Agency Aviation Safety Agency [2] instruction describes the characteristics of the incident but not the date. [1] https://www.theguardian.com/business/2025/nov/28/airbus-issues-major-a320-recall-after-recent-mid-air-incident https://www.theguardian.com/business/2025/nov/28/airbus-issu... [2] https://ad.easa.europa.eu/ad/2025-0268-E https://ad.easa.europa.eu/ad/2025-0268-E
- nQQKTz7dm27oZ 10mo ago[dead]
- tyingq 10mo agoWould guess "cosmic rays". Sun Microsystems had a batch of UltraSparc servers that were very sensitive to it, and it was a big issue. https://docs.oracle.com/cd/E19095-01/sf4810.srvr/816-5053-10/816-5053-10.pdf https://docs.oracle.com/cd/E19095-01/sf4810.srvr/816-5053-10... https://en.wikipedia.org/wiki/Cosmic_ray https://en.wikipedia.org/wiki/Cosmic_ray
- addaon 10mo agoI’d really, really like to know what microcontroller family this was found on. Assuming that this is a safety processor (lockstep, ECC, etc) it suggests that ECC was insufficient for the level of bit flips they’re seeing — and if the concern is data corruption, not unintended restart, it means it’s enough flips in one word to be undetectable. The environment they’re operating in isn’t that different from everyone else, so unless they ate some margin elsewhere (bad voltage corner or something), this can definitely be relevant to others. Also would be interesting to know if it’s NVM or SRAM that’s effected.
- jayanmn 10mo agoI am worried about a software fix for what looks like hardware problem.
- deleted 10mo ago[deleted]
- afavour 10mo agoIt could be as simple as storing multiple copies of the relevant data and adding a checksum, something like that. Hardware fix is the ultimate solution but it might be possible to paper over with software.
- deleted 10mo ago[deleted]
- kachapopopow 10mo agosoftware fixes are totally fine since the chance of two redundant pairs failing within the time it takes to correct these errors is more zero's than there are atoms in the universe. (each pilot has a redundant computer and because there's two pilots there's two redundant pairs)
- themerone 10mo agoGracefully handling hardware faults is a software problem. The Air France Flight 447 crash was the result of bad software and bad hardware.
- qaq 10mo agoHas BoFesc vibes "It's friday, so I get into work early, before lunch even. The phone rings. Shit! I turn the page on the excuse sheet. "SOLAR FLARES" stares out at me. I'd better read up on that..."
- owenthejumper 10mo agoA friend works at Jetblue. They are scrambling hard to do the updates.
- viiralvx 10mo agoI was traveling during this entire ordeal. My flight got delayed by 7 hours. Insane day, just now boarding my flight. American Airlines was in shambles today.
- jfoster 10mo agoI've noticed that some carriers seem to be suggesting that there might be no impact to flights, but isn't this an immediate grounding for each aircraft until the update is made? How is it possible that this wouldn't impact upon flight schedules?
- arrel 10mo agoN of 1, but I’m stuck in phoenix overnight because our flight was delayed an hour and a half by airbus maintenance and we missed our connection.
- icegreentea2 10mo agoThe grounding is for 6000 of 11000 A320 series. I believe it's some combination of software and hardware configuration that is at risk.
- jfoster 10mo agoThank you; that makes sense. I had the impression it was the entire fleet.
- julik 10mo agoIt depends on whether the ELAC is an LRU (line-replaceable unit, i.e. a box with ports that can be swapped at an airport) and whether a software update can be uploaded into a unit that is installed (not all aircraft have a "firmware update via cable or floppy", so to speak)
- simne 10mo agoIf possible for exact this plane, could make software update just as routine procedure. But as I hear, air transporters could buy planes in different configurations, so for example, Emirates airlines, or Lufthansa always buy planes with all features included, but small Asian airlines could buy limited configuration (even without some safety indicators). So for Emirates or Lufthansa, will need one empty flight to home airport, but for small airline will need to flight to some large maintenance base (or to factory base) and wait in queue there (you could find in internet images of Boeing factory base with lot of grounded 737-MAXes few years ago). So for Emirates or Lufthansa will be minimal impact to flights (just like replacement of bus), but for small airlines things could be much worse.
- kappi 10mo agoFollowing the Airbus A320 emergency airworthiness action, everyone will be talking about the ELAC (Elevator Aileron Computer) manufactured by Thales, which caused a sudden pitch-down without pilot input on JetBlue 1230 back in October. So here’s everything you need to know about ELAC. The ELAC System in the Airbus A320: The Brains Behind Pitch and Roll Control https://x.com/Turbinetraveler/status/1994498724513345637 https://x.com/Turbinetraveler/status/1994498724513345637
- ThePowerOfFuet 10mo agohttps://xcancel.com/Turbinetraveler/status/1994498724513345637 https://xcancel.com/Turbinetraveler/status/19944987245133456...
- minitoar 10mo agoWe flew too close to the sun
- rvz 10mo ago[flagged]
- joelthelion 10mo agoDo they really need to ground the entire fleet for that? One incident for ten thousand planes in the air for years. I'd think that giving airlines two months to fix it would be sufficient.
- mrpippy 10mo agoI don’t believe it’s been years, only the latest firmware version for the ELAC is affected. The fix is to downgrade (or replace hardware with a unit running earlier firmware)
- kijin 10mo agoI imagine it could help with Airbus marketing. "We take proactive measures, whereas our competitor only takes action after multiple fatal crashes!"
- brabel 10mo agoImagine an airplane crashed in these 2 months. I bet you would join the chorus and blame them for gross negligence.
- kijin 10mo agoThere's a huge difference between "manufacturer recommended updates, but airline waited until the last week to apply them" and "manufacturer didn't even acknowledge the issue" in terms of who the chorus is going to blame.
- probably_wrong 10mo agoI know someone who is stranded in another continent thanks to this. Trust me, all the understanding I could have as a technical user has been offset by the MASSIVE pain in the ass that is rebooking an international flight. And non-technical users have heard "the plane will not travel because it requires a software update", which does not inspire confidence. As far as I'm concerned it has not helped with their marketing.
- 65a 10mo agoThere's a great postmortem here about what might have been a similar SEU (single event upset--bitflip) here: https://www.atsb.gov.au/sites/default/files/media/3532398/ao2008070.pdf https://www.atsb.gov.au/sites/default/files/media/3532398/ao...
- pyb 10mo agoThe aerospace industry has had countermeasures in place against bit-flips for a long time, oftentimes thanks to redudancy Airbus/Thales's fix in this case appears to add more error checking, and to restart the misbehaving component. https://bea.aero/fileadmin/user_upload/BEA2024-0404-BEA2025-0020-BEA2025-0179-FR.pdf https://bea.aero/fileadmin/user_upload/BEA2024-0404-BEA2025-... ("une supervision interne du composant à l’origine de la défaillance ; - un mécanisme de redémarrage automatique de ce composant dès lors que la défaillance est détectée)
- nolist_policy 10mo agoThe linked document is not related to this incident.
- raverbashing 10mo agoApparently the fix is reverting to a previous version of the SW (see https://avherald.com/h?article=52f1ffc3&opt=0 https://avherald.com/h?article=52f1ffc3&opt=0 ) Curious what a sw change might have done in terms of resiliency. Maybe an incorrect memory setting or some code path that is not calculating things redundantly maybe?
- rootusrootus 10mo agoSo it's not just Boeing that can screw up software on an airplane. I guess now I have to be a little afraid of all the airliners.
- rene_d 10mo agoThe Aviation Herald has more technical details: https://avherald.com/h?article=52f1ffc3&opt=0 https://avherald.com/h?article=52f1ffc3&opt=0
- loxodrome 10mo agoThanks for the link. This line in particular is concerning. "This identified vulnerability could lead in the worst case scenario to an uncommanded elevator movement that may result in exceeding the aircraft structural capability."
- isodev 10mo agoWell, I think in the grand scheme of things (including on the ground), the range of safety faults that can be triggered by a simple bitflip at the wrong moment range from inconvenient to absolute disaster. So in that sense, I'm very happy that Airbus has managed to identify opportunities to improve their design to be even more resilient.
- nickdothutton 10mo agoI’d just like to point out that if you are in the computing industry long enough, you will get to see a few such incidents under different circumstances, not only in industries like aerospace. Mostly things like ECC save your a*, sometimes your software will be able to recognise a temporary spurious reading and disregard it because you had enough alternative checking logic, or in the case of realtime and safety critical maybe even your systems can take a vote between them. Got caught out by (cpu cache line) bit flips in the 90s, months of pain trying to track it down. Some of your will know :-)
- LadyCailin 10mo agoWe noticed this in our logs once! We service a huge amount of traffic, and as part of that, we log what is effectively an enum. We did a summarization of this field once, and noticed that there were a couple of “impossible” values being logged. One of my coworkers realized that the string that actually got logged was exactly one bit off from a valid string, and we came to the conclusion that we were probably seeing cosmic rays in action, either in our service, or in the logging service.
- deleted 10mo ago[deleted]
- tuetuopay 10mo agoI had a similar story on my NAS that got one btrfs path corrupt. Plopped in on the btrfs IRC, one of the devs noticed the inconsistency was one bitflip away from the right value. Incredibly they were able to give me the right commands to fix it! Got to give credit where it is due, btrfs took the safe path and refused to touch the affected directory until fixed, and has enough tooling to fix this. I won’t blame cosmic rays but more likely dying RAM. The NAS now runs ECC memory.
- Philip-J-Fry 10mo agoI also saw a similar thing. I also naively pointed at "cosmic rays". It wasn't until someone found the actual bug that I realised how unlikely that was. The actual bug was unsafe code somewhere else in the application corrupting the memory. The application worked fine, but the log message strings were being slightly corrupted. Just a random letter here and there being something it shouldn't be. The question really should have been, if this was truly cosmic interference, why only this service and why was the problem appearing more than once over multiple versions of the application? Cosmic rays are a great excuse to problems you don't yet understand. But the reality of them is extremely rare and it's like 99% a memory corruption bug caused by application code.
- supernova87a 10mo agoI wonder how the incident was diagnosed? Does the FDR record low level errors that might've contributed to this? I thought that it only recorded certain input parameters and high-level flight metrics but I'm no expert. If a radiation event caused some bit-flip, how would you realize that's what triggered an error? Or maybe the FDR does record when certain things go wrong? I'm thinking like, voting errors of the main flight computers? Anyway, would be very interested to know!
- yread 10mo agoFrom a comment on avherald: "Had the same problem with low power CMOS 3 transistor memory cells used in implantable defibrillators in the 1990s. Needed software detection and correction upgrade for implanted devices, and radiation hardening for new devices. Issue was confirmed to be caused by solar radiation by flying devices between Sydney and Buenos Aires over the south pole multiple times, accumulating a statistically significant different error rate to control sample in Sydney."
- jakub_g 10mo agoFrom newspaper reporting on this, they are rolling back a software update. I wonder what was the original cause or the update? How often are flight computers software updated and why?
- julik 10mo agoThis ELAC version is 100-something, and the A320 first flew around 1988. Why the updates - for example, there are updates to flight control law transitions, like after 1991 where the aircraft would limit flight control inputs during landing, thinking it would be preventing a stall - because it would not go into the flare law appropriately. See https://en.wikipedia.org/wiki/Iberia_Flight_1456 https://en.wikipedia.org/wiki/Iberia_Flight_1456 The cause could have also been an extra check introduced in one of the routines - which backfired in this particular failure scenario.
- 1970-01-01 10mo agoThey said the same thing at Toyota when the unintended accel problem was in the news, but never found a real world example. There are a lot more old Toyotas still on the road than Airbuses in the air, so distance to the sun makes all the difference here? I wonder if they only see issues when flying near the north pole?
- nubinetwork 10mo agoWhy would a CME disrupt a single brand and model of aircraft, when the entire planet is covered in computers that almost never have bitflip issues when a CME rolls through every few months?
- 1970-01-01 10mo agoI'm guessing EM shielding flaw or something electronic. See my comment on the Toyotas. It doesn't make sense from a raw probability perspective.
- squarefoot 10mo agoI would guess barely enough cable shielding paired with long enough paths along the aircraft so that the signals there would be more likely affected by EM induced currents.
- jpollock 10mo agoThe design of the system is very interesting, particularly how it expects to handle errors. In 90's Telco, you used to have a pair of systems and if they disagreed, they would decide which side was bad and disable it. In modern cloud, you accept there are errors. There's another request in ~10+ms. You only look when the error rate becomes commercially important. My understanding of spacecraft is that there would be 3 independent implementations and they would vote. The plane has a matrix of sensors and systems, allowing faults to be bubbled up and bad elements disabled independently. The ADIRU does compare values to detect failures (median of 3 sensors), but they could only detect errors that last >1s. The flight computer used the raw data - because the sensors aren't interchangeable (they won't have consistent readings in all flight modes)! Very nifty. One thing, they say "memorisation period", I don't think it's a memorisation period? From my reading of the algorithm, it should be more "last value retention period"? Or "sensor spurious fault reading delay"? Section 2.1 A330/A340 flight control system design "AOA computation logic" https://www.atsb.gov.au/sites/default/files/media/3532398/ao2008070.pdf#%5B%7B%22num%22%3A837%2C%22gen%22%3A0%7D%2C%7B%22name%22%3A%22FitH%22%7D%2C207%5D https://www.atsb.gov.au/sites/default/files/media/3532398/ao...
- jpollock 10mo agoFor example.... "Preliminary A330/A340 FCPC algorithm" "The algorithm did not effectively manage a specific situation where AOA 2 and AOA 3 on one side of the aircraft were temporarily incorrect and AOA 1 on the other side of the aircraft was correct, resulting in ADR 1 being rejected." So, you've got a system where _two_ of the three sensors are bad, and you need to deal with it.
- Loudergood 10mo agoI'm in awe of the fact that two sensors can be wrong AND agree with each other.
- Nextgrid 10mo agoThose being analog sensors measuring analog, physical things, they will never exactly agree with each other; so there's a plausibility window. As long as the fault causes the sensors to remain within said window they will be considered as valid.
- rishabhaiover 10mo agoI hope Airbus only uses Honeywell or Collins in their newer planes.
- p_l 10mo agoIt was Honeywell parts that failed in previous two events (2006 and 2008) EDIT: properly credited, it was Thales Honeywell I guess :)
- skx001 10mo agoThis video shows the the A320 computer and how the computer cooling system works https://www.youtube.com/watch?v=HQuc_HhW6VA https://www.youtube.com/watch?v=HQuc_HhW6VA
- oofbey 10mo agoThis is in response to JetBlue flight 1230 from Cancun to Newark on October 30, 2025, where a cosmic ray of some kind flipped a bit and caused a dangerous situation. At the time there was a minor (G1) geomagnetic storm - meaning more cosmic rays than normal. The Planetary K-index was at 5. These are somewhat elevated numbers - enough to produce a visible Aurora in Canada, but probably not even the northernmost US. But also this level of space weather is also very common. We hit G1 or higher about once a week. That's the really damning part. If it had happened in a G4 or G5 storm, then the engineers might have responded "we can't fix everything", but this level of reliability is clearly unacceptable.
- rossjudson 10mo agoMy armchair guess is that they had a new control pathway not properly participating in their integrity hand-off protocols, doing some kind of transformation outside of that protection. I once saw some HW engineers go nuts trying to find out why a storage device had an error rate several orders of magnitude higher than the extremely low error rate they expected (and triggering data corruption errors). It turns out to be one extremely deep VHDL-based control area for an FPGA that didn't properly do integrity. You'd have to flip a bit at an incredibly precise point in time for error to occur, but that's what was happening. When all the math was said and done, that FPGA control path integrity miss exactly accounted for the the higher error rate.
- albert_e 10mo agoWhat if future aircraft had "OTA" updates to software... using this as an example of avoidable downtime. OTA updates to cars makes me feel uneasy -- not knowing what new bugs it might introduce.
- p_l 10mo agoUpdates are distributed online, but applied by technician with a portable data loader device (a computer with special - partially standardized - interface cable). The actual update in this case is about 15 minute work, the often seen "2h" claim is the entire cycle of powering the aircraft down for maintenance, update, verification, etc.
- asdefghyk 10mo agoIntense solar radiation will be at a peak, since it is NOW the peak of the 11 year sunspot cycle. Good related reading on this page .... https://en.wikipedia.org/wiki/Radiation_hardening https://en.wikipedia.org/wiki/Radiation_hardening ... includes a range of mitigation effects. ( I would be interested to find out how they actually test these systems. What combinations of hardware hardening and software logic. ALso do they actually subject to system to radiation as part of the testing ) https://en.wikipedia.org/wiki/Radiation_hardening https://en.wikipedia.org/wiki/Radiation_hardening
- p_l 10mo agoIn early 1990s, when possibility of radiation caused upsets came to airliners, tests were conducted with various kinds of radiation including firing heavy ions from particle accelerators