16 ms·
British Airways: All flights cancelled amid IT crash
- crivabene 9y agoI feel close to their IT staff having to deal with this on a Saturday afternoon. Partially OT: anyone wanting to share any on-call horror stories? :)
- tyingq 9y agoAirlines are especially stressful, as the problem compounds with time. Once you hit 45 minutes or so of downtime, you start invalidating downstream flight connections. Two hours in, and you've issues with crews not being legal to fly. Four hours in, and there's not enough capacity the following day to fix the missed flights from today, etc.
- mtkd 9y agoare there commonly algos or decision support tools to help unravel that in an optimised way?
- DamonHD 9y agoI suspect that there's a lot of eating crow and bargaining with rivals for space on their flights to help clear the backlog, but each incident is going to have unique elements to resolve in it...
- user5994461 9y agoSadly that doesn't work for any serious outage. If it's only one airline down, you can get away by buying tickets on competitors. If it's multiple airlines, or the upstream reservation system, or a local meteorological incident, or an airport wide issue, etc... there are no flying plane.
- paulmd 9y agoIt's standard practice for airlines to fly each others' crews around even during normal operations (often at no cost). For example, the overbooking that led to the beating on the United flight was the result of an aircrew from another airline being booked on at the last second. I would imagine that when the shit hits the fan like this, the other airlines are very sympathetic and will do whatever it takes to get BA personnel where they need to go. After all, next month it could be their own systems that are down. (assuming it's not one of the shared services causing a global outage - not much you can do when everyone is down)
- objclxt 9y ago> For example, the overbooking that led to the beating on the United flight was the result of an aircrew from another airline being booked on at the last second. That's not quite what happened. The flight was operated by Republic Airlines, on behalf of United. Republic bumped the passenger to make way for more of their own crew - it's just that crew was going to to fly a Republic flight operating under another, different, carrier. But both crews were employed by Republic.
- tyingq 9y agoSort of. The software is called "automated re-accommodation". But you're solving for several intertwined problems. Which aircraft models / tail numbers fly which flights...they have different seating capacities, nautical range, etc. Which crews are assigned to which aircraft. They aren't all qualified to fly every model. And, they aren't all in the right city, so you have to "deadhead" them there. And, finally, which passengers go on which flights. They do have solver/optimizer algorithms, but you can imagine it's not a button press thing. There's a lot of human process, trial and error, etc. Oh, and federal laws about crew hours / legality plus union work rules. You can't just assign crews wherever you want for example...you have to consider seniority, their "home base" where they live, etc.
- mtkd 9y agoI was in the middle of some major rail cancellations a couple of months back over 2 days - on the 3rd day many of the assets were out of place around the country - they were clearly trying to solve it manually and the impact escalated over the day even though most of the original problems had cleared it's got to be a great problem to work on - and must be pretty rewarding to watch when you get it right
- gonzo 9y agoRemember on 9/11/2001 when the US FAA declared a nation-wide ground stop, and rerouted all inbound international flights to other countries (e.g. Canada https://en.wikipedia.org/wiki/Operation_Yellow_Ribbon https://en.wikipedia.org/wiki/Operation_Yellow_Ribbon)?
- CamTin 9y agoDo you have some airline/IT ops stories about that day and the aftermath? If so, I (and I'm sure others) would be delighted to hear about them.
- theoh 9y agoOne vendor of those optimization tools is Jeppesen, which about ten years ago bought up two previously independent companies in the sector, SBS (New York) and Carmen Systems (in Gothenburg, Sweden).
- Havoc 9y agoNote to self...never work in airline ops. That sounds like every day is on the edge of being a disaster.
- tyingq 9y agoShopping, flight status, etc, on their website is working, so their central reservations system isn't down. http://www.bbc.com/news/live/world-40069977 http://www.bbc.com/news/live/world-40069977 says "A BA captain has said the failure affects the passenger and baggage manifests" So it's an legality/operational thing. Passengers can get boarding passes, etc, but the plane isn't allowed to take off without proper manifests.
- tankenmate 9y agoIf that's true then it's probably their operations system, i.e. flight management, digital flight bag, etc. Without those systems you can't file flight plans easily, fuel the planes with the optimal amount for the weather conditions, etc.
- tyingq 9y agoI would guess some component of what they call "Departure Control".
- CiaranMcNulty 9y agoAccording to reports on flying forums checking is being processed manually with huge queues as a result, flight status screens aren't working, etc. Until an hour or so, ago login to the website was down for me. So it seems it's a very widespread outage that they're in the process of recovering from
- tyingq 9y agoLikely they turned both of those off on purpose. If you can't depart due to lack of a manifest, you turn off flight status and check-in. I was able to do a flight status a little while ago though.
- markonen 9y agoThe main KPI for BA's IT function is the percentage of jobs they have moved "nearshore" (to Krakow). Not that I think this is the Krakow workers' fault, mind you; rather this (and the string of similar IT incidents recently at BA) seems to fit the pattern of upper management focusing on driving down costs to the exclusion of everything else.
- crivabene 9y agoDo you happen to know if they outsourced or if they opened a subsidiary there and just moved the positions?
- markonen 9y agoIt's a group function of the International Airlines Group (BA's corporate parent). http://www.iaggbs.com http://www.iaggbs.com
- idlewords 9y agoKrakow is a big tech center, is in the EU, and it's not obvious why moving jobs there would have anything to do with a drop in quality.
- markonen 9y agoIt's not about Krakow per se. But when a company works towards specific targets like "90/10 supplier offshore ratio", rather than metrics based on quality or efficiency, I don't think quality comes first.
- kidjoedango 9y agoI'm curious if anyone can give insight on how the passenger backlog is resolved in these situations. How was it done before smart digital systems (probably circa 1980's?) and how is it handled today given all the intertwined applications. I can imagine that it must be fascinating and equally exhausting on a grand scale.
- emersonrsantos 9y ago> before digital systems (probably circa 1980's?) IBM offered ACP (Airline Control Program) with their new mainframes running OS/360 in 1965. Later they changed its name to ALCS and TPF, still used today by reservation systems.
- tyingq 9y ago>How was it done before smart digital systems Manually. The TPF reservation systems have a concept of "Queues". They would place travel records that needed to be reaccommodated onto a queue. Then, reservations agents would "work the queues" from their green screens, make phone calls, etc. >how is it handled today given all the intertwined applications Depends on the airline, and the system, but the general answer is "partially automated". The processes for less widespread issues is more automated. Think like a major storm in the northeast. Global outages are less automated because you're dealing with multiple issues, not just passengers.
- CiaranMcNulty 9y agoOne customer was told that the root cause was a "lightning strike on a datacentre", although it sounds pretty unlikely (there are storms today in the UK but surely there'd be a DR plan?)
- tyingq 9y agoThey do have DR plans, but the reality is that whole thing is a huge distributed system with parts from different vendors, in different locations, etc, with decades of legacy. A power outage the reboots a whole datacenter would screw any major airline for at least half a day. Thus far, nobody has found the economics of a modern, truly HA setup worth the cost. Outsiders greatly underestimate the complexity too. Think something like 40 disparate applications from different vendors, or some homegrown systems, in different geographical locations. Then, all the client applications are in buildings you don't own (airports) where you aren't allowed to control the infrastructure. If it were a high margin business, perhaps things would be different. It's not.
- jacques_chester 9y ago> If it were a high margin business, perhaps things would be different. It's not. Most of the loss here is not from lost bookings; a lot of people who prefer BA will probably just wait and book once the systems are restored. It's going to be from secondary losses -- lawsuits, ding to the reputation and so on. These numbers are estimatable, even for low-margin businesses.
- tyingq 9y agoI agree, but I've been around this sort of thing and seen the post incident analysis, been in the discussions, etc. Despite the huge costs of these outages, they pale in comparison to the costs of a real HA solution for all "needed to fly" applications. To give you some idea, ITA Software was bought by Google for $700 million. They were some of the best and brightest minds in this space, on par with any Silicon Valley darling. They successfully wrote a modern replacement for one popular airline function...shopping. They failed, however, at delivering a modern reservation system, despite tons of money and talent invested.
- BoiledCabbage 9y agoI have zero evidence, and am simply working through a though exercise, but it could be targeted? Maybe a group on behalf of a nation state is testing out its muscle. Or sending a warning. Maybe the UK did something recently a nation didn't like, and this is the new form of "diplomatic protest". Or sending a warning shot. I remember reading last year that the large Delta airlines outage was cyber terror related. It really feels like we're not too far off from a war between two nations without a single bullet fired. What happens when a country is hit with nation level ransomware? Ie not, "give us $300 bucks and you'll get your PC back", but "sign XYZ treaty and you get your country's water system back"? Or "we'll restore your internet and turn your powerplants back online"? How much will a wealthy country (like us in the US) be willing to stomach of seeing people going thirsty and hungry before they want their govt to capitulate? Of course the world will condemn and complain, but as has always been the case the country with the largest Army makes the rules. And we might be using the wrong measuring stick. Microsoft missed the boat by thinking (along with most of the industry) that measuring computing meant measuring PCs. It wasn't until it was too late that it realized that computing was about to be dominated by Mobile and smaller. They were using the wrong measuring stick. Are we on the verge of an era where an army should be measured in its digital strength and not its physical strength? It's a very scary thought. As an American I think first of the risks to my own country. How long could we sit with a nationwide blackout, and the internet down (for anyone using backup generators)? Food rotting in shipping containers because no cranes can offload it. No easy communication or flights for quick way to move people or things around. As a country (and as a world) I can't help but feel we're really not doing enough to take a risk like this seriously. It's just a feeling, but based on small incidents here and there it really feels to me like something in the next 5-10 years is gonna happen that'll make us all drop our jaws and say "I didn't think this was possible." Regardless of your political beliefs, a lot of people around the world didn't think last year's US election was possible. I was shocked by it as well, but it also opened my eyes to how much more really is in the realm of possible. As engineers we have much a deeper knowledge of the risks involved and as a result a much greater responsibility of raising awareness and getting the problem fixed. EDIT: I'm not sure of the original location I saw the Delta stuff, but a quick look now turns up this link. http://observer.com/2016/09/did-a-cyber-attack-ground-delta-airlines/ http://observer.com/2016/09/did-a-cyber-attack-ground-delta-... And yes, I'll repeat this above comment is purely speculative by me.
- tomschlick 9y agoAirlines should band together and form working group to redevelop the old system in modern open source tech. Yes it would probably take 5+ years to develop and roll out, but they can't keep maintaining these 50 year old mainframes that cost them tens of millions a year in downtime.
- benmarks 9y ago"Redevelop the old system in modern open source tech" is a tough sell when so many of these airlines are differentiating on technology. Myself and many other frequent fliers will not engage with companies with poor functionality and bad UX. We move or stay loyal when airline tech allows us to do & see most of the things we need.
- gjjrfcbugxbhf 9y agoThey can build a open-source back-end and differentiate in the front end. Or even build common components/ technologies together and then build their back-ends on top - that is basically what the rest of us are doing...
- deleted 9y ago[deleted]
- mbaha 9y agoTheir tech is not open enough to even allow for proper innovation and differentiation. That's the main issue with legacy IT. When you live in a world w/ Docker, REST (and every other open tech there is), you can build systems which are way more innovative.
- gaius 9y agoWhen you live in a world w/ Docker, REST (and every other open tech there is), you can build systems which are way more innovative That's quite funny because Docker and REST are just half-arsed reinventions of mainframe features from the 1970s.
- coldcode 9y agoAirlines these days are run by a complex set of systems most of which have to work in order to have the airline function. I remember a few years ago SABRE (one of the 3 major GDS companies in the world) had a 4 hour outage (I worked for a division). Half the world's airlines stopped functioning. Generally these things seldom go down but when they do, hell breaks loose. Often these systems not only do reservations, but crew scheduling, weight balancing, check-in, manifest creation and a whole host of other small but important things. Airlines also integrate some of their own pieces into the contracted ones, and some just contract almost everything. It usually comes down to money. Even Southwest Airlines which doesn't share its res data with anyone uses a GDS backend for a lot of things (in this case SABRE).
- tyingq 9y agoSouthwest was never on Sabre, per se. They had their own, separate reservation system, operated by Sabre the company, but not the main "Sabre Res System" / PSS. It was called SAAS, and sometimes "Cowboy". They moved off of that recently though, and are now on Amadeus/Altea.
- coldcode 9y agoWas true when I worked there, of course changed since then.
- drinchev 9y agoYour comment reminds me of how crazy duct-tape and bubblegum is in our IT industry.
- ams6110 9y agoPerhaps related to our calling people "engineers" who aren't anything of the sort...
- gmisra 9y agoFor reference, British Airways' parent company made a profit of €2.5 billion last year, and expects higher profits this year [0]. Without meaningful consequences at the top of the executive chain for sub-par IT/infrastructure quality, these kinds of incidents seem inevitable. But how do you hold people responsibly for "bad" software? We could adopt something akin to how PE licenses are required for civil engineering in the US. I suspect it is in the industry's best interest to address this need before a government entity decides to. [0] http://uk.reuters.com/article/uk-iag-results-idUKKBN1630MA http://uk.reuters.com/article/uk-iag-results-idUKKBN1630MA
- tim333 9y agoQuite a lot of mixed gossip in the Daily Mail article >Yesterday's issues are the fourth BA failure in the past month, with problems on June 19, July 7 and July 13 >Union leaders say hundreds of BA staff complained about 'FLY' system and most workers say it's not fit for purpose >A survey by GMB of 700 staff in June found that 89 per cent said training was poor, 94 cent suffered delays or system failures and 76 per cent said their health had suffered because of stress or anger aimed at them by frustrated passengers. etc http://www.dailymail.co.uk/news/article-3695151/Philip-Schofield-melts-Heathrow-check-computer-failure-hits-BA-passengers-starting-summer-holidays.html http://www.dailymail.co.uk/news/article-3695151/Philip-Schof...
- kbob 9y agoThat article is from July, 2016, which is why the dates "in the past month" don't compute.
- davidf18 9y agoBy using IT outsourcing in India (Tata?), Poland, BA and its parent is revealing some important information: They are cutting corners to the point of risking airline operations. The question is, where else are they unnecessarily cutting corners? Are they cutting unnecessary corners on airplane maintenance? Other quality airlines have outsourced IT or part of their IT operations, but they are careful about choosing their IT vendors. For example, Israel's national carrier, El Al, which also has extremely good security against hijacking, outsources at least its ticketing as a cost savings move, but this it does to Lufthansa the German national carrier, which uses the Amadeus Ticketing platform. El Al is saving money over running their own operations, but still using a reliable vendor. Quantas, the airline of Australia, outsources its IT operations to IBM. They did not choose the cheapest alternative, but a reliable one. Contrasting with Israeli and Australian flag carriers El Al and Quantas, BA the British flag carrier made some extremely unwise choices trying choose a "cheapest" solution, instead of a money saving cheaper solution. Compared with El Al and Quantas, BA management has shown us that quality of operations is not their top priority. The question is, where else in their operations is this a problem? BA management is revealing to customer and shareholder alike that customer service and shareholder value is not their highest priority. They are signaling to shareholders that it is time for a change before they further harm the BA brand and before some serious accident happens.
- ern 9y agoQuantas, the airline of Australia, outsources its IT operations to IBM. They did not choose the cheapest alternative, but a reliable one. I believe IBM was responsible for the Australian census debacle in 2016. Hardly a ringing endorsement of their reliability, and not the only high-profile instance of them messing things up. "No one ever got fired for choosing IBM" though.
- davidf18 9y agoWhile I can't speak to the exact circumstances, IBM when it takes over operations often hires the IT staff of the firm it is taking over operations from and runs systems on multiple-year contracts. At any rate, I believe few would contest that IBM in general is a more reliable vendor than Tata, other Indian suppliers or say, offices in Poland. EDIT: Details on IBM and Australian census. http://www.abc.net.au/news/2016-10-25/turning-router-off-and-on-could-have-prevented-census-outage/7963916 http://www.abc.net.au/news/2016-10-25/turning-router-off-and...
- mancerayder 9y agoIn the tech-related (but now fully mainstream, unlike a few years ago) news these days, there are a number of recurring themes. 1. Massive security breaches affecting major corporations, governments, and so on. 2. Massive IT outages, affecting corporations, governments and so forth. It doesn't take a lot of incidents to mark them as 'massive' since things are strongly centralized in this world. Meanwhile, in the last 15 years of my IT career, never have I seen such a strong push towards offshoring. One would think that with such a vulnerable IT landscape mixed with an unprecedented dependence on IT infrastructure, that CTO's would want to spend more and not LESS on operational costs. I'm slightly baffled.
- codeulike 9y agoThe reason is simple: short term thinking driven by short term investors
- ams6110 9y agoOr maybe it's what is repeated here so often: don't let the perfect be the enemy of the good. Perhaps it really is cheaper to have a few outages than it is to built the sort of systems that never go down. Every "nine" you add gets more and more expensive. Or perhaps the systems are so complex that it's just impossible for them to be fully reliable, and this is just one of the costs of having the ability to fly anywhere in the world, cheaply and at the drop of a hat, and to have it work 99% of the time.
- mancerayder 9y agoOr maybe it's what is repeated here so often: don't let the perfect be the enemy of the good. Is that what you tell hundreds of thousands of passengers sitting in a terminal with canceled flights and confused airline employees during one of the biggest holiday weekends? If your credit card number gets stolen or worse, your identity, and it's up to you to get that straightened out, are you just going to shrug your shoulders and write that off as the inevitable impossibility of perfection?
- riffic 9y agoHopefully, this will be a wake-up call to the industry - government regulation in IT is needed.
- al452 9y agoThe UK government (perhaps the relevant one in this case) does not have a stellar track record of figuring out how to make big IT systems and projects successful. This knee-jerk appeal to regulation is unconvincing to say the least.
- riffic 9y agoA modified form of ITIL would do a lot of good for important sectors. Let me put it another way, if the industry won't regulate itself, then they will be regulated by force.
- gaius 9y agoA modified form of ITIL would do a lot of good for important sectors You've never used ITIL for real, have you? Because if you had you would know that no amount of process can replace good engineers, and good engineers don't want to work in ITIL shops...
- Clubber 9y agoWhat's always worked in the past is hire good people, pay them well, treat them well and train them well. Pretty simple. Authoritarian type companies tend to do poorly at that though.
- id122015 9y agoif this is another case when they employed a horror coder who doesnt know to program, they'd better employ me !
- cauterize 9y agoThere might be a better word than "crashing" when describing airline computer systems malfunctioning. Nevertheless, was happy to hear it wasn't a plane.
- technofiend 9y agoMaybe they hired the managers who handled Deepwater Horizon. Although this time around, I don't think Donald Trump should send a nuclear sub to drop a nuke down the shaft on day two of the disaster. I'd wait at least a week.
- purpleostrich 9y agoWho is John Galt?
- GnarfGnarf 9y agoThis is a distant early warning of the impending demise of commercial aviation. The inexorable rise of the price of fuel will eventually make flying unaffordable for all but the military and civil servants. Airlines are already struggling to remain solvent, and are cutting costs everywhere they can: salaries, squashing passengers to increase seats/plane, nickel & dimeing us with baggage surcharges, etc. Offshoring IT is an example of cutting costs to the bone. Your grandchildren will only know of flying as an ancient legend.
- 65827 9y agoThis is nonsense, fuel prices have been plummeting amid new extraction tech and long term demand destruction. These are fundamental changes and if you're still droning on about peak oil you're just operating on old bad info.
- iaw 9y agoI wish there was a listing of company IT quality somewhere similar to the list of 2Fac financial institutions. I'm leaving my citi card for a chase one due to poor infrastructure but I could have avoided a lot of wasted time if I hadn't opened it in the first place.
- spydum 9y agoProblem is it's highly variable from engagement to engagement. A lot of the quality and poor performance could be bad process/management on the clients side (though a outsourcing company will profit hugely from).
- pbhjpbhj 9y ago>BA chief executive Alex Cruz said: "We believe the root cause was a power supply issue." // That shut down BA operations World wide? Does that seem likely? Possible? They don't have power fail-over and operational centres in different countries? I've been re-shaping my aluminium foil hat and wondering if there wasn't a specific terrorist threat that's been covered up; but then where I am there have been 2 cities suffer bomb threats (with attendance of bomb disposal and armed police) that seem to have been buried in the news completely. Also I've heard elsewhere that staff on the ground reported the incident as due to "hackers" almost immediately - the speculation being 'before they could possibly have known that' - which suggests some sort of disinformation process. /wild-speculation
- mkempe 9y agoSame thought here. Ten years ago the UK foiled a terrorist plot to bomb 6-7 airplanes in flight from the UK to the US. That's the origin of the ban on bringing large amounts of liquid on board. [1] The recent bomb in Manchester was apparently made using peroxide, too. Assuming they identified a real, immediate, and massive threat, I can see why they would prefer to ground all planes until they've sorted things out. [1] https://en.wikipedia.org/wiki/2006_transatlantic_aircraft_plot https://en.wikipedia.org/wiki/2006_transatlantic_aircraft_pl...
- rconti 9y agoCould have been a power feed to a datacenter. We can't always take newspaper descriptions as literally as we might like.
- forgottenacc57 9y agoAll those UML diagrams didn't make the software good.
- a-dub 9y agoLooks like their new (as of 2 yrs) CEO comes from the low-cost airline world: https://en.wikipedia.org/wiki/%C3%81lex_Cruz_(businessman) https://en.wikipedia.org/wiki/%C3%81lex_Cruz_(businessman) so that may be a hint... Although, I've seen this sort of thing before... Usually it starts with a middle management hiring of some alpha-male business asshole who is driven to advance his career and thinks that because he can use a spreadsheet that he's qualified to run IT. He'll then go on to sell upper management on some kind of ridiculous story straight out of some bullshit CIO magazine about how consolidating all the existing best-in-class systems into one system will cut costs and open opportunities for building and mining customer data to increase revenue. He'll get the greenlight and a shit-ton of capital, and then he has to make the decision about whether to build or buy. That question is irrelevant, as our intrepid hero has no idea what the fuck he is doing and will fuck things up regardless of which path he takes. Once the "new system" is fully half baked, he then shoves it out all over the company in some ridiculous balls to the wall no going back roll out plan. Subsequently there will be huge problems, massive lines, pissed off customers, pissed off employees... but this is where our intrepid hero really shines. His mastery of the art of bullshit successfully deflects all blame from himself and his incompetence onto the users/operators (eg, people who are responsible for revenue generating businesses) of his monstrosity. Somehow it is forgotten that outages never happened before and all the revenue loss and customer ill-will is blamed on the operators for not having well-tuned disaster recovery plans in place for manual operation. Of course the only disasters they ever dealt with were a result of his incompetence, and in a stunning feat of failing upwards, he's destined to helm the company (and then likely others) within a few short years.
- 2manyredirects 9y agoOne of the unions has been quick to attribute the issue to the outsourcing to India of some of the IT responsibility, which the right wing press here has been all too eager to publish, but BA have rebuffed this and said at this stage they believe the root cause was a power supply issue - sure, that could be attributed somewhere along the lines to an 'alpha male business asshole' (as I read in one of the comments here), but it's probably best to wait and see what the post-mortem really is rather than seek to blame someone, somewhere, be it a businessman, an Indian dev team or anything else. I am reminded of a post a while back regarding AWS' issues affecting multiple data centres (I forget the specifics), and how their post mortem didn't appropriate blame on anyone (which it really easily could have), but rather their own checks and balances, which allowed the issue to arise in the first place. I do hope that when the dust settles we see a measured response rather than a witch hunt.
- BillinghamJ 9y agoTheir system should be built in a way which makes it resistant to a power supply issue. That is the fault of however built it, whether onshore or off.
- rconti 9y agoYeah. I'm sure the "power supply issue" is an overly simplistic explanation -- it's highly unlikely a single computer PS could cause such major issues in such a huge organization. That said, having almost entirely dodged any outsourcing-related issues in the 90s, and worked with generally great offshore teams, seeing my current role impacted by an utterly shortsighted and ignorant attempt to offshore critical operations tasks is quite disheartening. It's rarely the fault of the teams, it's the fault of higher-ups who completely fail to grasp the complexity and consequences of the tasks they're offloading. Everything looks great for a few weeks or months or years until one of the dozens of things that have gone neglected rear their ugly heads. If they're lucky, they take the money and run before the crash happens and escape most blame.
- tonyedgecombe 9y ago
- CommanderData 9y agoI've managed enterprise IT applications shift support to India. I have to say when its done right, it can work well sometimes. But my experience has taught me this. The vast majority of time its a fragile egg shell waiting to crack. When it fails, it fails miraculously. An IT support team on this scale is one of those things you should keep close to your product or customer. Occasions like this I guarantee a bunch of execs somewhere in BA will pay ANYTHING to have some of their loyal IT staff back to take control of the situation.
- sgt101 9y agoBut within 4 months no one in the BA C-suite will remember this happened.
- chmaynard 9y agoHenceforward, BA passengers will be required to bring portable power supplies with them and make them available to airport personnel when asked to do so. Failure to do so will result in forced removal from the aircraft. Over and out.
- StavrosK 9y agoNah, BA isn't a US airline.
- faragon 9y agoTL;DR: yet another massive chaos because some "smart" PHB doing "cost reduction" in IT.
- m-j-fox 9y ago> "We believe the root cause was a power supply issue." I mean, fuck you. Whoever built this system to run an airline that can't tolerate someone tripping over a power cable. There should be a way to get planes in the sky using hand-written paper tickets and you know it, Jeremy. You know it god-damned well.
- oliv__ 9y agoCrash might not be the best word to put in a sentence containing British Airways. My brain was frozen for a couple seconds until I understood what had crashed.
- bitmapbrother 9y agoI can't help but be amused of the fact that all of the money British Airways thought they saved, by outsourcing their IT, was just squandered by the additional expenses it's going to cost them to fix this mess. And I'm not even factoring in all of the money this bad PR is going to cost them in the near and long term.
- GordonS 9y agoOne of the big take-aways from this is actually about how not to handle situations like this. Problems happen - even huge ones - and I think most customes understand that. What they don't understand is being given little or no information about what they should do, and also being given vague, contradicting and even false information. The BA Twitter feed seems to be the main source of dubious information. They are telling people to check the website for flight status info - but it only works intermittently, and different parts of the site say different things; for example, the flight status tool says my flight is cancelled, but the booking management tool says it's all fine. On the Twitter feed they are telling people that they are contacting people and they rebooking them automatically - but it's apparent this is happening for very few people. They are telling people to rebook on the website - which only works intermittently, and will not allow rebooking even when it is working. They are telling people to call them to rebook - but their call centres seemed to be working on normal working hours (rather than getting all hand on deck), so were not in operation between 20:00 and 09:00 (or whatever; it varies by country). During working hours, when calling any call centre anywhere in the world, you just get a recorded message and are then disconnected. Some people say they did get into the call queue, and have been waiting on hold for 7+ hours! They are sometimes telling people they shouldn't go to go to the airport to rebook, and sometimes telling them they should* go to the airport to rebook. Yesterday they were telling people they could book alternative travel with a different airline and then claim it back from them... and today they are saying they won't pay if you book alternative travel - this could cost passengers dearly. The CEO, Alex Cruz, also made a laughing stock of himself yesterday when he randomly donned a yellow high-visibility vest to do a recorded message in an office. Honestly, the whole thing has been a lesson on how not to treat your customers when things go wrong.
- known 9y agoCompanies ruined or almost ruined by Indians http://sammyboy.com/showthread.php?98021-Companies-ruined-or-almost-ruined-by-imported-Indian-labor-%28US%29 http://sammyboy.com/showthread.php?98021-Companies-ruined-or...
- haveopinion 9y agoThe facts are clear. No effective mirror system or failover system. I've architected major 365x24 systems repeatedly in my career: Removing the possibility of power supply failure across the primary and secondary sites (and even within single site) is absolutely fundamental. That clearly has not been done here, at least to the extent of making sure that the solution is effective- and the blame I would guess lie within a crazy management and decision structure that seems to increasing permeate IT within large companies. Tata do have some good staff (from my first hand experience over many years):However, from my more recent experience, there is an increasingly MASSIVE CLIFF EDGE in talent/competence in within Indian companies generally, (and to some lesser extent UK companies as well): So the possibility of having "very assertive, but stupid or incompetent" middle and even senior managers is becoming an ever greater real-world problem. The criteria and method for selecting candidates for IT related work does not help - in particular role, competency, and "modus operandi" of recruiting agencies, is pretty atrocious, as many a seasoned contractor will attest! But ultimately, I think Alex Cruz and his CIO must take responsibility for a lack of diligence and competence. They clearly do not understand the most basic truth of high reliability, mission-critical IT: i.e. They NEVER needed cheaper IT staff (with all the attendant risk within their industry) , but rather much fewer technical staff, of the HIGHEST QUALITY, for all BA's key systems. And that I’m afraid means probably not putting Indian companies at the top of the list.
- frik 9y agoApparently whole Heathrow airport, .. was shut down yesterday. United and Lufthansa has to cancel their flights there too. So it meant more passengers trying to get with other plans to neighbour countries.