top of page

Another Day, Another ATC Failure — And It’s Not South Africa

9 hours ago
6 min read

By Garth Calitz


South African aviation professionals and passengers have become accustomed to the occasional air traffic control system issue, technical glitch or infrastructure-related operational surprise; there is something strangely reassuring about discovering that it is not only happening here.

The UK's aviation system has just provided its own reminder that even countries with considerably deeper pockets, sophisticated infrastructure and an air traffic management system regarded as one of the world's most advanced can occasionally discover that computers, connectivity and aviation schedules do not always play nicely together. National Air Traffic Services (NATS) suffered another significant technical failure on Monday, 21 September, just 13 days after a separate software defect at its London control centre caused more than 2,000 flight cancellations and disrupted the travel plans of hundreds of thousands of passengers. This time, the problem was at NATS' Prestwick Centre in Scotland. And yes, before anyone gets too smug about it, the UK also has air traffic control problems. Not only in South Africa.

The latest NATS problem was described as a connectivity issue affecting part of its systems network at Prestwick. NATS said engineers resolved the problem during the morning and that the system subsequently returned to full capacity. The organisation was also quick to point out that Monday's problem was unrelated to the 8 September incident. Technically, that distinction is important. Operationally, however, passengers are unlikely to be comforted by the knowledge that the computer that failed this week was a different computer from the computer that failed last week. It is rather like being told by your mechanic that the car has stopped because the engine and gearbox are completely unrelated components. True, perhaps. Still not terribly useful when you're standing next to it.

The Prestwick problem affected operations across Scotland, Northern Ireland and northern England. British Airways and easyJet were among the airlines forced to cancel flights, while Cirium recorded 227 cancellations by late Monday afternoon. And that was despite the technical problem itself being fixed relatively quickly. This is one of the less glamorous realities of modern aviation: an air traffic control system can return to normal operation, but the airline network does not necessarily follow immediately. Aircraft are in the wrong places, crews are displaced, connecting passengers miss flights, airport slots disappear and flight rotations unravel. The computer says everything is fine. The departure board begs to differ.


There is a certain irony in watching British airlines complain about an ATC technical failure. For South African aviation, the words “technical issue”, “system failure”, “operational disruption” and “we are working to restore normal operations” have become sufficiently familiar that they could probably be added to the national aviation dictionary. We have experienced our own ATNS and airport infrastructure challenges, including ATC system outages and operational restrictions that have left airlines and passengers picking up the pieces. So when British Airways says it is disappointed about “yet another technical fault”, the South African response might reasonably be: “Welcome to the club. Ons gaan nor Braai”

The difference is that in Britain the outrage tends to arrive with a parliamentary inquiry, a regulator's statement and several thousand words in the newspapers. In South Africa, depending on the day, the first response may simply be: “Eish. The system is down again.” The underlying aviation principle, however, is identical. Air traffic control is not an optional service. If the system cannot safely handle aircraft, traffic restrictions have to be imposed, and safety comes first. Always. The problem comes afterwards, when somebody has to explain why restoring the system took minutes but restoring the airline schedule takes hours or days. That particular problem is not uniquely British, and anyone who has watched an airport departure board deteriorate from “On Time” to “Delayed” to “Cancelled” will understand the frustration.


The timing of Monday's Prestwick problem could hardly have been worse for NATS. Only three days earlier, the organisation had published its preliminary report into the major 8 September incident. That failure was caused by a software defect in a small subsection of the National Airspace System. According to NATS, the system allocates codes used to identify aircraft on radar. During the processing of a manual aircraft-code request, a higher-priority message interrupted the process. When the original request resumed, the software defect prevented it from restarting correctly, resulting in corrupted output that affected subsequent flight-data updates. The entire sequence apparently unfolded within milliseconds. Unfortunately, the consequences took considerably longer to disappear.

NATS imposed restrictions to maintain safety, with the problem affecting flights operating through the London Area Control Centre. The restrictions lasted approximately six hours. The technical problem was eventually resolved, but the resulting disruption continued for considerably longer, with more than 2,000 flights cancelled and hundreds of thousands of passengers affected. NATS subsequently identified a permanent software fix and said it was undergoing safety testing. Additional mitigation measures were also introduced to reduce the risk of another failure and improve recovery capability. Then, three days after that report was published, another NATS control centre developed a technical problem. Different system, different location, different fault — same industry-wide headache.


The airlines are understandably becoming less interested in the technical distinctions. easyJet said the latest disruption once again raised questions about the resilience of NATS' systems and called for firm action to prevent repeated failures. British Airways described it as another technical fault involving NATS and expressed its disappointment. Ryanair went further, renewing its call for NATS chief executive Martin Rolfe to resign, while saying more than 25,000 of its passengers had been affected and more than 140 of its flights had been delayed. The resignation demand is Ryanair's position, rather than an established finding about NATS management, but it illustrates just how irritated airlines have become.

Airlines are ultimately the organisations dealing face-to-face with passengers when things go wrong, even when the original problem has occurred somewhere entirely outside the airline. The passenger generally does not care whether the fault was caused by an ATC computer, an airport system or a communications network; they simply want to know why their flight is still sitting on the ground, and preferably when it is going to move. From the airline's perspective, meanwhile, an ATC failure can create high operational costs, disrupt aircraft utilisation, complicate crew planning and create a chain reaction extending well beyond the geographical area initially affected. In other words, a computer in Scotland can eventually become somebody else's problem in London, Manchester, Belfast or even further afield.

This is perhaps the most important lesson from the two September incidents. NATS can fix a technical problem, but that does not mean the airline network has instantly recovered. A single aircraft being delayed at the wrong point in a rotation can affect several subsequent sectors. A crew reaching its legal working limit can require another crew to be found. A missed connection can create another problem at another airport. Suddenly a technical failure in Scotland is affecting passengers who have never been anywhere near Scotland. That is the nature of a modern airline network, and it is also why resilience is becoming such an important issue in aviation infrastructure.


The system does not necessarily have to prevent every failure — that would be unrealistic for any complex technological network — but it needs to prevent a relatively contained failure from becoming a national operational crisis. This distinction is important because aviation is one of the few industries where a system failure cannot simply be managed by telling everybody to try again later. Aircraft have fuel limitations, crews have legal duty-time limits, airports have capacity constraints and air traffic controllers have safety requirements. Once the carefully balanced system begins to unravel, putting it back together is considerably more complicated than switching the offending computer off and on again.


The UK's Civil Aviation Authority has also weighed in on the passenger-rights issue. The CAA says the Prestwick incident is likely to constitute an “extraordinary circumstance” under passenger-rights regulations. That means passengers are unlikely to qualify for financial compensation for delays or cancellations directly caused by the incident. Airlines nevertheless remain responsible for looking after affected passengers and providing the applicable refund or rerouting options. The distinction between compensation and the airline's obligation to provide care is important, although it may be of limited comfort to somebody who has spent six hours in an airport watching their holiday disappear one cancelled flight at a time.


South African passengers may recognise another familiar concept here: “We apologise for the inconvenience.” Four of the most dangerous words in aviation. They are usually followed by a meal voucher, a revised departure time and the increasingly optimistic promise that the aircraft will be boarding “shortly”. Of course, airlines themselves are often victims of these disruptions, and it would be unfair to suggest that every cancellation is within their control. But passengers quite reasonably judge the aviation system by the experience they have in front of them, not by the technical architecture buried somewhere behind the scenes.


For South African readers, there is perhaps a small consolation. When the next ATC system failure occurs locally and somebody says, “These things only happen here,” we can now confidently point north. Apparently they don't. They happen in Britain too. Not only in South Africa. The difference is that when it happens in Britain, somebody drinks tea while writing the investigation report. When it happens in South Africa, we usually drink tea while waiting for the system to come back. Either way, the aircraft remain on the ground.


Comments


Archive

bottom of page