When digital communications became safety-critical: How a network switch failure affected control system logic and operator situational awareness
Although I am not an industrial networking specialist, this blog addresses industrial networking issues. I therefore invited my colleague Dominic Iadonisi, who has more than 25 years of industrial networking experience, to co-author this article.
Historically, communications networks merely transported information between otherwise independent control devices and networks. Modern integrated control systems increasingly rely on continuous digital communications for closed-loop control, distributed logic, synchronization, alarms and operator interfaces. Consequently, failures or delays in the communications infrastructure can directly influence process behavior rather than simply interrupting data display.
Industrial facilities increasingly depend on digital communications to monitor and control physical processes. When those communications degrade or fail, the consequences can extend far beyond information technology, directly affecting equipment operation, operator situational awareness and facility safety. As a result, digital communications systems have become safety-critical engineering components. Nuclear power plants provide perhaps the clearest illustration of the impacts because of their stringent safety and reporting requirements. A recent example occurred on Apr. 23, 2026, at the Edwin I. Hatch Nuclear Power Plant, Unit 1, near Baxley, Georgia.
I am a nuclear engineer and have been concerned with nuclear safety throughout my career. I supported the U.S. Nuclear Regulatory Commission through Pacific Northwest National Laboratory in developing NRC Regulatory Guide 5.71 on nuclear power plant cybersecurity. I also served for 12 years as the managing director of the international control system cybersecurity standards development effort – the ISA/IEC 62443 series of standards.
Generally, control system cyber incidents are a combination of technology and people issues. The Hatch Unit 1 case was both.
The plant was operating at 100% power which meant the turbine control valves modulate to maintain the reactor pressure at a setpoint. Approximately 40 minutes prior to the scram, plant staff replaced a failed network switch in the GE Mark VI turbine control network.
According to U.S. Nuclear Regulatory Commission (NRC) Edwin I. Hatch– Integrated Inspection Report dated July 16, 2026, there did not appear to be adequate testing to assure the configuration and integrity of the replacement network switch. The network switch was installed with Rapid Spanning Tree Protocol disabled, creating a network loop that produced a data storm. The data storm overwhelmed turbine controller communications and caused the turbine control valves to fail closed at 100% power. (A network data storm occurs when excessive broadcast, multicast, or other repeated Layer-2 traffic overwhelms network capacity.) The network data storm delayed communication between system parameters and the Mark VI control system. The main control room reported slow or frozen turbine/generator HMIs, controller faults and communication failures consistent with severe network problems. When the turbine control valves started closing there was an accompanying increase in reactor pressure. Because the Mark VI operator displays and alarms were degraded by the network problems, operators were not aware of the valve movement and the increasing reactor pressure in sufficient time to intervene before the automatic scram.
Unlike a failed sensor or actuator, a network infrastructure failure can simultaneously disrupt multiple controllers, HMIs, alarms, engineering workstations and operator situational awareness. Communications infrastructure therefore represents a potential common-mode failure affecting numerous plant functions at once.
GE Mark VI turbine control systems are widely deployed in power generation, oil and gas production, chemical processing, and renewable energy facilities. Consequently, the engineering lessons from this event extend well beyond the nuclear industry to any industrial facility that relies on networked turbine or other integrated control systems. The Hatch event also demonstrated the value of independent process sensor monitoring that is not dependent on the affected Ethernet network or operator displays. Such monitoring could have alerted operators that reactor pressure was increasing abnormally despite the loss of the Mark VI HMIs and alarms.
The Hatch Unit 1 event illustrates several safety implications of relying on digital communications for turbine/integrated control systems:
- Digital communications infrastructure can become a common-mode failure affecting control functions, operator information and plant operations simultaneously.
- Loss or degradation of communications through a network switch can affect plant operation such as unintended control actions, including unexpected valve movement.
- Operators may lose situational awareness when HMIs and alarms freeze or respond slowly.
- Automatic safety functions may continue to execute even while operators cannot observe plant response in real time.
- Loss of operator visibility can eliminate opportunities to intervene before an automatic safety action occurs.
- If manual control paths also depend on the affected network, recovery options become limited.
Similar vulnerabilities are likely to exist wherever turbine or integrated control systems rely on comparable Ethernet architectures and switch configurations. It also remains unclear what other plant systems, including turbine ancillary systems, could be affected by similar network failures.
Get your subscription to Control's tri-weekly newsletter.
Network switch issues including Spanning Tree failures have impacted networked control systems in multiple critical infrastructures. Spanning Tree failures have occurred during switch replacement, maintenance activities, topology changes, accidental cable loops and incorrect bridge priority configuration. These events have repeatedly produced broadcast storms that disrupted industrial control systems and enterprise networks. Examples of network communication issues that affected control systems include:
- Browns Ferry Unit 3 network storm that resulted in a manual reactor scram;
- Olympic Pipeline gasoline pipeline rupture with a broadcast storm affecting the SCADA system;
- Shutdown of an electric utility Energy Management System;
- Loss of SCADA communications at a water utility; and
- Conveyer belt shutdowns in an automotive assembly plant.
These events illustrated a common failure mechanism: disruption of digital communications produced significant physical or operational consequences without any malicious cyberattack.
Missing industry guidance
Standards such as ISA/IEC 62443 address the cybersecurity of industrial Ethernet switches including secure configuration, access control, firmware integrity and communications protection. However, ISA/IEC 62443 does not provide guidance for validating that network switches can safely support safety-critical control system communications under operational conditions. Similarly, NEI-08-09, NERC CIP, the NIST Cybersecurity Framework and CISA guidance do not provide engineering validation criteria for network switch integrity in safety-critical industrial applications. Specifically, I have not identified CISA guidance recommending operational validation testing of industrial Ethernet switches before deployment to ensure they will not introduce safety or operational problems.
They generally do not recommend operational testing of the switch in the process control environment beyond applying vendor-recommended mitigations. I have not found CISA guidance recommending testing whether a network switch can:
- become a common-mode failure;
- degrade controller performance;
- cause HMIs and alarms to freeze;
- cause operators to lose situational awareness;
- prevent automatic protective actions from continuing;
- limit recovery options; or
- correctly implement spanning tree, loop prevention and topology recovery under realistic operating conditions.
Collectively, these are engineering validation tests rather than cybersecurity assessments. They evaluate whether the communications infrastructure continues to support safe and reliable process operation under realistic failure conditions.
Engineering and regulatory implications
The Hatch event demonstrates that digital communications are no longer merely supporting infrastructure. In many facilities, they have become integral components of the process itself, directly influencing equipment behavior, operator actions and facility safety such as unexpected valve movements and the concurrent loss of operator displays and alarms.
Traditional functional safety analyses such as failure modes and effects analysis, hazard and operability, layer of protection analysis, and similar evaluations should explicitly treat failures of digital communications infrastructure as credible initiating events rather than assuming communications remain continuously available.
Yet industry and CISA guidance on network switch integrity is not available. Any industrial or manufacturing facility that depends on digital communications to monitor, control or protect physical processes faces similar engineering challenges. Protecting these facilities therefore requires more than defending against cyberattacks. It also requires treating digital communications with the same rigor traditionally applied to process sensors, actuators, controllers and other safety-critical equipment and, where appropriate, independent process sensor monitoring that obtains sensor data through communications paths that are independent of the affected control network.
As industrial control systems become increasingly networked, digital communications should be treated as engineering systems whose failure modes require the same validation, monitoring and engineering rigor traditionally applied to sensors, actuators, controllers and other safety-critical equipment—not simply cybersecurity protection.
In conclusion:
- Digital communications are no longer merely supporting infrastructure. In industrial and manufacturing facilities, they have become part of the control and safety functions themselves. Consequently, communication failures must be evaluated with the same engineering rigor as failures of sensors, actuators, controllers or other safety-critical equipment.
- Communications failures can produce physical and safety consequences without a cyberattack.
- Industry guidance provides little direction for validating that network switches can safely support safety-critical processes, highlighting the value of independent process monitoring that remains available when the primary communications infrastructure fails.
- I have not identified broadly adopted guidance recommending redundant network switch architectures specifically to mitigate single-switch failures in safety-critical industrial control systems.
Questions every facility should ask:
- Have replacement network switches been operationally validated under plant conditions?
- Have network failure modes been included in safety studies?
- Are operators able to independently verify critical process parameters if the primary control network fails?
- Do the communications architectures contain common-mode failure points?
- Is there a documented strategy for network redundancy and recovery?
About the Author
Joe Weiss
Cybersecurity Contributor
Joe Weiss P.E., CISM, is managing partner of Applied Control Solutions, LLC, in Cupertino, CA. Formerly of KEMA and EPRI, Joe is an international authority on cybersecurity. You can contact him at [email protected]

Leaders relevant to this article:
