What control engineers should understand about Ethernet APL reliability

A tale of fieldbus troubleshooting highlights why network architecture, operator confidence and fault isolation are just as important as the technology itself

Key Highlights

  • Point-to-point spurs can prevent a single instrument failure from disrupting an entire segment, but field switches, power supplies and uplinks become new common points of failure that require careful design.
  • Engineers should evaluate switch redundancy, power strategies, network segmentation and maintenance practices to minimize the operational impact of inevitable equipment failures.
  • Even a rare field switch failure that takes multiple instruments offline can undermine confidence in Ethernet-APL. Preparing operations teams with training, diagnostics and clear troubleshooting procedures is essential for long-term adoption.

It was Monday morning, and Emile couldn’t get through the control room for the morning ops meeting before Constantine intercepted him. “I think we must have a card failure,” said Constantine, an experienced operator whose curiosity (and worrisome nature) drove him to learn more about the distribute control system (DCS) interconnection with field devices. Emile immediately saw that the devices, having spurious or constant connection problems, all shared the same fieldbus segment. One of the devices was a control valve—installed and configured to provide pressure control for a distillation tower that ran close to a total vacuum, essential to keeping the product on spec.

Emile acknowledged it could be the card. But the card also provided an H1 fieldbus interface for another segment. He explained how another population of about 10 devices would also lose their connection to the DCS. “No worries,” he said, “everything will come back in the same state they were when the card was swapped.” But Constantine was an operator who always feared the worst and so agreed other troubleshooting avenues should first be pursued.

Emile already thought it was a troubled instrument causing enough communication problems to disrupt the remaining devices, like a meeting attendee on a Zoom call connecting from the bar on a cross-country train. Eventually, Emile isolated the problem device—a DP transmitter that functioned for 26 years atop the same distillation tower. Once it was disconnected, all communication was restored, and remained error-free after the problem device was replaced with a contemporary version.

Would such issues be minimized once he deployed advanced physical layer (APL), a two-wire Ethernet for instruments in hazardous locations? Each APL device would have a point-to-point spur back to the field switch, so many device- or spur-level faults should remain local to that device. The field switch, its power and its uplink would still be shared infrastructure, but each instrument would have its own single-pair connection back to the field switch. Should an instrument have communication issues, e.g., due to corrosion or moisture ingress, it should only affect that device. So, there’s an improvement in robustness thanks to the inherent spur-to-spur isolation.

Get your subscription to Control's tri-weekly newsletter.

But field switches introduced another worry. Compare a typical fieldbus coupler: along with its copper network wiring, the calculated mean time between failures (MTBF) was measured in centuries. The random failures he was seeing in this 25-year-old installation were still neither routine nor catastrophic—the plant kept running, the problems were annoyances.

He imagined how he’d deploy APL switches, which introduced active components, firmware and (most likely) local power supplies to the field. How would he minimize the impact of a random switch failure—for example, limit the loops that might be impacted when a switch reset or had a power supply issue? The more switches he deployed, the greater the likelihood of such random failures. He remembered fieldbus early adopters, some of whom kept critical loops hardwired over 4-20 mA. Should he follow suit for APL? Sadly, it was the critical loops for which the health and diagnostic information was most interesting.

The other vexing problem was how he’d prepare operations for switch hiccups. Constantine was able to diagnose a common-mode failure after a little study. How would the crew piloting the process nights, weekends, and holidays discern what was happening when six or 10 instruments went offline, and control valves went to their failure positions?

This more troubling failure mode might not be technical at all. Operators have little patience for mystery failures that disrupt their shift, especially when they cause or contribute to a shutdown. If an APL field switch takes several instruments down at once, it will not be remembered as an unfortunate random failure of an otherwise sound architecture. It will be remembered as the day the new networked instrument scheme created an unwelcome challenge for the crew. The fix will be obvious to everyone in the room: hardwire the affected devices. And whether fairly or not, the engineer who recommended the architecture will own that story for a long time.

Fieldbus was not always as robust as it is today. Emile remembered a time when a short circuit could fail an entire segment, for example. He continued to hope that similar fortification was on the roadmap for APL field switches.

About the Author

John Rezabek

Contributing Editor

John Rezabek is a contributing editor to Control

Sign up for our eNewsletters
Get the latest news and updates