Digital transformation, which, to me, means integration of the Purdue Enterprise Reference Architecture (PERA) from Level 1 through Level 4, arguably started with the internet and the increased connectivity it provides. The addition of intelligent fieldbuses and wireless sensor network devices increased the amount of available information from operational technology (OT) networks by orders of magnitude. However, conversations about digital transformation and its realized benefits haven’t progressed in proportion to the amount of data available. Perhaps the problem is not the data, but the vast amount of it and the ability to identify and find what’s important about it—securely.
I recently saw a graph showing enterprise data represents 60% of all data, and the volume of data created is growing at a compounded rate of 27% per year. What this means is organizations are collecting vast amounts of data in various data repositories (data warehouses, lakes, and lake houses) and not necessarily using it. However, because everyone else is, and they know data is valuable, they feel they they’d better capture it.
But capturing “everything” along with the Industrial Internet of Things (IIoT) compounds the problem. Datasets have become too large for any one physical repository to hold, so there is a shift toward virtualization and these decentralized architectures:
- Data fabric: A concept that uses metadata, machine learning and APIs to weave together existing, disparate data sources (lakes, warehouses, databases) into a single virtual access layer. Queries are made against the "fabric," and it fetches the data from wherever it lives behind the scenes.
- Data mesh: A decentralized, organizational approach, instead of a central IT team that manages a massive data lake, individual business domains (e.g., "logistics" or "sales"). Organizations using this method treat data as a product, hosting and managing their own localized repositories while exposing standard APIs for others to consume.
Many people are once again looking at artificial intelligence (AI), including large language models (LLMs), to save the day with friendly interfaces that make it possible to find the jewels amidst the dross. However, this means that you must pose the right question, and if you don’t know what treasures are hidden in the data, the right question is a challenge.
Get your subscription to Control's tri-weekly newsletter.
While attending the IEC TC65 plenary session, there was a presentation by Japan’s Open Data Spaces that included a graphic indicating that by 2030, LLM data consumption will equal the volume of human-curated text and proposed a model of how AI can help make sense of all the data, while respecting data ownership and security through the use of data spaces and an intermediate Architectural Quantal middleware to integrate the open public data with your qualified (internal and trusted third party) proprietary data. It is an interesting concept by this government-sponsored organization to make their concepts available to all.
However, another fly in the ointment, for the operational technology (OT) realm at least, are the many legacy systems that include a data historian and DMZ with a long-term historian so the data captured is limited to time-stamped PV and status information. I suspect traditional legacy systems with minimal data/diagnostic gathering capabilities are now the minority since the control systems for roughly 25 years have supported COTS concepts and digital communications with field sensor networks. Despite this, and based on my experience with IEC SC65E WG10, I believe at least 90% of systems do not capture this available data, and the reason is (not) knowing what data to capture, then how to make use of that same data to positively impact safe reliable operations to provide a positive return on investment.
There is no argument that we are living in a data-driven society, and sometimes it is scary how that data is being used. However, digital transformation requires data in context and unfortunately, despite drowning in data, we are suffering from a dearth of knowledge. Despite all the work in this space since introduction of the Reference Architectural Model for Industry (RAMI) model in 2015, the digital transformation of the OT realm at least will continue to flounder largely because without knowing which data can lead to information and how to integrate the same with the larger community knowledge there are more urgent fires requiring our limited time and resources.