Blog

The "Three Cs" of Industrial AI

Ben Simon with Leor Fishman

There are three steps ("three Cs") needed to unlock AI for industrial operations: Connectivity, Context, and Cognition. My core thesis in this piece—and our core thesis at Axilon with our platform, Axilon Prophet—is that context, not connectivity or cognition, is the key bottleneck.

Connectivity: a necessary but insufficient foundation

In the industrial context, connectivity refers to the data pathways that enable the extraction and transmission of telemetry data from physical equipment and the systems that control that equipment (e.g. PLCs).

There are various layers of the connectivity stack—the device layer, the data transport layer, and the aggregation/analysis layer. The point of this article is not to unpack each of these layers. Industrial data connectivity is already a well-understood concept and one that most major industrial enterprises have implemented in practice, to varying degrees.

The point is this: connectivity is necessary but insufficient for industrial AI. With connectivity, what you get is raw data from the plant floor. But that's all you get. Data collection itself is a means to the practical end of monitoring and analyzing key trends, detecting failures before they happen, and ultimately improving plant-floor efficiency, uptime, throughput, and quality.

Until now, it has been a struggle to bridge this gap. For most organizations, the post-connectivity phase never got past analytics like equipment efficiency dashboards and operator interfaces with alarms. Those with deeper pockets could hire teams of data scientists to build custom machine learning models, but these efforts rarely scaled. Advanced analytics are resource-intensive to build out and maintain as the plant floor changes.

Cognition: frontier AI is powerful and reliable, but also insufficient

In theory, LLMs should solve all of this. Frontier reasoning models like Anthropic's Claude and OpenAI's GPT are extraordinarily powerful. In minutes, these models can write code to analyze a dataset, find correlations that would take an analyst days, build a just-in-time dashboard on request, and explain their reasoning in plain language. They are also increasingly reliable: given real data to work with, these models reason over the numbers carefully rather than making them up.

So again, there should be nothing stopping LLMs from being able to proactively predict plant failures, root-cause and triage issues in real time, and identify bottlenecks for even the most complex industrial processes.

But in practice, what happens when you take Claude and hook it up to your industrial data? Does it just work?

No. The models can do accurate and advanced statistical work, but they cannot actually make sense of the raw industrial data.

Take for example a factory line with all of its inherent complexities. This sample line has a few hundred tags streaming from a mix of different PLCs: motor currents, valve positions, temperatures, pressures, flow rates, counts, and a long tail of status bits, most without descriptions or units. If you hook up the raw tag data and dump it into Claude, what you get back is statistical analysis: this tag is trending down, these two are correlated, that one just crossed a threshold, etc.

What you don't get is diagnosis, causation, or a recommended fix. The model has no reliable way to know which tags sit upstream of which, which are causes versus downstream symptoms, or which of the thousand correlations it can compute actually matter for throughput. Thus, when output drops, a dozen signals move at once. Without knowing how the line is actually controlled, the LLM will not know which tags are the cause of the drop, which tags reflect the consequences, and which are just noise. The critical contextual information cannot be gleaned from the raw data or the tag names. It lives elsewhere.

Context: the true bottleneck

The argument here is simple: if you were to give an LLM perfect context—if it knew everything about the industrial environment it was dropped into—it would produce near-perfect intelligence.

Perfect context doesn't just mean having access to tag relationships, equipment manuals, and plant documentation, though that would be a good start. It also includes all of the information that lives in the heads of operators, all of the past edge cases that were dealt with and process decisions that were made. If you can find a way to give all of that to an LLM, it will "just work."

If an LLM doesn't have all of that context, or even some of it, it will infer from limited context. This is where hallucination starts to creep in and where intelligence becomes educated guesswork. Smarter models do not solve this problem. A model with imperfect context will perhaps make better guesses about the missing information, but it will still be guessing.

Thus far, we have established the importance of the context layer. But for something to be a bottleneck, it must be both critical and difficult. This, perhaps, leads to my second, more controversial claim: even in the age of frontier AI, building and maintaining the industrial context layer remains a challenge.

Building the industrial context layer: LLMs on their own?

If you had infinite resources and time, you could achieve perfect context. You could pay every single operator at your plant to sit down, walk through their whole day, map all of the tags to equipment tag by tag, explain the quirks and subtleties of how their equipment works, and on and on. You'd have to do this not just once, but every time one of the lines, cells, or even individual pieces of equipment changes.

Obviously, this highly-manual approach is not remotely feasible or scalable.

Is AI the silver bullet solution, not just for the intelligence layer, but also for the context layer as well?

Yes and no. LLMs are an invaluable asset for parsing unstructured information like plant documentation, equipment manuals, and maintenance records. Dealing with this data before the advent of LLMs was a nightmare and doomed the vast majority of industrial contextualization initiatives.

But extracting context from unstructured information is only part of the puzzle. To build a robust and accurate context layer, you also need to tap into the industrial "systems of record," the living control and automation systems in the industrial environment that are actually executing logic and facilitating the manufacturing process. Supplementary information—the documentation referenced above—is always stale compared to the living, dynamic control systems that run the plant at the minute to minute and second to second levels.

These industrial systems of record—PLCs, SCADA, DCS, MES, ERP—are, first of all, undocumented and lack standard APIs for semantic interaction. This means that even extraction of relevant configuration data from these systems is a challenge.

But even more importantly, these systems generate such large volumes of data with such bespoke structure that even LLMs with large enough context windows are left to naively explore these control systems and prioritize what they infer to be relevant relationships in the data. Counterintuitive as it may sound, deterministic code/configuration analysis for these industrial systems has proven—in all of our internal evaluation at Axilon—to be a necessary precondition for reliable agentic operations. In sum, we have found that you need a preprocessing and static analysis layer, as well as the direct LLM processing of unstructured documentation and data.

In plain English: you cannot simply dump all of your control system code, configuration, and data into an LLM and expect it to provide reliably correct and useful answers. LLMs are an extremely important tool in building out the context layer, but they are not enough on their own.

It is beyond the scope of this article to provide a detailed explanation of exactly how Axilon approaches the problem of building scalable, comprehensive industrial data context. For that, you should reach out here over DM or visit our website. Suffice it to say, Axilon has adopted a hybrid deterministic and non-deterministic (i.e. LLM-based) method for building out the context layer efficiently, reliably, and scalably.

Conclusion: the opportunity at hand

The holy grail of industrial AI is finally within reach. To get there, we don't need smarter models. We don't need more data. We need more accurate, more deterministic, and more scalable contextualization.

More from the blog