Blog

Some Context on "Context"

Leor Fishman

A bottling line at a food and beverage plant goes down. An engineer on site hands the symptoms and the raw tag data to a frontier model. It tells them the issue is the capper. The capper's fault tags lined up precisely with the stoppage, and the model reasons that the capper sits upstream of the filler. Reasonable hypothesis, one problem. The capper is downstream of the filler. Its faults are a symptom, not a cause. The actual issue is the pasteurizer, which started failing 45 minutes earlier and which the model never checked, because nothing told it the pasteurizer was in the chain.

This is not a failure of observational data. All the data was there. It is not a failure of model intelligence either. We have validated that the problem persists across model generations, both empirically, through extensive simulation and evaluation, and conceptually: no model, however capable, can distinguish between causal structures from observational data alone. The missing ingredient is context. The model needs an accurate picture of what is running on the line and how each process feeds the next.

How, then, to build this context layer? A first pass would be to use plant diagrams and engineering documentation as the basis. While helpful, these sources go stale: processes and lines change, machines are added or removed. A second approach is brute force: sit every engineer, technician, and operator down for a few weeks and map every data point, and do it again each time the factory changes. Comprehensive and accurate in theory; in practice, almost no company has the time or resources. A third method is to extract context from asset models (UNS namespaces, PI Asset Framework, etc). This approach carries two issues. First, it is as costly as the interview approach, requiring manual effort and going stale as the factory changes. Second, even a perfect hierarchical model is focused on hierarchy, not causality. Nothing in a tag path tells you what tags gate which other tags.

These failure modes point at what an ideal causal extraction system looks like. It should draw on causal information already encoded in the plant, it should be guaranteed not to go stale, and it should be extractable mostly automatically, without hours of engineer time per line. Looked at this way there is a clear answer: extract context from the automation and control systems running those lines. Which data points matter, how are they derived from IO and sensor values, how do engineers see them day to day? The control logic and the HMIs answer all of these. Unlike documentation, control logic describes what the controller is actually executing. Unlike engineer interviews, it can be extracted in minutes, not weeks. Unlike asset models, it is focused on causality rather than asset location and hierarchy. For the baseline of a causal model we need something both automatable and authoritative, and only control systems provide this.

Concretely, control systems tell you how tags relate. If an operator watches a screen every day showing a tank level, a pump speed, a packager's output count, and whether a filter is running, that tells you those four points are what the operator cares about, together, as a unit. Tracing into the PLC then tells you the pump speed is set by a PID loop on that tank level and that the packager only runs if a vision system has approved fill levels. Derivation, control, and permissives all directly and automatically extracted.

This is a baseline, since controllers do not describe the whole plant. Every factory has dead code, unused screens, and orphaned tags. More importantly, physical coupling never appears in the logic at all: a shared header, a heat exchanger, or the forty-five minutes of product sitting on a conveyor between a pasteurizer and a filler. What controllers and HMIs do give you is a core that is authoritative wherever it speaks, even though it doesn't speak to everything. How agents can and should close that gap is its own article.

Set out like this, it seems obvious: agents should extract from the control systems they are talking to. Why isn't everyone doing this? The answer is translation and interoperability. One of the core pieces of Prophet is a deterministic translation and interoperation toolkit that takes in controls backups, deserializes them, and parses them into a generic graph model that forms the skeleton of our causal substrate. Dropping the controls logic into an LLM blindly is not enough, for two reasons.

First, the code usually isn't in a form an LLM can read. Vendor formats are proprietary binaries, naive exports are lossy and incomplete, and the abstractions don't map onto one another: an AOI is not a function block is not a UDT, even when they do the same job. Somebody has to make all of this legible before a model can touch it, and building the toolkits that turn these scattered sources into one model has been most of Axilon's engineering work since the start.

Second, deterministic workflows as tools for LLMs significantly outperform pure LLM interoperation with large bodies of data, and this holds even as models get smarter. Consider how GPT-6 Astra handles document modification: rather than writing extensive changes directly, it regularly writes a Python script to make them. Exact operations over large data are cheap and reliable in code and expensive and unreliable in attention, so the model spends its capacity deciding what to change and lets the script do the changing.

This is also the answer to an objection every technical reader has by now. Isn't a deterministic, vendor-specific contextualization toolkit exactly the hand-engineered domain knowledge the bitter lesson says gets steamrolled? No. Coding agents did not get good because models learned to emit machine code; they got good because they were handed compilers, type checkers, language servers, and test runners, and left to reason about what to do with the output. Nobody calls a coding agent running a linter a violation of the bitter lesson. That toolchain carries no opinion about how to write software. It makes the codebase legible and checkable, and every new model generation gets better at using it. Industrial AI hasn't had that moment because until now the toolchain didn't exist. The plant's structure is currently locked in proprietary binaries, stale drawings, and people's heads. A fine-tuned domain model is stranded by the next generation. The toolchain gets more useful every time the model improves. That is why we are building it at Axilon.

More from the blog