Blog

The Alarm Is Dead, Long Live the Alarm!

Ben Simon

Talk to enough controls engineers and plant managers and you'll hear some version of the same story.

At 2:40 in the morning, a tank level alarm pages the on-call tech. He drives 35 minutes to the plant, walks the floor, and finds nothing. The transmitter has been drifting since its last calibration, which anybody on day shift could have told him. He clears the alarm and drives home.

In truth, the idea behind industrial alarms is sound. The logic is essentially: if the process crosses a condition that matters in some way, then tell someone who can do something about it. Nobody should sit staring at trend screens waiting for trouble. If you were designing plant supervisory operations from scratch today, you'd still put event-based alarming at the center.

So what's broken? Why does everyone have their own version of the same middle-of-the-night-for-nothing alarm story?

First of all, every alarm is a threshold that a person picked. Most of the time, this alarm configuration happens at plant (or SCADA) commissioning and is never revisited. This means that staleness, incompleteness, and inaccuracy are all real problems that plague alarm management.

But the problem is deeper still. An individual alarm threshold carries no context. The same level reading or setpoint means different things during startup, turnaround, and steady state, but the alarm fires identically in all three. Plants accumulate tens of thousands of these alarms. Operators are continually drowning in them.

For decades the workaround has been human judgment. The alarm is “dumb,” so the person with the right knowledge and experience filters it. This solution works well enough when a plant has a few hundred alarms and a single operator can hold the whole unit in his head. But modern industrial environments are far too instrumented for that setup to be sustainable. This is how we inevitably get the scenario where there are countless alarms, some that are correct, some that are faulty, and some that are sometimes correct and sometimes faulty.

AI is the obvious solution. Frontier models are surely capable of both alarm configuration and rationalization; they already handle much more complex tasks without issue.

But deploying AI to fix alarm management isn’t so straightforward. The first-pass solution—point an agent at the control systems, stream in data from the historian, ask the agent to flag what matters—won’t work. Even a mid-sized plant produces more data in an hour than any LLM context window holds, and models degrade as the window fills. In practice, this naive approach ends up costing a lot in compute just for a model to search for needles in a massive haystack, miss all of them, and report that the plant is operating fine.

How, then, to deploy AI to fix alarm management? The architectural secret lies in the legacy alarm structure itself: Don't watch everything; instead, define the events that matter up front and act when an event fires. The key is that, in this case, “act” doesn’t mean immediately notify an operator; it means trigger an agent to investigate and analyze.

In this (correct) approach, an event fires, and an agent wakes up and investigates only the relevant slice of the plant: the loop that alarmed, its neighbors, the open work orders, the last time this pattern showed up. Using this information, the agent confirms or dismisses the underlying issue. Real problem? The tech gets a page with a diagnosis instead of a tag name and a timestamp. Drifting transmitter? Nobody's phone rings.

Legacy alarm structure with event-based triggers is sound; what’s needed to bring alarming into the 21st century is an agent layer that, when triggered, investigates, reasons, and ultimately decides whether to wake the sleeping operator.

In practice, designing and building out this event-based agentic system isn’t the only challenge. The other (arguably larger) task is arming AI agents with the comprehensive, ground-truth context they need to actually reason over raw factory data. With Prophet, we are building both of these layers: The context foundation and the agent harness, complete with event-based triggers.

At Axilon, we envision a future where humans work side by side with agents on the factory floor, with agents handling the first-pass investigation in every scenario, freeing up humans to focus on higher-leverage work—and ensuring that no on-call tech gets woken up for nothing ever again.

More from the blog