New Business Rules in the AI Agentic Age Get the free whitepaper

Agents on the Plant Floor

OT governance is not IT governance with different nouns. Rollback does not exist, the safety case is a legal artefact, and latency is a control parameter.

Agents on the Plant Floor

Most writing about AI agents assumes a world where mistakes are expensive but reversible. Refund the wrong customer and you claw it back. Post the wrong message and you delete it. That assumption is doing a lot of quiet work, and on a plant floor it is simply false.

Operational technology is the part of the business where software touches physical processes: valves, drives, furnaces, conveyors, grids, pumps. Agents are arriving there, for good reasons. Diagnosing a fault across historian data, maintenance records and vendor manuals is exactly the sort of synthesis work they are good at, and the people who used to do it are retiring faster than they are being replaced.

But the governance model that works in IT does not transfer, and the reasons are structural rather than cultural.

Four things that are different

There is no undo

IT governance leans hard on reversibility. You can restore a database, revoke a token, roll back a deploy. The implicit safety net under “move fast” is that most mistakes are recoverable.

Open the wrong valve and the product is contaminated. Change a setpoint on a furnace and you have a batch that has to be scrapped and possibly a lining that has to be replaced. Trip a line and restart takes hours, not seconds, because physical processes have inertia and ramp profiles.

This changes what a “hold for human approval” is worth. In IT it is a nice safeguard. In OT it is frequently the only safeguard, because detection after the fact buys you nothing.

The safety case is a legal artefact, and it is already written

Industrial plants operate under a documented safety case: hazard analyses, safety integrity levels, interlocks, defined operating envelopes. It is reviewed by regulators and, in serious sectors, by insurers. It is also specific: it states what may control what, under which conditions, with which independent protection layers.

Nothing in that document contemplates a language model choosing a setpoint. Which means an agent that writes to a control system is either outside the safety case or requires it to be reopened — a process measured in months and involving people whose job is to say no.

This is the single biggest practical difference. In IT, governance is something you add. In OT, governance already exists, has legal force, and your agent has to fit inside it.

Latency is a control parameter, not a user-experience metric

In IT, a slow check is an annoyance. In a control loop, timing is part of correctness. A decision that arrives late is not a slow decision, it is a wrong one, and a governance layer that adds unbounded latency into a loop with a cycle time is not a safety feature. It is a new failure mode.

The practical consequence is a hard line. Agents belong in advisory, diagnostic and supervisory positions, above the control loop, where a few hundred milliseconds is irrelevant. They do not belong inside it. Anything that must respond in a deterministic window stays with the PLC or the safety instrumented system, which are designed for exactly that and have been for decades.

Availability outranks confidentiality

IT security ranks confidentiality first. OT inverts it: availability and integrity come first, because the failure mode is not a data breach but a stopped line or an unsafe state.

This inverts the default for governance too. In IT, fail-closed is the safe choice: if the policy engine is unreachable, deny. On a plant floor, blocking a legitimate operator action during an upset can be the dangerous outcome. The correct behaviour is usually to fail out of the path — the agent loses its advisory capability, the humans and the existing control system carry on exactly as they did before, and the loss is degraded assistance rather than a halted process.

That has to be designed deliberately. It is the opposite of what an IT-shaped default would do.

Where agents genuinely help

None of this is an argument for keeping agents out. It is an argument for putting them where the value is real and the blast radius is bounded.

Diagnosis. Correlating a vibration signature with maintenance history, a vendor bulletin and three similar events from other sites is hours of specialist work and minutes of agent work. The output is a hypothesis for a human, which is exactly the right shape.

Procedure retrieval. Finding the right revision of the right procedure, for this equipment, at 3am, is a genuine problem in plants with decades of documentation. Read-only, high value, low risk.

Shift handover and reporting. Summarising twelve hours of alarms and interventions into something the next shift can act on. Nobody enjoys writing it and everybody relies on it.

Change preparation. Drafting the change request, assembling the impact assessment, pulling the relevant interlocks into one place. The agent prepares, the human decides, the existing process approves.

Notice what these share: the agent produces information for a person. The moment it produces an action on equipment, you are in a different regime and the safety case is the governing document, not your policy engine.

The Purdue model still decides the conversation

Most plants are organised, formally or informally, on something like the Purdue reference model: levels from the field devices and sensors at the bottom, through the control layer, up to supervisory systems, then site operations, then the corporate network.

The useful thing about that model, for this discussion, is that it already encodes how far something is allowed to reach. Nobody has to be convinced that a corporate-network service should not write directly to a PLC; the architecture has said so for thirty years.

So the first question about an agent is not “what may it do” but “what level does it live at, and what is it allowed to traverse”. An agent that reads from the historian at the supervisory level and writes nothing downward is a proposal most OT engineers can evaluate in an afternoon. An agent that reaches across levels is a proposal that needs the people who own the safety case in the room.

Framing it this way tends to unstick the conversation, because it uses a vocabulary the plant already has instead of importing an IT one that treats every system as equally reachable.

The alarm-flood problem, and why it is the honest test

Here is the scenario that decides whether an OT agent deployment is real.

Something upsets the process. Within ninety seconds the operator has several hundred alarms, most of them consequential rather than causal. This is a known, well-documented failure mode, and it is exactly where a summarising agent looks most valuable: rank the alarms, identify the probable first cause, surface the procedure.

It is also where every assumption is under maximum stress. The agent is reasoning about a state that is, by definition, outside normal operation. Its training and its retrieved context describe the plant as it usually behaves. The operator is time-pressured and will be inclined to trust a confident summary precisely when scepticism matters most.

Two design consequences follow, and they are not optional.

The agent must be visibly degraded-capable: when it is uncertain, or when its data is stale because the historian is lagging under load, it has to say so in a way an operator reads at a glance. A confident summary built on ninety-second-old data during a transient is worse than no summary.

And it must never be in the path of the response. If the governance layer, the network or the model is unavailable during the upset, the operator's tools must behave exactly as they did before the agent existed. This is the fail-out-of-the-path principle, and the alarm flood is the case that tests whether it was actually implemented or merely intended.

What governance looks like here

Read and write are different worlds. Treat them as separate grants with separate approval paths, not two permissions on one connection. An agent with historian read access and no write path is a straightforward proposition. Add one write tool and it becomes a safety-case conversation.

Bound it by operating state, not just by role. The same request is fine during normal operation and unacceptable during a startup, a trip or a maintenance window. Governance that only knows who is asking, and not what state the plant is in, is missing the variable that matters most.

Make the human approval real. A hold that routes to someone who is not in the control room, or arrives on a channel nobody watches during an upset, is theatre. The approval path has to match how the plant actually runs at 3am, which is usually the control room and usually not email.

Keep the record in OT terms. “Agent recommended X, operator approved, setpoint changed from A to B at T” is an entry an incident investigator can use. A generic API log is not, and NIS2 reporting has a habit of arriving with questions shaped like the first one.

Assume the network is not what the diagram says. Air gaps are aspirational in most plants. The vendor laptop, the remote support tunnel and the historian replication to the corporate network are all real, and an agent sitting in the IT layer can often reach further into OT than anyone intended. That reach is worth enumerating before it is worth governing.

The honest summary

Agents on the plant floor are not a smaller version of agents in the back office. The reversibility assumption is gone, the regulatory frame predates the technology by decades and has legal force, timing is a correctness property, and the safe default is inverted.

Which is why the advisory band is the right place to start, and why the boundary between advising and acting deserves to be an enforced, recorded line rather than a convention. That line is the whole design. Everything else is detail.

Book a demo