TL;DR: 76% of manufacturing AI projects fail — and the industry keeps misdiagnosing why. The real cause is not bad data or weak change management. It is that we are deploying AI agents that understand operations into environments that require operational context, and treating those two things as equivalent. An agent can read sensor telemetry, model failure sequences, and execute through an available API. What it cannot do is know that a valve is locked out because maintenance started three minutes ago, or that rolling back a patch today violates a compliance window closing at midnight, or that the anomaly in Pump 4 is expected because the process engineer changed the batch this morning. That gap is not a model limitation. It is an architectural failure — and unlike a crashed database, it produces its consequences six weeks later, on a production line, without an error log.
A technician loses access to a remote engineering system after a routine software update. She asks an AI agent to help restore connectivity. The agent identifies the recent update as the likely culprit and recommends rolling it back. The technician does. The connection returns.
The AI agent cannot see what the patch was for. It does not know the update contained a critical security fix for a known vulnerability in the plant's OT network. It cannot know that 'restore connectivity' and 'maintain operational integrity' are not the same instruction in an industrial environment — even when they look identical from the outside.
No attacker was involved. No model hallucinated. The agent performed exactly as designed.
The narrative in manufacturing AI right now is seductive. Siemens reports a 20% throughput increase from adaptive manufacturing agents at its Erlangen factory. Honeywell and Borouge are building what they call the industry's first AI-driven control room in Abu Dhabi. Rockwell Automation embedded agents across its Plex platform in August 2026. The ROI case is real — 171% average return in enterprises that have successfully deployed.
I am a believer in industrial AI. The productivity gains are documented, not projected.
But I think the industry has developed a blind spot so large it is hiding in plain sight. The Cisco 2026 State of Industrial AI report found that 76% of manufacturing AI projects fail. Eighty-four percent of those failures trace back to leadership decisions, not technical breakdowns. The conventional interpretation: executives don't champion it properly, change management is weak, data isn't clean enough.
I think that is the wrong diagnosis. I think 76% of manufacturing AI projects fail because we are shipping agents that understand operations into environments that require operational context — and treating those two things as the same.
Understanding operations means an agent can read sensor telemetry, identify anomalous patterns, model probable failure sequences, generate a recommended action, and execute it through an available API. It can do this faster than any human, across more variables, with more consistent pattern recognition. That capability is real and valuable.
Understanding operational context means knowing that a particular valve is in a locked-out state because maintenance started three minutes ago and the crew is still on the line. It means knowing that rolling back a patch today violates a compliance window that closes at midnight. It means knowing that this anomaly in Pump 4 telemetry is normal — because the plant is currently running a non-standard production batch that the process engineer scribbled on a whiteboard this morning.
An LLM trained on decades of industrial documentation does not know what happened this morning. And unlike enterprise software environments — where a wrong action produces a recoverable error in a staging table — an OT environment produces physical consequences. A bad recommendation that gets executed at machine speed inside a SCADA environment is not a bad query. It can be a safety incident, a compliance breach, or an equipment failure that takes a production line down for six weeks.
The agent cannot see the difference. And we are not building architectures that show it.
Here is what makes this particularly hard to see: the failure mode is usually silent.
When an AI agent takes a catastrophically wrong action in an IT environment — deletes a database, corrupts a file, sends a misconfigured payload — the incident is immediate, visible, and attributable. The agent did the thing. The thing failed. There is an error log.
In OT environments, the failure mode is deferred and diffuse. The agent recommends a setting change that is technically within parameters. The change gets applied. Three shifts later, a component starts degrading slightly faster than expected. Six weeks later, there is an unplanned maintenance event. The root cause investigation finds vibration anomalies. Nobody traces it back to the configuration change the agent suggested in Week 1, because nobody built a system that tracks the provenance of agent recommendations into the physical process chain.
This is what the 43% IT/OT collaboration gap actually means in practice. It is not a political problem between departments. It is an architectural problem: the agent was deployed into a context it cannot represent, and the infrastructure to make that context machine-readable was never built.
Three things — and none of them are about the model.
First: operational context as a first-class data primitive
In every plant I have worked in, operational context lives in three places: the shift handover log (usually a physical binder or an unstructured Teams message), the permit-to-work system (often a separate CMMS that is not connected to anything the AI agent can query), and the process engineer's head. Before you deploy an agent that can act on process data, you need a structured, queryable representation of what is happening in the physical environment right now — not historically, not on average. This is harder to build than the agent itself, and it is almost never scoped into the project.
Second: an explicit trust membrane between recommendation and execution
In enterprise software, this looks like a staging environment and a deployment gate. In OT, this looks like a mandatory OT-domain context check before any action touches a physical process. Not a generic confirmation step — a structured query against the operational state that verifies the recommended action is valid in the current physical configuration. If the permit-to-work system shows an active lock-out for the asset in question, the agent does not execute. Full stop. The check is architectural, not procedural. The same principle that keeps AI agents from inheriting execution authority they were never meant to have applies here with physical stakes attached. See: blog.siftit.dev/posts/agent-autonomy-not-agent-authority
Third: provenance chains, not audit logs
An audit log tells you what happened. A provenance chain tells you why the agent thought it was the right action, what context it had access to at decision time, and what context was missing. When a maintenance event occurs six weeks after an agent recommendation, the provenance chain is what makes that connection findable. Without it, you are flying blind on a feedback loop that makes the system progressively safer. And without progressive safety feedback, you cannot govern your agent fleet — you can only lock it down, which is not governance. See: blog.siftit.dev/posts/agent-governance-not-agent-lockdown
An AI agent is not dangerous because it is too capable. It is dangerous because it is capable without context. In a factory, that gap is measured in physical consequences.
The job is not to make the agent smarter. The job is to make the environment legible to the agent before the agent acts on it.
Frequently asked questions
Why do 76% of manufacturing AI projects fail? The conventional diagnosis blames data quality and change management. The real cause is architectural: we are deploying agents that understand operations into environments that require operational context, and treating those two things as the same. An agent can read sensor telemetry and identify anomalies faster than any human. But it cannot know that the anomaly is expected because the plant is running a non-standard batch, or that the asset it is about to touch is under an active lock-out. That gap — between operational data and operational state — is what produces the 76% failure rate.
What is the difference between operational data and operational context in manufacturing AI? Operational data is what the sensors, historians, and SCADA systems record: telemetry, process variables, anomaly patterns. Operational context is what those readings mean right now — a valve in locked-out state because maintenance started three minutes ago, a patch rollback that would violate a compliance window closing at midnight, a pump anomaly that is normal because the process engineer changed the batch this morning. Operational data is machine-readable. Operational context usually lives in a shift handover binder, a permit-to-work system not connected to anything the agent can query, and the process engineer's head. That is the gap.
Why are AI agent failures in OT environments harder to detect than in IT environments? In IT, a wrong agent action produces an immediate, visible, attributable error. In OT, the failure mode is deferred and diffuse. The agent recommends a setting change technically within parameters. Three shifts later, a component starts degrading slightly faster than expected. Six weeks later, there is an unplanned maintenance event. Nobody traces it back to the agent recommendation in Week 1 because nobody built a provenance chain connecting agent decisions to physical process outcomes. The failure is silent by design — and that silence is the most dangerous part.
What is operational context and why can't an LLM trained on industrial documentation provide it? Operational context is the real-time state of the physical environment: what maintenance is active, what permits are open, what non-standard conditions the process is running under right now. An LLM trained on decades of industrial documentation knows what a permit-to-work system is, what lock-out procedures require, and what a compliance window means. It does not know what happened this morning in your plant. That is not a model limitation — it is a data architecture problem. The context was never made machine-readable in the first place.
How should operational context be structured for AI agents in manufacturing? Operational context must be a first-class data primitive, not a sidebar. That means a structured, queryable representation of what is happening in the physical environment right now — not historically, not on average. In practice, this covers three sources: the shift handover log (usually an unstructured Teams message or physical binder), the permit-to-work system (often a separate CMMS disconnected from anything the agent can query), and the process engineer's knowledge of non-standard conditions. Building this is harder than building the agent itself, and it is almost never scoped into the project.
What is a trust membrane between agent recommendation and agent execution in OT? A trust membrane is an architectural check — not a procedural one — that validates a recommended action against the current physical state before any execution happens. In OT, this looks like a mandatory query against the operational state: is there an active permit-to-work for this asset? Is this action within the current compliance window? Does the current batch configuration make this anomaly expected? If the check fails, the agent does not execute. Full stop. The check lives in the architecture, not in a prompt instruction the agent can reason past.
What is the difference between an audit log and a provenance chain for AI agent decisions? An audit log records what happened. A provenance chain records why the agent thought the action was correct, what context it had access to at decision time, and what context was missing. In OT, this distinction is what makes it possible to connect a maintenance event six weeks later to an agent recommendation in Week 1. Without a provenance chain, you have no feedback loop. With one, the system can get progressively safer over time. Most production deployments build audit logs and call them provenance. They are not the same thing.



