AI Agents Don't Fail Because the Model Got It Wrong

A systematic analysis of 1,600+ agent execution traces identified 14 failure modes. Most didn't come from model errors. They came from the infrastructure layer between probabilistic output and deterministic business logic — the layer most organisations haven't built.

The enterprise AI problem isn't model quality. It's the infrastructure layer that converts probabilistic output into business-grade execution — and most organisations haven't built it.

Who owns the Determinism Stack in your organisation?

If that question doesn't have an immediate answer, you have a problem. Not a future problem. A current one that is accumulating cost and risk at exactly the rate your AI deployments are scaling.

Here's what I mean.

A systematic analysis of over 1,600 agent execution traces across seven popular frameworks, published this month, identified 14 distinct failure modes. Most of those failures did not originate in model reasoning errors. They originated in system design, coordination gaps, and verification failures — the infrastructure around the model, not the model itself. The researchers put it this way:

A payment cannot be 'probably executed once.' An authorization boundary cannot be 'usually respected.' A ledger cannot be 'mostly consistent.'

That's the finding the industry keeps misreading.

The prevailing response to AI agent failures is to upgrade the model. Try a frontier model. Fine-tune it. Add RAG. Improve the prompt. This logic assumes that the agent's probabilistic output is the problem. The evidence suggests otherwise: the problem is what happens to that output after the model finishes reasoning.

Businesses run on determinism. A transaction either completes or it doesn't. An authorisation either holds or it doesn't. An audit trail is either complete or it isn't. LLMs produce probabilistic output by design — confident approximations that are correct most of the time. The gap between "most of the time" and "every time" is not a model quality gap. It is an architectural gap. Something has to translate probabilistic reasoning into deterministic execution, and most enterprise agent systems have not been designed to do that translation reliably.

This is where agents are actually failing.

The cost you're not measuring

A research synthesis published on arXiv this month, drawing on three randomised field experiments across 4,867 developers, quantifies the downstream effect of this gap.

AI coding agents increased completed tasks at the assisted-coding stage by 26%. At the project level, the gains were 50%. Impressive numbers — until you examine what happens to the output. Only 44% of code produced by agents across 6,000 real development sessions survived into actual user commits. The other 56% was generated, reviewed, and discarded.

The paper names the downstream cost: the Verification Tax. Every piece of probabilistic agent output that enters a downstream system — CI pipeline, code review, security scanning, rework cycles — generates an assurance workload. That workload is not free. It compounds with deployment scale. Gartner's modelling, cited in the same paper, suggests AI coding costs could exceed the average developer salary by 2028 under growing token consumption alone — before accounting for the verification infrastructure required to make that output trustworthy.

The enterprise ROI model for AI coding agents is being constructed on the throughput number while the real cost accumulates in the verification queue. This is not a novel error. It is the same quality trap that manufacturing companies fell into before lean disciplines forced them to measure rework costs rather than production volume. The Verification Tax is the software equivalent of hidden rework cost. It appears downstream, it compounds, and it is invisible to the metric management is watching.

I've seen this pattern before — in hydro plants, in manufacturing lines, in market research operations. The number on the dashboard looks fine. The cost is building somewhere nobody is looking.

The missing layer

Put these two findings together and a specific architectural gap becomes visible.

There is a layer that every serious AI agent deployment requires: the infrastructure that converts probabilistic output into auditable, reliable, business-grade execution. It has an architectural dimension — the deliberate design of the interface between what the model produces and what the downstream system expects. It has an economic dimension — the cost of verification before that output can be trusted at scale. And it has a human dimension — the engineering role responsible for designing and maintaining that interface under production conditions.

I'm calling this the Determinism Stack. Not because the label matters, but because naming something is the prerequisite to building it. You cannot staff, fund, or audit an unnamed thing.

Most enterprises don't have a Determinism Stack. They have models and they have business processes, and somewhere between the two there is a gap filled with manual checks, duct-tape integrations, and the tacit assumption that someone is watching.

The counterargument worth taking seriously

The obvious objection: as models improve and agentic frameworks mature, the verification overhead will shrink. Better automated testing, AI-assisted code review, self-correcting agents — the Verification Tax is a transitional cost, not a permanent architectural requirement.

This may be partially right. Automated verification will improve. But the failure taxonomy makes a sharper point: the 14 failure modes map not to model capability gaps but to architectural interface problems. A more capable model does not, by itself, create a well-defined contract at the boundary between probabilistic output and deterministic execution. It produces higher-quality output that still needs to cross that boundary reliably.

The Determinism Stack is not a workaround for weak models. It is the engineering discipline that makes agent outputs trustworthy regardless of model capability. These are different problems with different solutions, and conflating them is why enterprises are investing in the wrong layer.

Three questions for your next AI architecture review

Where does probabilistic output meet deterministic business logic in your agent systems? Map every point where an agent's output — a decision, a generated artefact, a recommended action — feeds into a downstream system that requires a definite answer. These are your unengineered interfaces. Each one is a latent failure waiting for the right scale to trigger it.

Are you measuring the Verification Tax? If your AI ROI model has no line item for the assurance workload downstream of AI output, your numbers are understated. The tax exists whether or not you're tracking it. The 56% discard rate on agent-produced code is a concrete upper bound on what that cost can look like in practice.

Name the owner before your next deployment decision. This is an architectural responsibility — not a prompt engineering responsibility, not a model selection responsibility, and not something a Centre of Excellence can govern from a distance. If nobody owns it explicitly, it doesn't get built. And if it doesn't get built, the scale you're planning for will make its absence visible in the worst possible way.

Twenty years of building systems in manufacturing, energy, and market research has taught me one consistent lesson: the technical problem that brings a deployment down is rarely the one the vendor presentation warned you about. It's the interface — the point where two systems with different assumptions about reliability have to hand something off.

The AI industry has spent three years making models better at generating output. The next three years will reveal whether the enterprises deploying those models built the infrastructure that makes the output trustworthy.

Most haven't started yet.