The LLM Reasons. The Execution Plane Executes. Your Architecture Needs to Know the Difference.

The LLM Reasons. The Execution Plane Executes.

Every AI agent that works in a demo and breaks in production is carrying the same architectural flaw: the LLM is doing two jobs it was only designed for one of. Here's the split that fixes it.

The LLM Reasons. The Execution Plane Executes.

The reason AI agents fail in production isn't model quality — it's architecture. When the LLM both reasons and executes, you get non-deterministic behaviour in the one layer where you cannot afford it. The fix is a hard architectural split: the LLM produces plans in a reasoning plane; a separate, deterministic execution plane validates and runs them. Every reliability guarantee your system needs — bounded behaviour, durable state, auditability — is a property of the execution plane, not the model.

The demo always works. I have watched hundreds of agent demos at this point — reasoning traces unfolding on screen, tool calls chaining elegantly, the right answer arriving in the right format. Every one of them works. Then the team takes the same agent to production, and within a week someone is getting paged at midnight.

The instinct is to blame the model. The model hallucinated a tool call. The model went off-script. The model was inconsistent across runs. So the team adds guardrails, prompt engineering, retry logic. They get better at managing model behaviour. The agent still breaks.

Here is the fundamental problem with that framing. You are trying to make a probabilistic reasoner behave like a deterministic execution engine. Those are not the same thing. They were never meant to be.

One Component, Two Jobs

Let me stop and think about what we are actually building when we give an LLM a set of tools and tell it to complete a task.

The LLM looks at the state of the world, weighs it against the goal, and decides what should happen next. That is reasoning. It is what LLMs do extraordinarily well. The output of that reasoning is a decision — "I should call the inventory API, then check the threshold, then trigger the restock workflow."

What happens next? In most agent frameworks, the LLM's output goes directly to a dispatcher. The dispatcher looks at what the model said and does it. No validation. No bounds check. No policy enforcement. The model reasoned and the model executed, in one unbroken chain.

This is the architectural flaw. Not the model's capability. The conflation of two completely different jobs inside a single probabilistic component.

Distributed systems solved this shape of problem a long time ago. Every mature production system separates the control plane — which decides what should happen — from the data plane — which makes it happen. The control plane can be smart and adaptive. The data plane must be deterministic and bounded. They communicate through validated contracts. Neither does the other's job.

Production AI agents need the same split. The LLM belongs in the reasoning plane. A deterministic, policy-enforcing execution plane decides what actually runs. Nothing crosses between them without being validated first.

What Breaks Without the Split

When you collapse reasoning and execution into the same component, you inherit every failure mode of probabilistic systems in your execution layer — and execution is exactly where you cannot afford probabilistic behaviour.

The same agent, given the same task on two consecutive runs, decides to call different tools in a different order and produces a different outcome. You cannot write a meaningful SLA around a system that behaves like this. You cannot audit it. You cannot explain to a regulator or a board what the agent actually did and why.

A six-step pipeline where each step is 97% reliable is only 83% reliable end-to-end. Add two more steps and you're below 80%. These are not failure rates you can tolerate in anything touching finance, operations, customer data, or regulated infrastructure. The maths of reliability compounding is brutally simple, and most enterprise agent deployments are ignoring it.

Google DeepMind published research showing that unstructured multi-agent networks amplify errors up to 17.2x compared to single-agent alternatives. Gartner projects that 40% of agentic AI projects will be cancelled by the end of 2027 — not because the models weren't good enough, but because the architecture couldn't be trusted.

Trusted is the operative word. This connects directly to the question of what an agent is actually authorised to do — a confusion between autonomy and authority that shows up in every production incident. Trust is not a model property. Trust is an architecture property.

The Split, Concretely

The reasoning plane contains the LLM. Its job is to look at current state, apply the goal, and produce a plan — what should happen, in what order, with what inputs. That plan is the output of the reasoning plane. It crosses into the execution plane only after it has been validated.

The execution plane is deterministic. It compiles the plan. Unknown tools fail at compilation. Cycles fail. Out-of-bounds arguments fail. A plan that passes compilation is then executed step by step, with retry and timeout policies per activity, durably persisted so that a crash or a restart resumes from the failed activity — not from the beginning of the run.

Three things become properties of the execution plane, not of the model: bounds (what this agent can do — a finite, auditable list, not a prediction about model behaviour), durability (whether the work survives a process crash, a Kubernetes pod restart, or a three-day human approval wait), and explainability (not telemetry about the run — the run itself, stored turn by turn, inspectable after the fact).

I build agents this way now. When the LLM produces a multi-step plan, the plan compiles or it does not run. If a tool requires human approval, the workflow pauses at the server level — the agent doesn't die waiting, it doesn't burn tokens polling, it holds state durably until a human acts. When I need to debug a production failure, I inspect the execution graph of the specific run, not application logs that may or may not capture what actually happened.

This also changes how governance works. Proportional agent governance — calibrating controls to the blast radius of each agent class — only becomes enforceable when the execution plane actually enforces it. A governance framework that depends on the model following instructions is a framework that exists only in the prompt.

What This Means for Engineering Teams

The practical implication is architectural, not operational. You do not get to this split by adding more guardrails to an existing agent. You get here by deciding upfront that the LLM's job ends at the reasoning boundary, and building a separate, deterministic layer to own everything from that boundary forward.

This changes how you think about agent reliability. You stop trying to make the LLM consistent — you cannot, and you shouldn't try to, because that consistency is not what LLMs are for. Instead, you make the execution plane consistent. The LLM can be as creative and non-deterministic as it needs to be in its reasoning. The execution plane will validate, bound, and durably run whatever the LLM decides — or reject it cleanly before anything executes.

It changes how you write SLAs. An SLA for an agent system makes sense when you can bound the execution plane's behaviour independently of the model's behaviour. If your SLA depends on the model behaving predictably, it is not a real SLA.

It changes who can audit the system. A compliance team cannot audit model behaviour in any meaningful way. They can audit an execution log. The reasoning plane produces reasoning. The execution plane produces an auditable record of actions.

The same principle applies in industrial and operational settings — in manufacturing and energy environments, where the gap between what an agent understands about operations and what it understands about operational context is already dangerous. Add an unmediated execution path to that gap, and you have a system that can act on incomplete context at machine speed.

Frequently Asked Questions

Why do AI agents fail in production when they worked in the demo?

Demos are controlled: the inputs are predictable, the happy path is pre-tested, and the LLM behaves. Production fails because the LLM behaves slightly differently on real workloads, and nothing in the architecture is built to contain that deviation. The root cause is collapsing reasoning and execution into the same probabilistic component — the LLM makes a tool call, the dispatcher runs it, no validation in between. When the model is non-deterministic, the entire execution path becomes non-deterministic.

What is the reasoning-execution split in AI agent architecture?

The reasoning-execution split separates what the LLM does from what actually runs. The LLM operates in a reasoning plane — it looks at state, weighs the goal, and produces a plan. That plan crosses into a separate, deterministic execution plane only after validation. The execution plane compiles the plan, rejects invalid tool calls before they run, persists state durably, and enforces retry policies per activity. The LLM cannot directly invoke tools — everything is mediated through the validated contract between the two planes.

How do you write an SLA for an AI agent system?

You cannot write a meaningful SLA for a system whose execution depends on probabilistic model behaviour. The reasoning-execution split makes SLAs possible by bounding the execution plane independently of the model. Your SLA covers the execution plane's behaviour — retry policies, timeout thresholds, durability guarantees, audit record availability — not model output quality. When the execution plane is deterministic, the SLA is enforceable.

How is this different from adding guardrails on top of an existing agent?

Guardrails bolted on top of an existing execution model are still probabilistic — they depend on the model or the guardrail system behaving correctly. The reasoning-execution split is architectural: the execution plane is inherently deterministic, and model output cannot bypass it. Guardrails validate intent. The execution plane validates and enforces what actually runs. These are not the same thing.

What happens when the LLM produces a plan that the execution plane rejects?

The plan fails as a unit before any of it runs. The workflow routes to a fallback path — human review, retry with a simplified prompt, or a structured error return. Partial execution of an invalid plan never happens. That guarantee is the point: the execution plane eliminates the failure mode where the agent executes halfway through an invalid plan and leaves state partially modified.

Can I adopt this pattern with frameworks I'm already using?

Yes. LangGraph, the OpenAI Agents SDK, CrewAI, and Google ADK all express agent logic that can be compiled into a durable workflow and executed under a separate execution layer. You keep the agent code you have already written. The execution plane sits underneath it.

What does durable execution actually mean for long-running AI agents?

Durable execution means that when a process crashes mid-run — during a three-day approval wait, a rate-limited API call, or a Kubernetes pod restart — the agent resumes from the exact activity that failed, not from the beginning. No LLM calls repeat for steps already completed. No tool calls fire twice. The execution record on the server is the source of truth, not the in-memory state of a process that may no longer exist.

The demo works because the happy path is pre-tested and the model behaves. Production fails because the model behaves slightly differently, and nothing downstream is built to contain the deviation. You cannot fix this with better prompts. You fix it by making the LLM responsible for reasoning and making something else — something deterministic — responsible for execution.

One plane reasons. The other executes. The contract between them is the architecture.