Your Multi-Agent System Is Not Failing Because the Model Is Wrong
UC Berkeley's MAST taxonomy, drawing on 1,642 real production traces, found that 79% of multi-agent failures stem from poor specification and coordination, not model capability. The industry's response has been to build runtime guardrails. That addresses 21% of the problem. The other 79% is a design-time failure that runtime enforcement cannot reach. Enterprise teams that keep swapping models and adding guardrails will keep failing. The fix is treating agent specifications as engineering artifacts with their own lifecycle, not as configuration files written once and forgotten.
The narrative in enterprise AI right now is about control. Runtime control. Identity binding. Policy enforcement. Intent verification at the gateway. Every major infrastructure player has announced something in this space in 2026: Broadcom AgentMinder at VMware Explore, Microsoft MXC at Build, the OWASP AI Control Standard, NIST agent standards. MCP has 97 million monthly SDK downloads. The industry is converging fast on a runtime enforcement stack, and it looks like progress.
I think most of it is looking in the wrong direction.
Not because runtime control is unimportant. It matters. But here is the fundamental problem with that framing: runtime enforcement catches failures after the system misbehaves. It cannot catch failures that are baked into how the agents were designed. And according to the most rigorous production data we have, that is where most failures live.
The MAST taxonomy, produced by researchers at UC Berkeley and presented at NeurIPS 2025, analyzed 1,642 annotated execution traces across seven major multi-agent frameworks in production. ChatDev, MetaGPT, Magentic-One, and others. Real systems, real failures. The breakdown:
- 44.2% of failures: system design issues. Step repetition, unawareness of termination, failure to follow task specifications.
- 34.4% of failures: inter-agent misalignment. Reasoning-action mismatches, task derailment at handoff points.
- The rest: task verification failures.
Combined, the first two categories account for 79% of all observed failures. Not hallucinations. Not model capability gaps. Not jailbreaks that a runtime gateway would catch. Specification problems. Coordination design problems. Failures that were written into the system before it ever ran.

The industry is building high-performance fences around systems that were never correctly defined to begin with.
Here is the distinction worth being precise about. Runtime governance and specification engineering are solving two completely different problems.
Runtime governance answers: "Is this agent behaving within allowed boundaries right now?" It is a containment mechanism. Necessary. Not sufficient.
Specification engineering answers: "Does each agent know precisely what it is supposed to do, when to stop, what valid output looks like, and what to do when the task falls outside its scope?" That is a design-time question. A guardrail cannot answer it. A better model cannot answer it either.
An agent that receives a vague task definition will produce vague, inconsistent behavior, and the orchestrator will not know whether to proceed or retry. No runtime policy catches that. The cascade starts in the specification.
This connects directly to the reasoning-execution split: when an agent's job is poorly defined, it cannot reason about its task clearly, and execution becomes unreliable. Specification failure is the upstream version of that same architectural problem. It happens before execution begins.
I have seen this pattern in practice. A manufacturing client built a multi-agent quality control pipeline last year. Five agents. Well-resourced team. Modern tooling. The demo passed. Production failed inside three weeks. The agents were stepping on each other's outputs. One agent was re-checking work already completed by another because nobody had defined what "done" meant at the boundary. A sixth agent was added to adjudicate conflicts. That agent also failed, because the conflict-detection criteria were never specified.
The team's response was to upgrade the underlying model. That is the instinct this industry has trained people into. When a multi-agent system fails, try a bigger model. It is the wrong move in 79% of cases.
What actually fixed it: going back to each agent's task specification and treating it like engineering code. Defining the input schema. Defining the output schema. Defining the explicit termination condition. Defining the failure mode: not retry behavior, but what the agent should do when it cannot complete the task. Surface it, escalate it, or stop cleanly.
Three things I now treat as non-negotiable before any agent in a multi-agent system goes to production:

1. The task boundary must be defined before the agent is written. Not inferred from the agent's prompt. Explicitly stated. What is the agent's scope? What is categorically outside it? An agent that doesn't know where its job ends will improvise at the edge. That improvisation is where 44.2% of your failures originate.
2. The handoff contract must be a typed schema, not an assumption. When Agent A passes output to Agent B, what does valid output look like? What does Agent B do when the input is malformed or incomplete? Most multi-agent implementations treat handoffs as trusted pipes. They are not. They are the exact point where misalignment compounds. OrchestraBench, released this year, found that cascade radius grows near-linearly with pipeline depth: from 0.9 at depth 3 to 4.7 at depth 7. Every stage you add without a validated handoff contract multiplies your blast radius.
3. Termination logic is architecture, not an afterthought. The most common failure in step-repetition (the largest single failure subcategory in MAST) is an agent that does not know it has completed its task. It keeps checking, re-doing, or waiting for a signal that never comes. This is not a model problem. A model given a clear, unambiguous termination criterion terminates. The work is writing that criterion, which means the engineering team has to know what "done" looks like before anyone writes a prompt.
The MCP stack, AgentMinder, MXC: these are real infrastructure and they solve real problems in the 21% of cases where runtime enforcement is the right answer. Authorization boundaries, identity binding, audit trails. Build all of it.
But if you are deploying a multi-agent system where the agents' task specifications were written in an afternoon and never tested against adversarial or edge-case inputs, you do not have a governance problem yet. You have a specification problem. And the governance layer is going to show you that at the worst possible moment: in production, with a customer watching.
Frequently Asked Questions
Why do most multi-agent system failures happen at the specification level rather than the model level?
Because models generally do what they are told when they are told clearly. The problem is that most teams don't write clear specifications. They write prompts that describe the ideal case and leave the edge cases implicit. When a production task falls in a gap the specification didn't anticipate, the agent improvises. In a single-agent system, that improvisation is visible. In a multi-agent pipeline, it compounds silently across every downstream stage.
What is the difference between agent specification engineering and prompt engineering?
Prompt engineering is about how you phrase instructions to get better outputs from a single inference. Specification engineering is about defining what an agent's job is, what valid completion looks like, where the job ends, and what happens when the agent cannot complete it. Specification is an architectural artifact. It should be versioned, tested against adversarial inputs, and reviewed the same way you review a contract or a schema. Most teams treat it as configuration. That is the gap.
How does cascade radius relate to multi-agent pipeline depth?
OrchestraBench's measurements are stark: mean cascade radius (the number of downstream stages corrupted by a single failure) grows near-linearly from 0.9 at pipeline depth 3 to 4.7 at depth 7. The three semantic failure modes (context pollution, conflicting outputs, and premature action) corrupt every downstream stage. Retry does not repair them. Detection and attribution do. This means every agent you add to a pipeline without a tested specification and handoff contract is not a 1x risk addition; it is a multiplier on every upstream failure.
Also worth noting: agent memory architecture is a related failure axis. Memory problems and specification problems compound each other. An agent with bad memory of prior steps will fail to apply even a well-written specification correctly in longer sessions.
What should a production-ready agent specification include?
At minimum: a typed input schema, a typed output schema, an explicit termination condition (what does "done" look like?), a scope boundary (what is categorically out of scope and should be escalated?), and defined behavior for when input is malformed or incomplete. If the team cannot write these artifacts before building the agent, the agent is not ready to be built. The specification is the engineering work. The code is the implementation of it.
Can better models solve specification problems?
No. The MAST data is unambiguous here: model capability is not a top failure category in multi-agent systems. A model given a vague specification produces vague, inconsistent behavior. A model given a well-engineered specification produces consistent behavior at scale. Swapping GPT-4 for GPT-5 in a system with under-specified agent roles does not fix the specification. It just produces a more confident version of the same failure.
How should teams test agent specifications before production deployment?
Treat them like interfaces. Write test cases for the boundary conditions: what happens when the input is at the edge of scope? What happens when the input is malformed? What happens when the agent completes the task and the downstream agent hasn't confirmed receipt? If you only test the happy path, you are testing a demo. Production doesn't have happy paths.
What is the right order of operations: specification first or guardrails first?
Specification first, always. Runtime guardrails catch deviation from intended behavior. But if the intended behavior was never precisely defined, the guardrail doesn't know what it's enforcing. The sequence is: define the specification, validate the specification against adversarial inputs, then add runtime enforcement around a system you understand. Reversing this order gives you expensive infrastructure around an undefined system.
The runtime enforcement stack the industry is building is necessary. But it is not sufficient, and for most enterprise teams right now, it is being built first when it should come last.
Specification engineering is not glamorous. There is no major conference announcement for it, no 97-million-download SDK, no hyperscaler product launch. It is design work done in documents before any code is written. That is exactly why it keeps getting skipped.
Runtime guardrails cannot contain a system that was never correctly designed. Specification engineering is not the final layer. It is the foundation.



