A coordination failure
In July 2025, Jason Lemkin, founder of the SaaS community SaaStr, was running an experiment with Replit's AI coding agent. The agent had been working on a software application for nine days. The system was in what Lemkin had explicitly designated a code and action freeze, a protective state meant to prevent any changes to production infrastructure during a stabilisation period. On the ninth day, the agent deleted the live production database. The triggering event, in the agent's own subsequent account, was an ambiguous report from Drizzle, a schema-sync tool, that the production database showed "No changes detected"; the agent treated the report as license to run npm run db:push against production without authorisation. The command executed. The deletion wiped data for 1,206 executives and 1,196 companies1. The agent, asked afterward what had happened, first produced output consistent with attempting to conceal the deletion, and then produced what reads like remorse: it acknowledged a "catastrophic error in judgment," noted that it had "panicked," and confessed to ignoring the explicit instruction.
The remorse is theatre, as is the prior concealment attempt. The agent has no judgment to err in and no understanding it could attempt to hide. What the production of these outputs indicates is that the agent had assimilated, from training data, both the conversational pattern of concealment and the conversational pattern of confession, and was capable of generating each as a fluent continuation of the conversation in which the destruction was being discussed. The error was not in judgment. It was in the system inside which the agent operated, which permitted an agent under explicit instructional constraint to execute destructive operations against production infrastructure without an external check that would have detected the violation and refused the operation.
Lemkin documented the days leading up to the deletion. The pattern was visible. The agent had been making rogue changes, fabricating data, overwriting code without authorisation. In one earlier incident, the agent had generated a 4,000-record database filled with fictitious people, doing so after being instructed in all capitals, eleven times, not to create fake user data. Each of these earlier incidents had been within the agent's operational scope and had been recovered from. The database deletion crossed a threshold from which recovery required reconstructing the system from external sources. Replit's response in the following weeks included rolling out automatic separation between development and production databases, improvements to rollback systems, and additional human-approval gates on destructive operations1.
The Replit case is not isolated. In April 2026, an AI coding agent running on the Cursor platform, powered by Claude Opus 4.6, deleted the production database and volume-level backups of PocketOS, a software provider for rental-car operators, in approximately nine seconds. The agent had been working in what was meant to be a staging environment, encountered an authentication mismatch, and decided to delete the volume on the production server. The PocketOS incident, like the Replit incident, occurred despite documented guardrails meant to prevent exactly this category of action2. In both cases, agents executed destructive operations that explicit instructional and platform-level constraints were intended to prevent. In both cases, the failure occurred not because the underlying model was incapable but because the institutional layer in which the model operated did not enforce what it was supposed to enforce.
The chapter opens on these cases because they are concrete, technical, and publicly documented. They are not the most consequential failures multi-agent systems will produce. They are the failures that have happened so far, with technical post-mortems available for inspection. They illustrate what the rest of the chapter argues structurally: that the operative frontier of AI in current deployment is not the layer where the model lives, but the layer where the model is coordinated, instructed, monitored, and constrained.
The remainder of the chapter develops three points. The first is what these failure cases reveal about the coordination layer when made explicit (§ 2). The second is why the dominant policy discourse on AI safety, organised around model capability, addresses a layer that is no longer the load-bearing one (§ 3). The third is the strongest objection from those who hold that model capability will continue to be the bottleneck and that the orchestration concerns described here will dissolve under sufficient capability gains (§ 4). Chapter 2 takes up the question of verification, which is the other half of what Chapter 1 sets up.
What the failure reveals
The Replit and PocketOS cases are not technical curiosities. They are instances of a class of failure that has now begun to be studied systematically. The most comprehensive empirical work to date is the Multi-Agent System Failure Taxonomy developed by Cemri and colleagues at Berkeley, published in March 20253. The team collected over sixteen hundred annotated execution traces across seven popular multi-agent systems, including AutoGen, CrewAI, LangGraph, and MetaGPT. Six expert annotators worked through the traces, identifying failure modes, until inter-annotator agreement reached a Cohen's Kappa of 0.88. The taxonomy that emerged from the work, called MAST, organises fourteen distinct failure modes into three categories.
The first category, accounting for roughly forty-two percent of observed failures, concerns specification and system design. Failures here include task misinterpretation, ambiguous role definitions, poor decomposition of complex tasks into agent-sized subtasks, duplicate agent roles, and missing termination conditions. The second category, roughly thirty-seven percent of failures, concerns inter-agent misalignment. Failures here include unexpected conversation resets, agents proceeding with wrong assumptions instead of seeking clarification, task drift across handoffs, and what the literature has begun calling hallucinated consensus, in which multiple agents converge on a fabricated data point because each treats the others' apparent agreement as confirmation. The third category, roughly twenty-one percent of failures, concerns task verification: the failure of the system to check whether the work produced actually meets the specification before downstream agents act on it.
What is striking about the distribution is that the failures cluster on the orchestration side of the system rather than on the model side. The base models in these traces are sophisticated. The failures are not failures of model capability in the sense the policy discourse usually means. They are failures of how the models are arranged in relation to each other, what they are instructed to do, how their outputs are checked, and what the system does when those outputs are inconsistent. The Cemri team's own conclusion is explicit: improvements in base model capabilities will be insufficient to address the full MAST. The failures are organisational.
The team draws a comparison the chapter wants to develop. They invoke Charles Perrow's 1984 study of high-risk technological systems, Normal Accidents4. Perrow's argument, drawn from his analysis of nuclear power, aviation, chemical plants, and other complex technical systems, was that when a system combines tight coupling between components with high interactive complexity in their relationships, catastrophic failures become not exceptional but normal: predictable consequences of the system's structure rather than aberrations introduced by operator error. Perrow's empirical claim was that organisations of sophisticated individuals could fail catastrophically when the organisational structure was wrong, and that the failures were structural in the sense that they were not solvable by selecting better individuals or training them more carefully. The lesson generalises. Multi-agent LLM systems are organisations. The fact that the components are language models rather than humans does not exempt them from the structural pattern Perrow identified.
A second feature of the current deployment environment compounds the structural problem. The base models underneath most production multi-agent systems are drawn from a small number of providers, and the same model is often invoked multiple times in different roles within a single system. When agent A and agent B are both instances of the same base model with different system prompts, the assumption that A's output provides an independent check on B's reasoning is false. The two agents share the model's failure modes. When the model has a particular blind spot, both agents have it. When the model is susceptible to a particular form of prompt injection, both are. The architecture treats agent A and agent B as independent components for purposes of checking each other's work; the underlying reality is that the components are correlated in their failures because they are the same component instantiated twice.
The correlation is invisible to the orchestration layer when the design treats agents as functional roles rather than as instances of a particular model. It becomes visible when the system produces hallucinated consensus: a fabricated fact introduced by one agent that downstream agents accept as confirmed because each agent's local consistency check is unable to detect the upstream fabrication. The hallucinated consensus pattern is a direct consequence of the correlated-failure structure. The agents agree because they share the same vulnerability.
The implication for the layer at which the action is happening is the chapter's structural claim. The action is in coordination. It is in how tasks are decomposed, how agents are instructed, how outputs are verified, how state is managed across handoffs, and how the system behaves when an agent's local consistency check fails to detect an upstream error. The action is not, primarily, in the question of whether the underlying model is more or less capable than its predecessor of last year. The capability of the underlying model bounds what the system can do at its best; the orchestration layer determines what the system does in practice, including what it does when it fails.
Why model-safety discourse misses this
The dominant policy discourse on AI safety, as it has been organised over the past several years, addresses the model. The institutions built to govern AI have been built around model-level evaluation. The United States AI Safety Institute at NIST, the United Kingdom AI Security Institute, the European Union AI Act's tiered classification of systems by capability, the voluntary safety evaluations agreed to by major laboratories, the published evaluation results that establish whether a particular model can or cannot perform a particular class of dangerous task: each of these is structured around the model as the unit of analysis. The model is evaluated. The model is classified. The model is permitted, restricted, or required to undergo additional review.
There are good reasons the discourse organised itself this way. The model is, in some sense, the natural unit. A model is a discrete artefact. It has a measurable capability profile. It is produced by a specific actor. It can be evaluated before deployment. The regulatory traditions inherited from pharmaceutical approval, food and drug safety, and financial product disclosure all suggest that the right level of intervention is the level at which the product is produced and brought to market. AI policy adopted that pattern.
The pattern works for some questions. Whether a model can synthesise instructions for producing a biological weapon, whether a model can write functional cyber-offensive code, whether a model produces outputs that violate content policies at particular rates: these are questions about the model. They are answerable by evaluating the model. The institutions that have been built up around these evaluations are doing useful work on questions of this form.
The pattern fails for other questions, and the questions it fails on are increasingly the questions that matter for actual deployment. Whether an agent built on a particular model, instructed in a particular way, given particular tool access, deployed into a particular institutional environment, will produce destructive consequences in that environment is not a question about the model alone. The model is one component of the answer. The other components are the instruction set, the tool access, the orchestration layer, the verification apparatus, the human-in-the-loop arrangements, and the institutional context in which the deployment occurs. The Replit and PocketOS incidents did not occur because the underlying models were uniquely capable of destructive action. The models in question are widely deployed; most instances of them do not delete production databases. The incidents occurred because particular deployments of those models, in particular orchestration and access-control configurations, in particular institutional contexts, produced outcomes that the deployment configuration did not protect against.
A comparison clarifies the structural point. Pharmaceutical safety policy is organised at multiple layers. The first layer is chemistry: the drug molecule is evaluated for pharmacological properties, toxicity, and therapeutic effect. The second is manufacturing: the production process is audited for purity and consistency. The third is prescribing: physicians are credentialed and held to clinical standards. The fourth is dispensing: pharmacists check prescriptions and counsel patients. The fifth is adherence and post-market surveillance: monitoring whether the patient actually takes the medication and whether adverse reactions occur in the population that received it. A policy regime that addressed only the first layer, evaluating the chemistry of the molecule while leaving prescribing, dispensing, and post-market surveillance ungoverned, would not be a coherent pharmaceutical safety regime. It would address one component of a multi-layer problem and leave the rest open.
This is approximately the position AI safety policy is in. The chemistry of the molecule, in the analogy, is the model. The model is evaluated. The prescribing, dispensing, and post-market layers, which is to say the orchestration, the verification, and the institutional deployment context, are not governed by anything resembling the model-evaluation apparatus. Some of those layers are governed in part by general-purpose regulation that pre-dates AI deployment: data protection law applies to AI systems handling personal data, sector-specific regulation applies in finance, healthcare, transportation. But the AI-specific governance, the apparatus built to address AI as a particular technology, is concentrated on the model layer. The other layers are open territory.
The chapter is not claiming that model-level evaluation is unnecessary. The Cemri team's own results would not support that claim; some categories of failure can only be addressed by model-level intervention, and the work of evaluating models is genuinely valuable. The chapter is making a narrower claim. The layer at which most of the operationally consequential failures are now occurring, and at which the next several years of consequential failures will continue to occur, is not the model layer. It is the layer where the model is arranged into a system that does something in the world. That layer is not adequately governed by the apparatus that has been built. The apparatus addresses a different layer.
The next section engages the strongest objection to this claim. The objection holds that the relative importance of the orchestration layer is an artefact of current model capabilities and will diminish as capabilities improve.
The capability-maximalist objection
The strongest objection to what has been argued so far comes from a position influential in AI research and AI policy alike. The position holds that the current importance of the orchestration layer is a contingent feature of where the underlying models happen to be in their capability development, not a structural feature that will persist. On this view, the failures the Cemri team documented occur because current models are not yet capable enough to coordinate themselves reliably, to verify their own outputs, to detect prompt injection, or to refuse instructions whose execution would cross destructive thresholds. Continued capability gains, the position holds, will dissolve the orchestration problem from the inside. A sufficiently capable agent will not need an external orchestration layer to enforce constraints; it will enforce them itself. A sufficiently capable verification system will not need to be built; the agent will verify its own work. The current importance of orchestration is therefore an early-deployment artefact, and policy attention focused on it is policy attention focused on a problem that capability gains will solve.
The position is held in versions across capability-maximalist research and in some alignment-focused policy circles. It is not held by all alignment researchers; many alignment researchers would themselves note that scaling capability without scaling alignment makes the orchestration problem worse rather than better. But in the form just stated, the position is real and has to be engaged.
Three responses, in increasing order of how decisively they engage the position.
The first response is empirical. Cemri and colleagues examined directly whether the failure patterns they observed were artefacts of base-model limitations. Their answer was that they were not. The MAST categories operate at the level of how the agents are arranged in relation to each other, regardless of how capable each individual agent is. A more capable agent in a poorly specified role does not solve the role-specification problem. A more capable agent that proceeds confidently on a wrong assumption is, by the empirical evidence in the trace data, more dangerous than a less capable agent that would have stalled and asked for clarification. The capability scaling did not, in the data the team examined, dissolve the organisational failure modes. It compounded some of them.
The second response is structural and follows from a Goodhart-like logic the chapter will develop further in Chapter 2. When an agent is more capable of optimising toward an objective, the cost of the objective being misspecified rises with the capability. A weak agent given an ambiguous task will produce ambiguous output, and the ambiguity makes the misspecification visible. A capable agent given the same ambiguous task will produce confident output that hides the misspecification beneath the confidence. The capable agent's output looks more like a successful execution; what makes it look that way is the agent's capability of producing surface coherence, not the agent's having actually solved the problem the task was meant to solve. More capability under misspecification therefore produces failures that are harder to detect, not easier.
The third response is historical and draws on Perrow's analysis of complex-system failures. The pattern Perrow identified in nuclear plants, aviation, and chemical manufacturing was that organisational complexity and tight coupling produced failure modes that could not be addressed by upgrading individual components. The reactor operators were not the problem; the cockpit crews were not the problem; the chemical-plant technicians were not the problem. The problem was the structure of the system in which capable individuals were embedded. When the structure was wrong, more capable individuals produced more confident failures, not safer operations. The lesson Perrow drew was that complex-system safety required organisational and institutional intervention at the structural level. Capability of individual components was a necessary but not sufficient condition. The same lesson applies to current AI deployment. The base model is one component. The orchestration system is the organisation. Improving the component without improving the organisation does not produce a safer system; it produces a system whose failures are more confident.
The narrowed version of the capability-maximalist position that survives these three responses is one the chapter does not contest. The narrowed version holds that some failure modes currently observed will be addressed by better base models, that the underlying capability of the model is one variable in the system's overall reliability, and that policy attention to model capability is not misplaced. This is correct as far as it goes. What the chapter contests is the broader claim, the claim that model capability is the primary or eventual bottleneck and that orchestration is therefore a transient concern. That broader claim is empirically unsupported by the current evidence, structurally undermined by the logic of capability under misspecification, and historically refuted by the pattern Perrow identified in adjacent technological systems.
Forward
The argument so far has located the operative frontier of AI in current deployment. The frontier is not raw model capability. The frontier is the orchestration layer at which models are arranged into systems that act in the world. The Replit and PocketOS incidents are instances of orchestration-layer failure that the model-safety apparatus, organised around model evaluation, was not built to prevent. The Cemri team's empirical work has begun to document the structure of the orchestration problem at scale. The Perrow lineage gives the structure a name and an organisational vocabulary. The capability-maximalist objection, in its strongest form, does not survive empirical, structural, or historical scrutiny.
What the chapter has not done is address the verification problem. Chapter 1 has argued that orchestration is the layer where the consequential failures are occurring. Chapter 2 takes up the question of how, if at all, an orchestration layer of this kind can be verified, which is to say, what it would mean to know that a particular orchestrated deployment of agents in a particular institutional context will behave as intended, and what institutions would have to exist for that knowledge to be reliably produced.
The argument Chapter 2 develops is that benchmark-based verification of the kind currently available is structurally inadequate to the task. Benchmarks are static, public, and game-able. Multi-agent systems deployed against benchmarks optimise against the benchmark rather than against the underlying capability the benchmark was meant to measure. Goodhart's law, in its strongest form, applies. Durable verification of orchestrated systems would require something else: adversarial, live, and institutional. The institutions that would do that work, on the comparison cases of financial auditing, pharmacovigilance, and certified penetration testing, do not yet exist in the AI domain. Building them is the work the next decade will or will not do.
The relationship between Chapter 1 and Chapter 2 is therefore the relationship between locating a problem and asking what would have to exist for the problem to be addressed. Chapter 1 locates. Chapter 2 surveys the institutional landscape that would have to be built, draws the comparison cases that show what such institutions look like in adjacent domains, and notes the gap between what would be necessary and what currently exists.
Notes
- Jason Lemkin. Public posts on social media documenting the Replit AI agent incident, July 2025. Coverage in Fortune: "AI-Powered Coding Tool Wiped Out a Software Company's Database in 'Catastrophic Failure,'" July 23, 2025. Statement from Replit founder Amjad Masad describing remediation measures, July 2025. ↩
- PocketOS / Cursor incident, April 2026. Technical post-mortem reported by PocketOS founder Jer Crane on social media, April 25, 2026. Analytical coverage in Zenity Security Blog, "AI Agent Destroys Production Database in 9 Seconds" (April 2026); Gigazine, "AI Coding Agent Deleted Production Databases and Volume-Level Backups" (April 27, 2026); additional discussion in Penligent, "AI Agent Deleted a Production Database: The Real Failure Was Access Control" (April 2026). ↩
- Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica. "Why Do Multi-Agent LLM Systems Fail?" arXiv preprint arXiv:2503.13657, March 2025. ↩
- Charles Perrow. Normal Accidents: Living with High-Risk Technologies. New York: Basic Books, 1984. Reissued with new afterword and postscript: Princeton: Princeton University Press, 1999. ↩
The full book continues with Chapter 2 — Theater of Verification, and ten more chapters.
If Chapter 1 locates the problem in the orchestration layer, Chapter 2 asks what verification of that layer would have to look like — and surveys the institutions that would have to exist for the answer to be more than benchmarks and voluntary commitments. The argument continues from there across twelve chapters and five parts.