Automating a Broken Process Makes It Fail Faster
Automating broken processes exposes latent failures at machine speed and scale.

Automation does not break processes. It takes a process that is already broken and runs it at a speed no one can intervene on in time. Most enterprise AI deployments fail because organizations automate processes that were never designed to work reliably in the first place, with recurring patterns like undefined decision points, inconsistent inputs, and untracked exceptions driving the failure.
Why automation accelerates failure instead of preventing it
Bill Gates made the underlying point back in 1995, in The Road Ahead, long before anyone was talking about large language models or agentic workflows: automation applied to an efficient operation magnifies the efficiency, and automation applied to an inefficient one magnifies the inefficiency in exactly the same way. Michael Hammer said much the same thing in 1993, more bluntly: automate a mess and what you get is an automated mess. Neither observation depended on the technology of the moment. Both describe a property of automation itself: it takes whatever logic already governs a process and runs it faster, at greater volume, with less opportunity for anyone to notice something has gone wrong.
The reason this matters now, and didn't quite land the same way three decades ago, comes down to what automation actually removes from a workflow. A broken process staffed by humans still fails, but it fails slowly. Someone notices a mismatched invoice and holds it for review. A customer service rep senses that a scripted answer doesn't fit the situation and escalates instead of reciting it. Those pauses, judgment calls, and informal checks absorb errors before they spread. Automation strips exactly those buffers out. It executes the broken logic at machine speed and machine scale, with none of the friction that used to catch the problem early.
This is not a fringe read on the data. MIT, McKinsey, RAND, Gartner, BCG, S&P Global, and Deloitte converged on the same diagnosis between 2024 and 2026: the model is almost never the failure point, the process is. The process was. Different institutions, different datasets, same conclusion, which is about as close to consensus as this kind of research gets.
How widespread process unreadiness is, and what failure looks like in practice
Enterprise AI deployments are failing at a striking rate, and the technology is rarely the reason. Organizations are automating processes that were never built to run reliably in the first place, and no model, however capable, corrects for that on its own.
RAND Corporation's analysis of enterprise AI initiatives found that only a minority of failures traced back to technical shortcomings. The bulk of them came from poor strategy, unclear governance, and organizational design that never accounted for processes already broken before any automation touched them. McKinsey's research points the same direction from the other side: the organizations getting real value from AI are disproportionately the ones that redesigned their workflows end to end first, then integrated automation into the redesigned version.
Three patterns occur repeatedly in the post-mortems. Decision points that get resolved with "it depends" instead of a defined rule. Inputs that arrive in inconsistent formats depending on who or what generated them. And exceptions nobody bothered to track because they were assumed to be rare. That last assumption tends to be the expensive one. When more than roughly one in five real-world cases actually requires a judgment call, the "standard" path being automated no longer covers the bulk of the volume, and the cases that fall outside it generate manual cleanup that quietly erases whatever efficiency the automation was supposed to deliver.
None of this appears at launch. It becomes visible in the data two or three weeks in, once edge cases pile up and the data stops matching what the workflow was built to expect. By then, the team has usually already built manual workarounds around the automation rather than through it, and what was meant to be one process has quietly become two, run in parallel, both needing upkeep.
Agentic systems make the arithmetic worse. As companies chain AI agents into longer autonomous workflows, the outcome depends on the reliability of every step in the chain. A workflow can fail even when each individual step performs well on its own, simply because the chain has enough links for one weak point to sink the rest. The governance side of this is already producing casualties. The Gravitee State of AI Agent Security 2026 report found that most technical teams have moved past planning into active testing or full production, but only a small fraction of those agents went live with complete security and IT approval. Agent sprawl of that kind is projected to cause SLA breaches and cascading failures wherever automations cross department lines.
What the case record shows: Target Canada, Air Canada, McDonald's, and Meta
The historical record is a repeating pattern, not a scattering of edge cases, in which the technology did what it was told and the process it was given was the problem.
Target Canada is the clearest example on record. The company ran an automated SAP supply chain to manage inventory, distribution, and supplier payments, and the technology executed flawlessly on product data that was only roughly 30% accurate. Empty shelves, distribution centers overflowing with unsold stock, and eventual closure followed: billions of dollars lost and over a hundred stores shut down. The automation didn't create the bad data. It scaled it with total precision, which is arguably worse.
Tesla's 2018 experience runs a parallel course in manufacturing. Elon Musk stated publicly that excessive automation at Tesla had been a mistake, and that humans were underrated, after an over-automated assembly line produced quality problems severe enough to require pulling people back into the line. Software can be patched overnight. A physical assembly line jammed on bad automation logic has to stop, and stopping a production line costs real money for every hour it sits idle.
Air Canada's chatbot offers the legal version of the same story. The customer-service bot gave a passenger incorrect guidance on bereavement fares, and the airline was later found liable for negligent misrepresentation. The bot had been layered onto an underlying process that was never clearly defined or consistently executed, and the automation converted that gap directly into legal liability rather than into any efficiency gain.
McDonald's voice-ordering pilot with IBM shows what the same failure looks like at high volume and in public view. The pilot ended in June 2024 after documented ordering errors and a wave of viral complaints, though McDonald's official statement framed the decision as a desire to explore voice-ordering options more broadly rather than an admission of failure. Deploying a customer-facing agent onto an inconsistently defined ordering process doesn't average out over volume. It compounds, and it compounds in front of customers holding phones.
Meta's 2026 incident points to the governance dimension specifically. An internal AI agent error briefly exposed sensitive internal data, a failure that traces back to access controls that were never formally defined before the agent went live. Poorly governed agent systems turn out to be fragile in exactly the places nobody thought to lock down in advance.
Across all five cases, the common thread is the absence of process clarity, data accuracy, or governance architecture at the moment the automation was switched on.
What process readiness requires before automation begins
Process readiness is a structural interrogation of whether the work actually flows coherently in the first place, before any automation logic gets layered on top of it.
Three layers of workflow clarity need to be established first. The structural layer means mapping what actually happens, not the idealized flowchart version: every handoff, every tool dependency, every cycle time, and the cost of delay at each misaligned step. The behavioral layer means locating where human judgment, improvisation, or quiet workaround behavior already occurs: every one of those workarounds is a hidden process failure, and automation will proliferate it rather than eliminate it. It will proliferate it. The outcome layer means confirming that whatever metric the process produces actually connects to measurable business value; automating a process without first redefining what "good" looks like just produces inefficiency faster and at a bigger scale.
The bar for process consistency is agreement: one defined path, consistently executed, so the automation isn't left choosing between two parallel versions of the same task. Left to make that choice on its own, it will pick one and break on the other, every time. Undefined decision points, any step still resolved with "it depends", are the single most common failure point in this whole story. Automation facing one either forces a default answer, which is often the wrong one, or it stalls outright.
Exception mapping deserves to be treated as core work rather than overhead. A Harvard Business Review analysis found that organizations running structured process redesign before automation achieve substantially higher ROI than those that automate first and clean up later, largely because exception mapping is part of that redesign rather than an afterthought bolted onto it. MIT Sloan Management Review found that most digital initiatives fail because organizations digitize before they rationalize. Sequence matters here as much as the tooling does.
In practice, this starts with a process audit: operations leaders walk through how a process runs today, including the informal fixes nobody wrote down, rather than the version on the declared flowchart. Only once that picture is honest does it make sense to decide what to rebuild and in what order. The organizations that consistently generate value from AI reflect this discipline in how they allocate effort. Roughly 70% of it goes to people and process, with markedly smaller shares going to data and to the algorithms themselves. Most organizations spend in the reverse order, and the case record above is largely a record of what that reversal costs.
Construction and logistics as the sectors where the cost of skipping this sequence is non-recoverable
Physical operations don't get the luxury of rolling back a mistake once it has compounded at machine speed. Software can patch overnight. A miscoordinated construction schedule or a corrupted freight routing decision can't simply be reverted once the downstream consequences have already played out on a job site or in a warehouse.
The fail-fast model only holds where failure is cheap: Target Canada could not "iterate" its way out of empty stores, and Tesla could not A/B test a jammed assembly line. The operative rule is to fail fast where failure is genuinely cheap, and to stabilize first everywhere failure compounds. Construction and logistics sit firmly in the second category.
Construction procurement makes the point concretely. A Deloitte audit of procurement practices across mid-size contractors found an average of roughly 17% duplicate vendor records per firm. Duplicate supplier entries split purchase orders across records that should be unified, and in doing so they hide patterns, repeated invoice uplifts, off-contract buying, that would otherwise stand out. Layering a spend-analysis or invoice-checking automation on top of a vendor master in that condition means the tool enforces those patterns instead of surfacing them. It enforces the corrupted data at scale, with total consistency. The vendor list has to get fixed before any automation built on top of it can be trusted.
Adoption numbers reflect that same gap. A large share of construction firms report zero AI implementation, and almost none have reached organization-wide adoption. The tools involved exist and work; the constraint is process and data readiness, not availability.
Logistics runs into a related wall. Most logistics operators remain stuck in ad-hoc experimentation, held back less by a shortage of capable tools than by process and data readiness gaps that block any real structured rollout. The average enterprise now runs well over a thousand discrete applications, and only roughly a quarter to a third of them are actually integrated with each other. Automate any layer of a workflow while that much fragmentation persists, and every integration point becomes a place where faults get introduced rather than caught. BCG's research on manufacturing finds a similar ceiling: only about a third of manufacturing digital transformations achieve real operating-model impact, and success depends on whether the operating model got rebuilt before automation was layered onto it.
Successful sequencing: Kuehne+Nagel and the EY/Salesforce/JPMorgan pattern
The organizations that generate consistent, scalable value from AI share one discipline that has nothing to do with which model they use. They design the decision architecture, routing rules, confidence thresholds, escalation paths, governance, before they turn automation on, not after.
Kuehne+Nagel's customs classification work shows what that looks like in practice. The company applies customs classification at scale across dozens of countries, built on a tiered confidence-scoring architecture that was designed before deployment. High-confidence declarations process automatically. Mid-confidence cases route to expedited human review. Low-confidence cases go to specialist brokers rather than getting forced through the automated path regardless. The differentiator wasn't the underlying model. It was the pre-automation decision architecture, the tiering, the routing logic, the escalation paths, all defined and validated before a single case moved through the system.
EY, Salesforce, and JPMorgan show a version of the same discipline at much larger scale. The common element underneath these deployments has been described as governance and orchestration architecture built before scale, not retrofitted once problems started surfacing.
The infrastructure layer underneath these deployments matters as much as the governance decisions sitting on top of it. Agentic systems confined to a single team or an experimental sandbox rarely produce lasting value without deep integration into the organization's core business systems and explicit escalation paths back to human oversight. The ERP and orchestration layer is what separates deployments that hold up under real production load from the ones that stall out. Effective AI governance in 2026 looks less like a policy document sitting in a compliance folder and more like an operating model in its own right: clearly defined boundaries on what autonomous systems can act on, transparent validation of the models and decisions involved, and auditability that scales across complex, cross-system workflows rather than breaking down at the first handoff.
Every organization in this section made the same investment before it made the automation investment: process clarity and governance architecture built up front, not deferred until something failed. That sequencing is what makes the automation that follows worth trusting at scale, and it is the same sequencing that Target Canada, Tesla, Air Canada, McDonald's, and Meta each skipped, in their own way, before the record caught up with them.
Sources
- AI Process Improvement: 5 Hard Truths Before You Automate
- Why AI Projects Fail in Enterprises (and How to Avoid It in 2026)
- Stop Automating Old Processes. Design New Ones Instead.
- Case Study 7: The $2.5 Billion Cross-Border Expansion Mistake by Target - Henrico Dolfing
- Case Study of Air Canada's Chatbot Misleading on Bereavement Fares
- Why AI Isn't Delivering ROI in Logistics in 2026 | BCG
- Kuehne+Nagel Deploys AI for Enhanced Supply Chain Visibility | AI Magazine