Listen to this article
Jump to a section
In this article
Executive Summary
A broken process, accelerated by AI, produces bad results faster and at scale. Every failure pattern below traces back to missing engineering discipline, not model quality.
$62M
spent before IBM Watson for Oncology was shut down, deployed into clinical workflows that were never standardised
$500M+
lost when Zillow’s iBuying algorithm couldn’t account for local market nuance human experts navigate intuitively
$440M
lost by Knight Capital in 45 minutes when a monitoring gap let a software error run undetected
Core conclusions
- Fix the underlying process before automating it. AI speeds up whatever workflow it’s dropped into, broken or not.
- The costliest failures share a root cause: no defined success metrics, no human review step, or no observability, not poor model quality.
- Every case here was preventable with basic operational safeguards defined before deployment.
All nine, with the price tags
Save it before your next deployment review, or send it to whoever owns the budget.










Failure 1 is automating a broken process
A bad process, accelerated by AI, produces bad results faster and at scale.
Think: automated email replies with no triage logic, so customers get contradictory answers. Or AI-generated reports that nobody reads because no one defined what decision they were supposed to inform.
Real example: IBM Watson for Oncology at MD Anderson Cancer Center. The system was deployed into clinical workflows that weren’t standardised. Clinicians couldn’t act on its recommendations because the process underneath was inconsistent. The project was eventually shut down after USD 62 million was spent.
Fix it first. Map your decision points, remove the inefficiencies, then introduce AI.
Failure 2 is having no definition of success
If you can’t measure it, you can’t manage it. And you can’t justify it.
AI deployments without clear KPIs drift. You end up with tools that are “being used” but not delivering value. And when budget reviews come around, you can’t defend the investment.
Real example: The UK Government’s early chatbot deployments, including support bots for GOV.UK , saw growing usage, but limited evidence of improved resolution rates or reduced costs. Without meaningful success metrics, the programmes were redesigned multiple times.
Before launch: set a baseline.
Failure 3 is trusting AI output without checking it
AI outputs are probabilistic. That means they’re sometimes confidently wrong.
When teams treat AI as authoritative and skip the review step, small errors compound into serious problems.
Real example: In the 2023 US court case Mata v. Avianca, lawyers submitted AI-generated legal citations. The cases didn’t exist. The court issued sanctions. This became one of the most-cited cautionary tales in the legal profession.
Human-in-the-loop is not optional for high-stakes decisions. Build review into the workflow, not as an afterthought, but as a designed step. As these review layers mature, some organisations are deliberately redesigning them: see From Human-in-the-Loop to AI-on-the-Loop.
Failure 4 is ignoring data quality and bias
AI learns from data. If your data reflects past biases, your AI will reproduce and amplify them.
It shows up in hiring, lending, lead scoring, and recommendations.
Real example: Amazon built an AI recruiting tool between 2014 and 2017. It learned from historical hiring data, which was heavily male-dominated. The model systematically downgraded CVs from women. Amazon scrapped it before launch, but only after internal audits caught the bias.
Audit your training data before you train your model. Check for bias continuously. Build retraining cycles from the start.
Failure 5 is building with no visibility into what’s happening
If an AI system fails and you have no logs, no traces, no alerts, you have no way to fix it.
Worse, you might not even know it failed.
Real example: Knight Capital’s 2012 trading system collapse isn’t strictly AI, but it’s the defining automation disaster. A software deployment error went undetected because there were no proper monitoring or rollback mechanisms. The firm lost USD 440 million in 45 minutes and was effectively put out of business.
Log inputs and outputs. Set performance thresholds. Build alerts. Treat observability as a first-class requirement.
Failure 6 is using AI to replace expert judgment
AI is a force multiplier for expertise. It is not a substitute for it.
When organisations deploy AI in domains they don’t understand well, the model operates without any check on its reasoning. That’s when the expensive mistakes happen.
Real example: Zillow’s iBuying programme (Zillow Offers) used an algorithm to buy and sell homes at scale. The model couldn’t account for the nuanced, local market dynamics that experienced real estate professionals navigate intuitively. Zillow lost over USD 500 million and shut down the programme in 2021.
AI should work alongside experts, not instead of them. If your team doesn’t have domain knowledge, the model won’t compensate for that gap.
Failure 7 is AI that creates more work
This one catches teams off guard.
If the surrounding workflow isn’t redesigned, AI can increase cognitive load. Agents reviewing, correcting, and contextualising AI suggestions every cycle is a more complicated manual process, not automation.
Real example: Early rollouts of AI-assisted customer service tools at large call centres, including some reported in post-implementation reviews of major US telecom providers, found that average handling times increased. Staff spent more time validating AI suggestions than they saved by using them.
Measure end-to-end task time, including review and correction. Don’t benchmark AI in isolation from the full workflow.
Failure 8 is choosing the wrong tool
Popularity is not a technical specification.
Grabbing the most talked-about AI platform without checking whether it fits your actual requirements (latency, accuracy, scale, security and integration) leads to expensive replacements.
Real example: Many banks and financial services firms adopted generic conversational AI platforms in the early wave of chatbot adoption, only to replace them within 18–24 months with domain-specific systems. The generic tools couldn’t handle intent complexity, compliance constraints, or core system integration requirements.
Start with requirements. Work backwards to the tool. Not the other way around.
Failure 9 is gradual degradation no one notices
This is the quietest failure mode, and often the most costly.
AI models degrade over time as the world changes. Without continuous monitoring, performance slips gradually, and by the time someone notices, the damage is already done.
Real example: Meta’s content recommendation algorithms have undergone multiple corrective interventions after researchers and regulators identified gradual amplification effects: content optimised for engagement that drifted toward increasingly extreme material over time. The degradation wasn’t visible through standard performance dashboards.
Set explicit performance thresholds. Monitor drift. Build alerts that fire before the problem becomes visible to users.
There’s a thread running through every failure above.
The absence of engineering discipline around the AI is the real problem. If this list resonated, the individual-behaviour version of it is 9 Common Mistakes Intelligent People Make with AI. Successful deployments treat AI as part of an integrated operational system with process design, measurable outcomes, human oversight, and continuous monitoring built in from the start.
AI fails inside poorly designed systems.
And if this was useful, share it with someone deploying AI right now. They’ll thank you later.
Evidence & Methodology
Five of these nine failures are on the public record, with a filing, a court docket, or a named reporter attached. The other four are patterns I have seen described in post-implementation reviews, without one specific source I can point to. I would rather tell you which is which than let all nine look equally documented.
| Claim | Source | Grade |
|---|---|---|
| IBM Watson for Oncology, Zillow’s iBuying programme, and Knight Capital’s trading collapse each cost tens or hundreds of millions of dollars | Forbes reporting, Zillow’s own SEC filing, and the SEC’s Knight Capital order | Documented |
| Amazon’s AI recruiting tool systematically downgraded CVs from women and was scrapped before launch | Reuters reporting on Amazon’s internal audit findings | Documented |
| Lawyers who submitted AI-hallucinated legal citations were sanctioned by the court | Mata v. Avianca court docket | Court record |
| Early UK government chatbots, telecom call-centre AI, and bank chatbot rollouts underperformed or got replaced | Described in the piece from post-implementation reviews, without one specific named source | Anecdotal |
Sources
- Herper, M. (2017). MD Anderson benches IBM Watson in setback for artificial intelligence in medicine. Forbes.
- Dastin, J. (2018). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters.
- U.S. Securities and Exchange Commission. (2013). In the matter of Knight Capital Americas LLC (Release No. 70694).
- Zillow Group, Inc. (2021). Zillow Group reports third-quarter 2021 financial results & shares plan to wind down Zillow Offers operations.
- Mata v. Avianca, Inc., No. 1:22-cv-01461 (S.D.N.Y. 2023). Case docket and order on sanctions.
Free tool
TRACE Agent Evaluation
Every failure above maps to a missing check (traceability, reversibility, acceptance, compliance, or escalation), so score your next candidate task before it becomes failure 10.
Was this useful?







