← All Articles

The Lean Way to Build AI: Small Bets, Real Evidence, Fewer Failed Pilots

13 min readAI StrategySharePDF

Listen to this article

The Lean Way to Build AI: Small Bets, Real Evidence, Fewer Failed Pilots

0:00
Jump to a section

Executive Summary

Lean thinking was built to stop factories and startups from betting everything on an idea before anyone tested its smallest possible version. Applied to AI, the same discipline that shortens time to a working product also happens to be the safest way to adopt a technology that nobody in the room fully understands yet.

Core conclusions

  • Most AI pilots do not fail because the model performed badly. They fail because the first deployment was scoped like a rollout rather than an experiment, so there was nothing to learn from when it did not work.
  • The Toyota Production System and the Lean Startup method were built on the same insight decades apart: small batches with a real feedback loop catch expensive mistakes while they are still cheap to fix.
  • Five Lean practices translate directly to AI adoption, and the one most programmes are missing is the andon cord, a named person with the standing authority to stop a deployment the moment it produces a defect.

Most AI pilots fail before the model gets a chance to be wrong

I sit in a lot of steering committee meetings where the AI pilot is six months old, the budget has been spent, and nobody can point to a number that moved. The conversation that follows almost always turns to the model. Wrong vendor, wrong architecture, wrong prompt strategy. In my experience that is rarely the actual fault line.

The pattern I see instead is a pilot that was scoped and governed like a finished rollout from day one. Every department that might eventually use the tool got a seat on the steering committee before a single real user had touched it. The success criteria were written as a business case, with a payback period, rather than as a hypothesis with a date attached for checking it. By the time the thing was live, so much organisational weight was resting on it succeeding that nobody involved could afford to notice early signs it was not working, let alone say so out loud.

This is not a rare failure. MIT’s NANDA initiative surveyed enterprise generative AI deployments in 2025 and found that 95 percent of pilots produced no measurable profit and loss impact, despite tens of billions of dollars in enterprise spending. Gartner had already predicted the shape of this a year earlier, forecasting that at least 30 percent of generative AI projects would be abandoned after proof of concept, citing poor data quality, unclear business value, and escalating cost as the recurring causes. Neither number is a story about bad models. It is a story about how the projects were built.

There is an older discipline that solved exactly this problem, just not for AI, and it is worth borrowing on purpose.

Lean was built to solve exactly this problem, just for factories and startups

The Toyota Production System, developed from the 1950s onward, was a response to a manufacturer that could not afford Detroit’s approach of building at enormous scale and fixing defects downstream. Toyota’s answer was to shrink the batch, put the person doing the work in direct contact with the outcome, and give every worker on the line the authority to stop production the instant something looked wrong. That last piece had a name and a physical form: the andon cord, a rope running the length of the assembly line that any worker could pull.

Half a century later, Eric Ries described the same logic for software startups in The Lean Startup, formalising it as build, measure, learn. Build the smallest version of the idea that can produce real evidence. Measure the thing you said you would measure before you built it. Learn from the result, including the result you did not want, and let that evidence decide whether you continue, change direction, or stop.

What connects a car factory in Nagoya and a software startup in Silicon Valley is not the product. It is the belief that the most expensive mistake in the room is the one nobody notices until the batch is large and the sunk cost is real. Small batches are not a courtesy to nervous executives. They are how you keep a mistake cheap enough that finding it does not feel like failure.

AI has all the conditions that made this discipline necessary in the first place. The systems behave in ways nobody fully specified, the cost of a large-batch deployment is high, and the people closest to where it will run a workflow are usually the last to be consulted. Lean was not built for AI. It was built for exactly this kind of uncertainty, which is why it transfers with almost no translation required.

Five Lean practices translate directly into an AI programme

Go and see the actual workflow before you automate any of it. Toyota called this genchi genbutsu, go to the real place. Most AI use cases I am asked to evaluate were identified from a slide describing what the department does, not from watching what happens at the desk where the work gets done. The gap between the two is usually where the AI ends up failing, automating a step that looks correct on the process map and skips a judgment call the person doing the job makes without thinking about it.

Specify value from the user’s problem, not the model’s capability. A recurring pattern in failed pilots is a team that starts from “we have access to a strong model, what should we build,” rather than from a specific, named point of friction someone already complains about. The first question produces a demo. The second produces a use case you can measure.

Work in one real workflow, with real users, before you scope anything wider. This is the direct analogue of Toyota’s small batch. A pilot covering one team, one workflow, and a small number of real transactions gives you evidence you can trust. A pilot covering four departments at once gives you a rollout wearing a pilot’s name, and rollouts do not have a real off switch.

Decide your stop or go metric and your review date before you start. This is build, measure, learn with the measuring part taken seriously. Write down, in advance, the specific number that would make you continue and the specific number that would make you stop. Put a date on the calendar to look at it. Without that discipline, the pilot’s own momentum becomes the evidence, and momentum is not a metric.

Give someone a real andon cord. This is the piece most AI governance frameworks describe in policy language and almost none of them build as an actual mechanism. It needs to be one named person, not a committee, with the explicit authority to pause a deployment the moment it produces an error the organisation was not prepared to accept, and no requirement to build a business case first for having pulled it. Toyota’s line workers did not need to justify stopping the line. They needed the stop to already be theirs to call.

Free tool

AI Use Case Prioritisation Matrix

Before you scope a small batch, you need to know which candidate deserves to be it. Score up to five initiatives and get a ranked sequence instead of a debate.

Small batches reduce risk before they increase speed

Executives usually adopt Lean thinking for the speed. The reason I keep recommending it to boards is the risk reduction, which arrives first and matters more for AI specifically.

A small batch limits your exposure while you still have almost no evidence about how the system behaves in your organisation’s actual conditions, with your actual data, against your actual edge cases. That is precisely the period when you know the least and can least afford a large-scale mistake. A four-department rollout that goes wrong is a governance incident with a paper trail and a board conversation attached. A one-workflow pilot that goes wrong is a Tuesday, and you fix it before lunch.

There is a second, quieter benefit. A small, honestly measured pilot is the only kind that can fail without becoming an organisational crisis. When failure is survivable, people report it accurately. When failure would mean writing off a large, visible investment, the incentives point everyone toward finding a way to call it a success regardless of what happened. I have watched teams redefine what counts as a win six weeks into a struggling pilot, not out of dishonesty, but because admitting the truth had become more expensive than bending the metric. Lean’s small batches keep the cost of honesty low enough that people can still afford it.

Table comparing a rollout-scoped AI pilot against a Lean-scoped AI pilot across stakeholders involved, success criteria, batch size, review mechanism, and the cost of being wrong
The scoping decision, not the model, is what determines whether a mistake stays cheap.

What this looks like in the first few weeks

Pick one workflow with a named owner and a real, recurring point of friction, not the most strategically impressive candidate on the list. Go and watch the work being done before anyone writes a requirements document. Write down, on one page, the specific metric that will tell you whether this worked, the threshold that means stop, and the date you will look at it. Name the person who holds the andon cord for this pilot specifically, and confirm out loud, in the room, that everyone accepts they can pull it without first building a case.

Then build the smallest version that can produce real evidence against real transactions, not a demo against curated examples. Run it with actual users doing actual work. Measure what you said you would measure, on the date you said you would look. If the evidence says stop, stop, and count that as the pilot having worked exactly as designed, because a small experiment that correctly rules something out has done its job.

None of this requires new governance infrastructure or a change management programme. It requires deciding, before you start, that you would rather find out something is wrong in week three of one workflow than in month nine of four.

Evidence & Methodology

These four claims are not equally solid. Two are measured. One is a forecast. One is mine, from client work, not a survey. Here is which is which, so you can decide how much weight to put on each.

ClaimSourceGrade
95% of generative AI pilots show no measurable P&L impactMIT NANDA, 2025. Surveyed 300+ enterprise deployments and 150+ executivesMeasured
30% of GenAI projects abandoned after proof of conceptGartner, 2024. A forecast, not a completed countForecast
Small-batch, andon-cord discipline reduces AI deployment riskToyota Production System and Lean Startup, built for factories and software, never tested on AIAdapted
Pilots fail on scoping, not model qualityMy own pattern from governance and audit-readiness engagementsMy call

Sources

  1. Fortune. (2025). MIT report: 95% of generative AI pilots at companies are failing, reporting on MIT NANDA’s “The GenAI Divide: State of AI in Business 2025” study.
  2. Gartner. (2024). Gartner predicts 30% of generative AI projects will be abandoned after proof of concept by end of 2025.

Where to start this quarter

Look at whatever AI initiative on your roadmap is currently scoped the widest, the one with the most departments, the most stakeholders, and the most impressive-sounding business case. Ask what the smallest version of it would look like: one workflow, one team, one metric, one named person holding the andon cord.

Start there instead. Not because the bigger version is wrong, but because you do not yet have the evidence to know whether it is, and a small batch is how you get that evidence while it is still cheap to act on.


If the AI initiative on your desk is scoped bigger than the evidence for it, that is usually the first thing worth fixing before anything else. My consulting work is built around getting that scope right early, before it becomes an expensive lesson.

Was this useful?

Terence Kok
Before You Go

I did not expect to end up recommending a manufacturing philosophy from 1950s Japan to boards worried about generative AI, but the fit is closer than it looks. Toyota built the andon cord because nobody trusted any one person's judgment about when a defect was serious enough to stop the line, so they built a mechanism instead. That is exactly the gap I keep finding in AI governance conversations. If you have found a different way to get that same discipline into an AI programme, I would like to hear it.

Terence Kok