← All Articles

Why Autonomy Readiness Should Cap Your Agentic AI Score

11 min readGovernance & RiskSharePDF

Listen to this article

Why Autonomy Readiness Should Cap Your Agentic AI Score

0:00
Jump to a section

Executive Summary

The six-dimension scoring method for agentic AI projects has a quiet flaw if you apply it as a straight average: a project can score high everywhere else and low on autonomy readiness, and still come out on top. Averaging treats a governance gap like just another weakness to be offset by strength elsewhere. It is not. This page covers why autonomy readiness needs to work as a gate on what the project is allowed to do, not a sixth number added to the pile.

Core conclusions

  • An average lets five strong scores cancel out one dangerous one. A gate does not let that happen, because it caps what the project is cleared to do.
  • The fix is not to disqualify a project with low autonomy readiness. It is to launch it at a lower level of independence than the pitch deck assumed, and re-score it once the gap closes.
  • Present two numbers to leadership, not one: the project’s business merit, and the level of autonomy it is cleared to run at today. A single blended score erases that distinction on purpose.

The problem with adding six numbers together

Picture a collections agent: it emails overdue customers, negotiates short payment plans within pre-set limits, and updates the ledger automatically. Score it against the six dimensions from the original scoring model, and here is a plausible result.

DimensionScore (1-10)
Strategic Fit8
Data Readiness9
Ease of Implementation7
Expected ROI9
Speed to Value8
Autonomy Readiness2
Straight average7.2

A 7.2 out of 10 is a strong score. In a room full of competing ideas, this project wins. And yet the one dimension it scored worst on is the one that determines whether it should be allowed to email a customer a legally sensitive payment offer with nobody checking first. A 2 on autonomy readiness means there is no meaningful logging, no approval step, and no fast way to catch or undo a bad message before it reaches someone’s inbox. None of that gets fixed by the project also having excellent data and an obvious ROI. Those are true and irrelevant to the specific risk.

This is what an average does structurally: it lets strength in one place cancel weakness in another, on the assumption that all six dimensions are the same kind of thing. Five of them are inputs to how good an idea this is. The sixth is a precondition for whether it is safe to run the way the pitch describes. Blending a precondition into an average hides it exactly when it matters most.

Two ways to fix the math

There are two workable fixes, and they solve slightly different problems.

A hard floor on autonomy level. Set a threshold, for example anything scoring below 4 on Autonomy Readiness cannot launch at its intended level of independence, full stop, regardless of the total score. The project is not disqualified. It is automatically capped to a lower autonomy tier until the gap closes: an agent that was pitched as fully autonomous instead launches as one that drafts an action and waits for a person to approve it, or one that only checks a system and reports back. This keeps the project alive and keeps the ranking honest at the same time.

A multiplier on the total, the way the original model already treats confidence. Instead of averaging six equal inputs, apply Autonomy Readiness as a scaling factor on the rest of the score, similar to the confidence discount described in the base scoring method. A project scoring well everywhere else but a 2 on autonomy readiness gets pulled down hard. This keeps the ranking as a single number, which some leadership teams prefer, while making sure that number cannot lie about the governance gap underneath it.

Either fix works. The floor is easier to explain and enforce operationally: it changes what the project is cleared to build. The multiplier is easier to keep inside a single spreadsheet column if your team is committed to one ranked number. What does not work is leaving autonomy readiness as a seventh input weighted the same as the other five, because that is exactly the setup that let the collections agent above land a 7.2.

The AI Use Case Prioritisation Matrix linked below now runs the floor version of this fix automatically. Score all six dimensions and it separates the two numbers for you: a Business Merit Score built from the other five, and an autonomy clearance tier, checks and reports only, drafts and waits for approval, or checks and acts on its own, calculated from Autonomy Readiness alone. Nothing to compute by hand.

Free tool

AI Trust, Risk & Governance Dashboard

Get an honest number for your actual guardrails, the input the gate above depends on, before you apply it to any project’s score.

What the collections agent looks like with the gate applied

Run the same project through a hard floor set at 4, and the picture changes without the project disappearing from the roadmap.

Without a gateWith a gate at 4
Ranking number shown to leadership7.2, ranked first7.2 business merit, but capped
Autonomy level cleared to launch atFull autonomy, as pitchedDraft-and-approve: a person reviews every payment offer before it sends
What ships in week oneThe agent negotiates and emails on its ownThe agent drafts the offer; a collections officer approves each one
What re-scores the project upward laterNothing changes automaticallyLogging and an approval audit trail get built during the draft-and-approve phase, which is exactly what autonomy readiness measures

Nothing about the project’s underlying business case changed. What changed is that the ranking now tells the truth about what it is safe to build first. The project still goes ahead. It goes ahead at a level of independence its actual controls can support, and the very act of running it at that lower tier is what generates the evidence needed to earn a higher autonomy readiness score on the next pass.

Separate the ranking from the clearance

The clearest fix, in practice, is to stop asking one number to answer two different questions. Present leadership with the business merit score, everything except autonomy readiness, alongside the autonomy tier the project is cleared to run at today: checks and reports only, drafts and waits for approval, or checks and acts entirely on its own. A project can have outstanding business merit and still be cleared for the lowest of the three tiers on day one. That is not a contradiction. It is the entire point of scoring autonomy separately in the first place.

Most of the disagreement that shows up in a prioritisation meeting is a disagreement between these two questions being asked at once. Someone arguing “this is clearly our best idea” is usually right about business merit. Someone arguing “we are not ready for this” is usually right about the autonomy tier. A single blended average forces them to argue past each other over one number that was never built to answer both questions.

Set a floor before you rank anything

Decide the autonomy readiness threshold below which a project cannot launch at its intended tier, before you see any actual scores.

Cap, don’t disqualify

A low autonomy score moves a project to a lower tier. It does not remove it from the roadmap.

Show two numbers, not one

Business merit and autonomy clearance are different questions. Let leadership see both, separately.

Free tool

AI Use Case Prioritisation Matrix

Score your projects and get two numbers automatically: a Business Merit Score, and the autonomy clearance tier Autonomy Readiness earns on its own.

Where to go next on this site


Still deciding where to set your own autonomy floor? That is exactly the conversation my consulting work begins with.

Was this useful?

Terence Kok
Before You Go

I have watched a leadership team approve a project with an 8 out of 10 total score, wave through the plan, and only ask 'wait, what happens if it's wrong' after the kickoff meeting was already scheduled. The score was not lying to them. It was just built to hide the one number that mattered most inside an average of five others. Fixing the math takes an afternoon. Finding out the hard way takes a lot longer.

Terence Kok