← All Articles

How to Calibrate Every Dimension of Your Agentic AI Scorecard

12 min readAgentic AISharePDF

Listen to this article

How to Calibrate Every Dimension of Your Agentic AI Scorecard

0:00
Jump to a section

Executive Summary

A six-dimension score for your first agentic AI project only works if it produces the same number twice. Hand the same project to two people with only “rate this 1 to 10” as guidance, and you will get two different scores, not because either person is wrong, but because neither one was told what a 5 is supposed to look like. This page is the rubric that closes that gap: concrete, checkable anchors for what a 2, a 5, and an 8 mean on each of the six dimensions.

Core conclusions

  • An unanchored 1-to-10 scale is not a scoring method. It is a survey question wearing a spreadsheet.
  • Three of the six dimensions can be scored against something you can point to: an audit, a calendar, a signed-off number. The other three are closer to informed judgement, and should be scored by more than one person.
  • Anchors are not permanent. As your data, tooling, and governance mature, what counts as an 8 today should get harder to earn.

Why the same project gets two different scores

Score any candidate agentic AI project against the six-dimension model, strategic fit, data readiness, ease of implementation, expected ROI, speed to value, autonomy readiness, and you will notice something within the first ten minutes of the exercise. Ask two people to rate the same project on “ease of implementation,” and one comes back with a 3, the other with a 7. Neither is being dishonest. They are answering different questions, because “how hard is this” has no fixed meaning until someone defines what hard looks like.

This is the part that gets skipped when a scoring method gets adopted quickly. The six dimensions are the easy half. The rubric underneath each one, the thing that tells a scorer what evidence justifies a 2 versus a 5 versus an 8, is the half that makes the number defensible in a room. Without it, you have replaced one gut call with six gut calls and called it rigour.

What a 2, a 5, and an 8 mean, dimension by dimension

Use this table as the shared reference before anyone scores anything. Read it out loud in the room if you have to. The goal is for two people looking at the same project to land within a point or two of each other, not to land on an identical number.

DimensionA 2 looks likeA 5 looks likeAn 8 looks like
Strategic FitSolves a problem nobody above your level is trackingSupported by one department head, not yet on a leadership scorecardNamed in this year’s leadership priorities or board deck
Data ReadinessThe data does not exist yet, or lives in someone’s inboxData exists but needs real cleanup, consolidation, or access work firstData is already accessible, structured, and used for reporting today
Ease of ImplementationCustom model work, multiple system integrations, heavy change managementOne clear integration and a moderate change to how a team worksLargely off-the-shelf, one system, minimal retraining needed
Expected ROIValue is theoretical, and nobody has tried to size itA rough estimate exists but rests on assumptions nobody has testedA specific, defensible number, ideally run through an ROI calculation, not a guess
Speed to ValueOver six months before any measurable result is visibleA visible result in eight to twelve weeksA measurable result within four weeks of go-live
Autonomy ReadinessIt could act on real systems today with no logging, no approval step, and no way to catch a bad action fastLogging exists; a person reviews output before anything reaches a customer or a system of recordGuardrails, logging, and an approval point are already built and tested for this exact level of autonomy

Three columns is a starting point, not a ceiling. If your team keeps landing between two anchors, that is useful information on its own: it usually means the project needs a short discovery step before it deserves a confident score at all. The full rubric below fills in the seven anchors this table skips, a 1, 3, 4, 6, 7, 9, and 10 for every dimension, so a scorer never has to guess what sits between the columns.

Free tool

AI Score Calibration Guide

The full 1-to-10 rubric for all six dimensions. Set a dimension and a score, see exactly what evidence it’s supposed to represent.

Free tool

AI Use Case Prioritisation Matrix

Score your candidates against these same six dimensions, with anchor guidance built into each one, and get a ranked business merit sequence plus each one’s autonomy clearance tier.

Score three dimensions against evidence

Not all six dimensions carry the same kind of uncertainty, and the rubric works better once you notice which is which.

Data Readiness, Speed to Value, and Autonomy Readiness can usually be scored against something concrete: a data audit finding, a calendar count of weeks to a pilot, an actual inventory of what logging and approval controls exist right now. If someone’s score on one of these three cannot be traced to a specific piece of evidence, treat the score as provisional.

Strategic Fit, Ease of Implementation, and Expected ROI are closer to informed judgement, especially on a first pass. Strategic fit depends on reading leadership’s actual priorities correctly. Ease of implementation depends on guessing at integration and change management costs before anyone has scoped the work. Expected ROI, before a real calculation, is often closer to a hopeful estimate than a number. These three deserve more than one scorer, and they deserve the confidence discount described in the original scoring guide: rate how sure you are, and let that pull an overconfident score back down before it wins on paper.

Free tool

AI ROI Calculator

Turn an Expected ROI guess into the kind of number that earns an 8 on this rubric, not a hopeful 5.

A short example of two scores converging

A mid-market logistics company scored a proposed dispatch-scheduling agent twice: once before anyone had agreed on anchors, and once after.

DimensionReviewer A, no rubricReviewer B, no rubricReviewer A, with rubricReviewer B, with rubric
Ease of Implementation3745
Autonomy Readiness6234
Expected ROI8566

Before the rubric, the two reviewers’ scores on autonomy readiness were four points apart, enough to swing whether this project should even be considered for a first launch at full autonomy. Reviewer A had been picturing the agent’s long-term potential. Reviewer B had been picturing what would happen the first time it double-booked a driver with no one watching. Neither was wrong about the project. They were answering different implicit questions until the rubric forced them onto the same one: what does the evidence in front of us show, right now, not what could this become.

After the rubric, the scores did not become identical. They became close enough to have a real conversation about the actual one-point gap, instead of arguing about which imagined version of the project was correct.

Anchors should get harder to earn as your programme matures

The anchors in the table above are calibrated for a company scoring its first or second agentic AI project. That calibration should not stay fixed.

Once your organisation has shipped three or four agents into production, an 8 on Autonomy Readiness should require more than “logging exists.” It should require a track record: incidents caught before they reached a customer, an approval process that has been exercised under a real mistake. The same logic applies to Ease of Implementation. Integration work that looked hard on your first project should look routine by your fourth, if you built anything reusable along the way. If your anchors never move, your scoring model is rewarding the same level of caution forever, even after your organisation has earned the right to expect more of itself.

Revisit the rubric itself, not just the individual project scores, roughly once a year or after any project that seriously under- or over-performed its score. That review is a five-minute conversation. Skipping it is how a five-year-old scoring model keeps telling a mature AI programme it is still a beginner.

Agree the anchors before you score anything

Five minutes reading the table above as a group saves hours of arguing about what a 6 means.

Score the judgement calls with more than one person

Strategic fit, ease of implementation, and expected ROI benefit from a second opinion far more than data readiness does.

Revisit the rubric once a year

What earns an 8 today should get harder to earn as your data, tooling, and governance mature.

Where to go next on this site


Still not sure whether your team’s scores are calibrated to the same reality? That is exactly the conversation my consulting work begins with.

Was this useful?

Terence Kok
Before You Go

The question I get after someone tries the six-dimension score for the first time is almost never about the method. It is 'my head of ops gave this a 7 and I gave it a 4, who's right.' Nobody is right until you both agree what a 4 and a 7 mean on that dimension. That agreement is this page. It is duller than the scoring method itself, and it is the part that makes the number worth trusting.

Terence Kok