Agent Risk Assessment Matrix

Eight questions place each AI agent on a four-by-four matrix of autonomy against consequence. The result is an oversight tier, the controls that tier requires, and the ISO/IEC 42001:2023 clause each control satisfies, printed as a register an auditor can read. Load the worked example to see five anonymised agents already filled in, then replace them with your own. No email required.

Listen to this briefing

The Agent Risk Assessment Matrix

0:00

Why agents need their own risk register

ISO/IEC 42001 asks for an AI risk assessment and an impact assessment for every system in scope. It does not say how to rate an agent that reads supplier emails, plans a sequence of actions across three systems, and posts a journal entry without asking. Most registers I am shown rate that agent the way they rate a chatbot, because the template was built for a model that answers questions rather than one that takes actions.

The fix is to separate two things the standard bundles into "risk": how much the agent does on its own, and how bad a wrong action is. Rate them on two axes and the oversight decision becomes visible, defensible and repeatable across every agent you register.

Figure 1Two axes instead of one score. A point apart on autonomy, three tiers apart once consequence is counted.
Two axes instead of one scoreA four-by-four matrix of autonomy (Assisted, Supervised, Bounded, Autonomous) against consequence (Low, Moderate, High, Severe). Each cell names an oversight tier from 1 Monitor to 4 Hold. Two example agents sit one autonomy band apart: a meeting summariser with no review lands at Supervised autonomy and Low consequence in Tier 1 Monitor; a benefit payment agent sampled after the fact lands at Bounded autonomy and Severe consequence in Tier 4 Hold.ConsequenceAutonomySevereTIER 2SuperviseTIER 3ApproveTIER 4HoldTIER 4HoldHighTIER 2SuperviseTIER 2SuperviseTIER 3ApproveTIER 4HoldModerateTIER 1MonitorTIER 2SuperviseTIER 2SuperviseTIER 3ApproveLowTIER 1MonitorTIER 1MonitorTIER 1MonitorTIER 2SuperviseAssistedSupervisedBoundedAutonomous22Meeting summariser, no reviewReads and reports across several systems; nobodychecks the output unless a problem is raised.Supervised × LowTier 1 Monitor11Benefit payment agent, sampled afterwardsPays claimants within a limit across several systems;money sent, regulated data, a decision on entitlements.Bounded × SevereTier 4 Hold
Figure 2How the eight answers become a register row: two banded scores, one matrix lookup, one tier with its control set and holds.
How the eight answers become a register rowThree answers about how much the agent does on its own are scored 0 to 3 each and summed to an autonomy score out of 9, banded Assisted, Supervised, Bounded or Autonomous. Five answers about how bad a wrong action is are summed to a consequence score out of 15, banded Low, Moderate, High or Severe. The two bands are looked up in a four-by-four matrix to give an oversight tier from 1 to 4. The register row records the tier and its oversight mode, the control set (7 controls for every agent plus more by tier or by a single answer), and any holds.1. Eight answersA. How much it does on its ownAction scopeHuman checkpointTask chainscored 0–3 eachB. How bad a wrong action isReversibilityData sensitivityBlast radiusExposure to untrusted inputRegulatory weightscored 0–3 each2. Two scores, bandedAutonomy 0–9AssistedSupervisedBoundedAutonomousConsequence 0–15LowModerateHighSevere3. Matrix lookup2344223412231112Autonomy →↑ Consequence1 Monitor2 Supervise3 Approve4 Hold4. Register rowOversight tierTier 1 to 4 and thereview it demandsControl set7 for every agent,more by tier or by asingle answerHoldsHalt never tested,no named owner,no approver at Tier 2+Each control carries its ISO/IEC 42001:2023 clause, so the row reads into the risk treatment plan and the statement of applicability.

The clause references are to the published standard. The band thresholds and the matrix are my calibration from applying the standard to agent deployments; the basis table further down shows what each rule rests on.

Describe one agent

Answer the eight questions for the agent as it will run in production, not as the vendor demonstrated it.

A How much it does on its own 0 / 9

Retrieves, summarises or answers. Nothing changes in any system.

Nothing executes until a person approves it.

A single call into a single tool or system.

B How bad a wrong action is 0 / 15

Undo is a click and nobody outside the team notices.

Published material or internal information with no access restriction.

The person using it sees the error and corrects it.

Structured data from systems you control, entered by staff.

General law only; no sector or activity rule names this task.

C Register fields

Your agent risk register

Agents you add appear on the matrix and in the register below. Everything stays in this browser; nothing is sent to me.

Consequence
Severe13–15
T2 Supervise
T3 Approve
T4 Hold
T4 Hold
High9–12
T2 Supervise
T2 Supervise
T3 Approve
T4 Hold
Moderate5–8
T1 Monitor
T2 Supervise
T2 Supervise
T3 Approve
Low0–4
T1 Monitor
T1 Monitor
T1 Monitor
T2 Supervise
Assisted0–2
Supervised3–5
Bounded6–7
Autonomous8–9
Autonomy

The outlined marker is the agent in the form. Numbered markers are agents in the register.

#AgentAutonomyConsequenceTierControlsHoldsActions
No agents yet. Add one above, or load the worked example.

The PDF opens a register with one page per agent: the eight answers, the tier, every required control with its ISO/IEC 42001 reference, and a sign-off block for the owner, the approver and the residual-risk acceptance. The worked example is an illustrative composite; none of the five agents is a client record.

This is a self-assessment against your own description of each agent, not an independent risk assessment, an audit or a legal opinion. The tier and the control set are a starting point for your own risk treatment plan and statement of applicability, not a substitute for them. Your entries stay in your browser and are not sent to Terence Kok or reviewed by anyone.

The seventeen controls and where ISO/IEC 42001 asks for them

Seven controls apply to every agent, whatever its tier. The rest are pulled in by the tier, or by a single answer that makes them necessary at any tier. The register names the reason against each control.

Figure 3How the control set stacks. Tier 4 carries the same controls as Tier 3 and adds a hold on deployment.
How the control set stacks by tierFour columns, one per oversight tier. Tier 1 carries the 7 controls that apply to every agent (C01–C07). Tier 2 adds 5 (C08–C12). Tier 3 adds 4 (C13, C14, C15, C17). Tier 4 carries the same 16 controls and deployment is held until the agent is redesigned. Below the columns, the answers that pull a control in at any tier: C08: partly or fully irreversible actions; C09: customers, the public or third parties affected, or personal data; C10: statutory or licensed activity; C12: reads or receives untrusted content; C14: personal or regulated data; C15: statutory or licensed activity; C16: instructs other agents (this answer alone, at any tier).C01–C077Tier 1 Monitor7 controlsC01–C077C08–C125Tier 2 Supervise12 controlsC01–C077C08–C125C13, C14, C15, C174Tier 3 Approve16 controlsC01–C077C08–C125C13, C14, C15, C174Do not deploy as scoredTier 4 Hold16 controlsEvery agent: 7 controlsTier 2 adds: 5 controlsTier 3 adds: 4 controlsPulled in at any tier by a single answerC08: partly or fully irreversible actionsC09: customers, the public or third parties affected, or personal dataC10: statutory or licensed activityC12: reads or receives untrusted contentC14: personal or regulated dataC15: statutory or licensed activityC16: instructs other agents (this answer alone, at any tier)
Show all seventeen controls with their ISO/IEC 42001 references
ControlWhat the register row has to evidenceAppliesISO/IEC 42001:2023
C01 Register the agent with a named ownerAn inventory entry stating purpose, intended use, the systems it touches, the model and platform behind it, and one accountable person by name. Also: Governance Baseline Q2.All tiers4.3 Scope; A.3.2 AI roles and responsibilities; A.9.4 Intended use
C02 Enumerate the action allowlistA dated list of every action the agent can take without approval, covering read, write, external communication and financial commitment. Anything not listed is denied. Also: Governance Baseline Q1; RA control B3.All tiersA.9.4 Intended use; A.4.4 Tooling resources; A.10.2 Allocating responsibilities
C03 Log every run end to endInputs, retrieved sources and their versions, each tool call, model and prompt version, approvals, and outputs, retained for the period your risk criteria set. Also: TRACE T; RA control E1.All tiersA.6.2.8 Recording of event logs; 9.1 Monitoring, measurement, analysis and evaluation
C04 Set acceptance criteria and test against them before releasePre-agreed, measurable criteria independent of the agent’s own output, with the evaluation results filed before go-live. Also: TRACE A.All tiersA.6.2.2 Requirements and specification; A.6.2.4 Verification and validation
C05 Control changes to model, prompt and policyModel, prompt and policy versions held in a register; every change is a release with an approver and a re-run of the acceptance tests. Also: RA control E3.All tiersA.6.2.7 Technical documentation; 8.1 Operational planning and control
C06 Tell users what it is and what it is forUsers and affected parties are told they are dealing with an AI system, what it is intended to do, and how to raise a concern.All tiersA.8.2 System documentation and information for users; A.3.3 Reporting of concerns
C07 Demonstrate the halt before go-liveA person has stopped the agent mid-task and the stop worked, with the date and the operator recorded. Repeat the exercise on a defined interval. Also: TRACE E; RA control E4.All tiersA.6.2.6 Operation and monitoring; A.8.4 Communication of incidents
C08 Put a human checkpoint before consequential actionsThe actions that wait for a person, and the thresholds that define them (amount, record type, recipient), are written down and enforced in code rather than in the prompt. Also: Governance Baseline Q2; RA control D2.Tier 2 and abovePartly or fully irreversible actionsA.9.2 Responsible use; A.3.2 AI roles and responsibilities
C09 Record an AI system impact assessmentA documented assessment of the effect on individuals, groups and society, reviewed when the scope or the affected population changes.Tier 2 and aboveCustomers, the public or third parties affectedPersonal data6.1.4 and 8.4 AI system impact assessment; A.5.2, A.5.3, A.5.4
C10 Write the incident procedure and name the suspension authorityWho is notified, what is reverted, how it is logged, who can suspend the agent, and what counts as a reportable event under the rules that apply. Also: Governance Baseline Q3.Tier 2 and aboveStatutory or licensed activityA.8.4 Communication of incidents; A.8.3 External reporting; 10.2 Nonconformity and corrective action
C11 Keep the manual fallback exercisedStaff can do the task by hand if the agent is stopped, suspended or changes behaviour after a model update, and have done so within a defined period. Also: Governance Baseline Q4.Tier 2 and aboveA.6.2.6 Operation and monitoring; A.4.6 Human resources
C12 Screen untrusted input and treat retrieved content as dataInbound content is classified for injection patterns and stripped of secrets; retrieved text can never widen the agent’s permissions. Also: RA controls A2, C2.Tier 2 and aboveReads or receives untrusted contentA.6.2.6 Operation and monitoring; 6.1.2 AI risk assessment
C13 Require a named approver per consequential action with the trace attachedEach consequential action waits for a specific person who sees the reasoning trace, and the approver queue is measured so review does not decay into rubber-stamping. Also: RA control D2.Tier 3 and aboveA.9.2 Responsible use; A.3.2 AI roles and responsibilities
C14 Enforce data access at retrieval and provenance on every sourceDocument-level access control applied at query time, every passage tagged with source, version and trust tier; access control itself under ISO/IEC 27001. Also: RA control C2.Tier 3 and abovePersonal or regulated dataA.7.5 Data provenance; A.7.4 Quality of data; A.7.2 Data for AI systems
C15 Allocate supplier and platform responsibilities in writingThe model provider, the agent platform and any integrator each have documented responsibilities for security, availability, change notice and incident reporting.Tier 3 and aboveStatutory or licensed activityA.10.2 Allocating responsibilities; A.10.3 Suppliers
C16 Bound delegation between agentsAn agent that instructs other agents passes on a narrower permission set than its own, and the full chain is visible in one log. Also: RA risk register, agent-to-agent escalation.By trigger onlyInstructs other agentsA.6.2.2 Requirements and specification; A.4.4 Tooling resources; A.6.2.8 Event logs
C17 Independent review before go-live and after material changeInternal audit or an independent function reviews the register row, the evidence behind it and the residual-risk acceptance before deployment and after any material change.Tier 3 and above9.2 Internal audit; 9.3 Management review; 6.1.3 AI risk treatment

Annex A references are to the control objectives and controls of ISO/IEC 42001:2023. Clause references without an "A." prefix are to the management system requirements. Access control, logging integrity and supplier security are ISO/IEC 27001 controls, referenced and not duplicated.

What each rule in the matrix rests on

The tool answers design questions with published practice rather than with my own preference. Where a rule is my calibration, the table says so.

Show the ten rules and their sources
RuleSource
Risk is assessed per AI system, with likelihood and consequence, and treated through documented controlsISO/IEC 42001:2023 clauses 6.1.2, 6.1.3, 8.2 and 8.3 Source →
An impact assessment on individuals, groups and society is recorded separately from the organisational risk assessmentISO/IEC 42001:2023 clauses 6.1.4 and 8.4; Annex A.5 Source →
Each control in the register is referenced to the Annex A control or clause it evidencesISO/IEC 42001:2023 Annex A and the statement of applicability (clause 6.1.3) Source →
Reversibility decides the human checkpoint: irreversible actions wait for a named personTRACE Framework v1.1, criterion R and the three reversibility tiers Source →
An agent that cannot be halted mid-task is not deployed, whatever its model performanceTRACE Framework v1.1, criterion E Source →
The four register fields an auditor asks for first: allowed actions, named reviewer, error procedure, manual fallbackFour-Question AI Governance Baseline v1.1 Source →
Excessive agency is contained by least-privilege tool permissions, a human checkpoint by tier and output guardrailsOWASP Top 10 for LLM Applications 2025, LLM06; OWASP Agentic AI Threats T2, T3, T10 Source →
Controls C01 to C17 correspond to the control groups of the Governed Agentic RAG reference architectureGoverned Agentic RAG reference architecture v1.0, control set and risk register Source →
Map, measure and manage AI risk across the life cycle, with monitoring after deploymentNIST AI Risk Management Framework 1.0, MAP, MEASURE and MANAGE functions Source →
The band thresholds (Assisted 0–2, Supervised 3–5, Bounded 6–7, Autonomous 8–9; Low 0–4, Moderate 5–8, High 9–12, Severe 13–15) and the tier matrixPractitioner calibration by Terence Kok from applying ISO/IEC 42001 to agent deployments. Not prescribed by the standard. Re-cut the bands to your own risk criteria under clause 6.1.1 if your appetite differs.

How to read the matrix and use the register

Two axes instead of one score

Purpose

Autonomy measures how much the agent does before a person sees it: action scope, where the checkpoint sits, and how far one run reaches. Consequence measures how bad a wrong action is: reversibility, data, blast radius, untrusted input and regulatory weight. The tier is read where the two meet.

Why It Matters
  • Two agents with the same single score can need opposite treatment. A high-autonomy, low-consequence summariser needs monitoring. A low-autonomy, high-consequence payment drafter needs a named approver
  • The cheapest way to move an agent out of Tier 4 is almost always along the autonomy axis: put the person earlier in the loop. Cutting consequence usually means changing the task
  • The matrix makes the trade explicit, which is what a risk committee needs to accept residual risk under clause 6.1.1
How to Use
  1. Rate the agent as it will run in production, including the review that will actually happen, not the review the design document promises
  2. If the tier surprises you, change one answer at a time and watch the marker move
  3. Record the answers you disagreed with in the register notes so the next review can test them

Where the register lands in the AIMS

Purpose

Each register row is a clause 6.1.2 risk assessment and, where personal data or the public are involved, a clause 6.1.4 impact assessment for one agent. The control set is the input to the risk treatment plan under 6.1.3. The ISO references against each control are what go into the statement of applicability.

Why It Matters
  • The first thing a certification auditor asks for is the list of AI systems in scope and the assessment behind each one. A register with a row per agent is that list
  • A register row with a named owner, a dated halt test and an approver satisfies three of the four Governance Baseline questions on its own
  • The standard does not distinguish agents from other AI systems. Your register has to, or the controls for a chatbot get copied onto an agent that moves money
How to Use
  1. Print the register and file it with the AI system inventory under clause 4.3
  2. Copy each control's ISO reference into the statement of applicability with "applied" or "not applied" and the reason
  3. Re-run the row after any model, prompt, scope or supplier change, and date the re-run

Holds that override the tier

Purpose

Three conditions hold deployment regardless of where the agent lands: nobody has stopped it mid-task and proved the stop works, no single person owns it, and a Tier 2 or higher agent has no named approver. The register lists them as holds rather than folding them into the score.

Why It Matters
  • A documented kill switch that has never been pulled is the most common finding I see on agent deployments, and the one that takes an afternoon to close
  • "The team owns it" is not an owner. The standard asks for roles and authorities assigned to people, and an auditor will ask that person to describe the agent
  • A hold is cheaper to clear before go-live than after the first incident report
How to Use
  1. Clear every hold before the row goes to the risk committee
  2. Schedule the halt exercise on the same cadence as the tier review
  3. If a Tier 4 agent is already live, treat the row as an open nonconformity under clause 10.2 and start the redesign
ISO/IEC 42001 Lead Auditor badge

About This Tool

This matrix was built by Terence Kok, a certified ISO/IEC 42001 Lead Auditor and AIGP, who has unified AI governance across three national regulatory environments, Singapore's IMDA among them, under a single standard. The control set follows the Governed Agentic RAG reference architecture, whose machine-readable form is in the public AI governance toolkit.

See the full certification list →

This matrix rates agents you have already decided to build. To test whether a task should be handed to an agent at all, run it through the TRACE Agent Evaluation first. To check whether the management system around the register would survive certification, use theISO 42001 AIMS Readiness Checklist. To test whether your board could answer for the agents in the register, use theBoard AI Oversight Checklist.

Taking the register through to a statement of applicability

The matrix tells you which tier each agent sits in. Agreeing the risk criteria, closing the holds and writing the treatment plan is work I do with clients directly.

Book a private session
Terence Kok
Before You Go

I built this after seeing the same register three times in one quarter: a chatbot template with the word agent written over the top. The two-axis matrix is the fix I use in my own work. It is not clever; it just refuses to let a summariser and a payment agent share a row. Load the worked example first. The public-sector agent in Tier 4 is there on purpose, because that is where I most often find agents already running.

Terence Kok