Agent Risk Assessment Matrix
Eight questions place each AI agent on a four-by-four matrix of autonomy against consequence. The result is an oversight tier, the controls that tier requires, and the ISO/IEC 42001:2023 clause each control satisfies, printed as a register an auditor can read. Load the worked example to see five anonymised agents already filled in, then replace them with your own. No email required.
Listen to this briefing
The Agent Risk Assessment Matrix
Why agents need their own risk register
ISO/IEC 42001 asks for an AI risk assessment and an impact assessment for every system in scope. It does not say how to rate an agent that reads supplier emails, plans a sequence of actions across three systems, and posts a journal entry without asking. Most registers I am shown rate that agent the way they rate a chatbot, because the template was built for a model that answers questions rather than one that takes actions.
The fix is to separate two things the standard bundles into "risk": how much the agent does on its own, and how bad a wrong action is. Rate them on two axes and the oversight decision becomes visible, defensible and repeatable across every agent you register.
The clause references are to the published standard. The band thresholds and the matrix are my calibration from applying the standard to agent deployments; the basis table further down shows what each rule rests on.
Your agent risk register
Agents you add appear on the matrix and in the register below. Everything stays in this browser; nothing is sent to me.
The outlined marker is the agent in the form. Numbered markers are agents in the register.
| # | Agent | Autonomy | Consequence | Tier | Controls | Holds | Actions |
|---|---|---|---|---|---|---|---|
| No agents yet. Add one above, or load the worked example. | |||||||
The PDF opens a register with one page per agent: the eight answers, the tier, every required control with its ISO/IEC 42001 reference, and a sign-off block for the owner, the approver and the residual-risk acceptance. The worked example is an illustrative composite; none of the five agents is a client record.
This is a self-assessment against your own description of each agent, not an independent risk assessment, an audit or a legal opinion. The tier and the control set are a starting point for your own risk treatment plan and statement of applicability, not a substitute for them. Your entries stay in your browser and are not sent to Terence Kok or reviewed by anyone.
The seventeen controls and where ISO/IEC 42001 asks for them
Seven controls apply to every agent, whatever its tier. The rest are pulled in by the tier, or by a single answer that makes them necessary at any tier. The register names the reason against each control.
Show all seventeen controls with their ISO/IEC 42001 references ▸
| Control | What the register row has to evidence | Applies | ISO/IEC 42001:2023 |
|---|---|---|---|
| C01 Register the agent with a named owner | An inventory entry stating purpose, intended use, the systems it touches, the model and platform behind it, and one accountable person by name. Also: Governance Baseline Q2. | All tiers | 4.3 Scope; A.3.2 AI roles and responsibilities; A.9.4 Intended use |
| C02 Enumerate the action allowlist | A dated list of every action the agent can take without approval, covering read, write, external communication and financial commitment. Anything not listed is denied. Also: Governance Baseline Q1; RA control B3. | All tiers | A.9.4 Intended use; A.4.4 Tooling resources; A.10.2 Allocating responsibilities |
| C03 Log every run end to end | Inputs, retrieved sources and their versions, each tool call, model and prompt version, approvals, and outputs, retained for the period your risk criteria set. Also: TRACE T; RA control E1. | All tiers | A.6.2.8 Recording of event logs; 9.1 Monitoring, measurement, analysis and evaluation |
| C04 Set acceptance criteria and test against them before release | Pre-agreed, measurable criteria independent of the agent’s own output, with the evaluation results filed before go-live. Also: TRACE A. | All tiers | A.6.2.2 Requirements and specification; A.6.2.4 Verification and validation |
| C05 Control changes to model, prompt and policy | Model, prompt and policy versions held in a register; every change is a release with an approver and a re-run of the acceptance tests. Also: RA control E3. | All tiers | A.6.2.7 Technical documentation; 8.1 Operational planning and control |
| C06 Tell users what it is and what it is for | Users and affected parties are told they are dealing with an AI system, what it is intended to do, and how to raise a concern. | All tiers | A.8.2 System documentation and information for users; A.3.3 Reporting of concerns |
| C07 Demonstrate the halt before go-live | A person has stopped the agent mid-task and the stop worked, with the date and the operator recorded. Repeat the exercise on a defined interval. Also: TRACE E; RA control E4. | All tiers | A.6.2.6 Operation and monitoring; A.8.4 Communication of incidents |
| C08 Put a human checkpoint before consequential actions | The actions that wait for a person, and the thresholds that define them (amount, record type, recipient), are written down and enforced in code rather than in the prompt. Also: Governance Baseline Q2; RA control D2. | Tier 2 and abovePartly or fully irreversible actions | A.9.2 Responsible use; A.3.2 AI roles and responsibilities |
| C09 Record an AI system impact assessment | A documented assessment of the effect on individuals, groups and society, reviewed when the scope or the affected population changes. | Tier 2 and aboveCustomers, the public or third parties affectedPersonal data | 6.1.4 and 8.4 AI system impact assessment; A.5.2, A.5.3, A.5.4 |
| C10 Write the incident procedure and name the suspension authority | Who is notified, what is reverted, how it is logged, who can suspend the agent, and what counts as a reportable event under the rules that apply. Also: Governance Baseline Q3. | Tier 2 and aboveStatutory or licensed activity | A.8.4 Communication of incidents; A.8.3 External reporting; 10.2 Nonconformity and corrective action |
| C11 Keep the manual fallback exercised | Staff can do the task by hand if the agent is stopped, suspended or changes behaviour after a model update, and have done so within a defined period. Also: Governance Baseline Q4. | Tier 2 and above | A.6.2.6 Operation and monitoring; A.4.6 Human resources |
| C12 Screen untrusted input and treat retrieved content as data | Inbound content is classified for injection patterns and stripped of secrets; retrieved text can never widen the agent’s permissions. Also: RA controls A2, C2. | Tier 2 and aboveReads or receives untrusted content | A.6.2.6 Operation and monitoring; 6.1.2 AI risk assessment |
| C13 Require a named approver per consequential action with the trace attached | Each consequential action waits for a specific person who sees the reasoning trace, and the approver queue is measured so review does not decay into rubber-stamping. Also: RA control D2. | Tier 3 and above | A.9.2 Responsible use; A.3.2 AI roles and responsibilities |
| C14 Enforce data access at retrieval and provenance on every source | Document-level access control applied at query time, every passage tagged with source, version and trust tier; access control itself under ISO/IEC 27001. Also: RA control C2. | Tier 3 and abovePersonal or regulated data | A.7.5 Data provenance; A.7.4 Quality of data; A.7.2 Data for AI systems |
| C15 Allocate supplier and platform responsibilities in writing | The model provider, the agent platform and any integrator each have documented responsibilities for security, availability, change notice and incident reporting. | Tier 3 and aboveStatutory or licensed activity | A.10.2 Allocating responsibilities; A.10.3 Suppliers |
| C16 Bound delegation between agents | An agent that instructs other agents passes on a narrower permission set than its own, and the full chain is visible in one log. Also: RA risk register, agent-to-agent escalation. | By trigger onlyInstructs other agents | A.6.2.2 Requirements and specification; A.4.4 Tooling resources; A.6.2.8 Event logs |
| C17 Independent review before go-live and after material change | Internal audit or an independent function reviews the register row, the evidence behind it and the residual-risk acceptance before deployment and after any material change. | Tier 3 and above | 9.2 Internal audit; 9.3 Management review; 6.1.3 AI risk treatment |
Annex A references are to the control objectives and controls of ISO/IEC 42001:2023. Clause references without an "A." prefix are to the management system requirements. Access control, logging integrity and supplier security are ISO/IEC 27001 controls, referenced and not duplicated.
What each rule in the matrix rests on
The tool answers design questions with published practice rather than with my own preference. Where a rule is my calibration, the table says so.
Show the ten rules and their sources ▸
| Rule | Source |
|---|---|
| Risk is assessed per AI system, with likelihood and consequence, and treated through documented controls | ISO/IEC 42001:2023 clauses 6.1.2, 6.1.3, 8.2 and 8.3 Source → |
| An impact assessment on individuals, groups and society is recorded separately from the organisational risk assessment | ISO/IEC 42001:2023 clauses 6.1.4 and 8.4; Annex A.5 Source → |
| Each control in the register is referenced to the Annex A control or clause it evidences | ISO/IEC 42001:2023 Annex A and the statement of applicability (clause 6.1.3) Source → |
| Reversibility decides the human checkpoint: irreversible actions wait for a named person | TRACE Framework v1.1, criterion R and the three reversibility tiers Source → |
| An agent that cannot be halted mid-task is not deployed, whatever its model performance | TRACE Framework v1.1, criterion E Source → |
| The four register fields an auditor asks for first: allowed actions, named reviewer, error procedure, manual fallback | Four-Question AI Governance Baseline v1.1 Source → |
| Excessive agency is contained by least-privilege tool permissions, a human checkpoint by tier and output guardrails | OWASP Top 10 for LLM Applications 2025, LLM06; OWASP Agentic AI Threats T2, T3, T10 Source → |
| Controls C01 to C17 correspond to the control groups of the Governed Agentic RAG reference architecture | Governed Agentic RAG reference architecture v1.0, control set and risk register Source → |
| Map, measure and manage AI risk across the life cycle, with monitoring after deployment | NIST AI Risk Management Framework 1.0, MAP, MEASURE and MANAGE functions Source → |
| The band thresholds (Assisted 0–2, Supervised 3–5, Bounded 6–7, Autonomous 8–9; Low 0–4, Moderate 5–8, High 9–12, Severe 13–15) and the tier matrix | Practitioner calibration by Terence Kok from applying ISO/IEC 42001 to agent deployments. Not prescribed by the standard. Re-cut the bands to your own risk criteria under clause 6.1.1 if your appetite differs. |
How to read the matrix and use the register
Two axes instead of one score
Autonomy measures how much the agent does before a person sees it: action scope, where the checkpoint sits, and how far one run reaches. Consequence measures how bad a wrong action is: reversibility, data, blast radius, untrusted input and regulatory weight. The tier is read where the two meet.
- Two agents with the same single score can need opposite treatment. A high-autonomy, low-consequence summariser needs monitoring. A low-autonomy, high-consequence payment drafter needs a named approver
- The cheapest way to move an agent out of Tier 4 is almost always along the autonomy axis: put the person earlier in the loop. Cutting consequence usually means changing the task
- The matrix makes the trade explicit, which is what a risk committee needs to accept residual risk under clause 6.1.1
- Rate the agent as it will run in production, including the review that will actually happen, not the review the design document promises
- If the tier surprises you, change one answer at a time and watch the marker move
- Record the answers you disagreed with in the register notes so the next review can test them
Where the register lands in the AIMS
Each register row is a clause 6.1.2 risk assessment and, where personal data or the public are involved, a clause 6.1.4 impact assessment for one agent. The control set is the input to the risk treatment plan under 6.1.3. The ISO references against each control are what go into the statement of applicability.
- The first thing a certification auditor asks for is the list of AI systems in scope and the assessment behind each one. A register with a row per agent is that list
- A register row with a named owner, a dated halt test and an approver satisfies three of the four Governance Baseline questions on its own
- The standard does not distinguish agents from other AI systems. Your register has to, or the controls for a chatbot get copied onto an agent that moves money
- Print the register and file it with the AI system inventory under clause 4.3
- Copy each control's ISO reference into the statement of applicability with "applied" or "not applied" and the reason
- Re-run the row after any model, prompt, scope or supplier change, and date the re-run
Holds that override the tier
Three conditions hold deployment regardless of where the agent lands: nobody has stopped it mid-task and proved the stop works, no single person owns it, and a Tier 2 or higher agent has no named approver. The register lists them as holds rather than folding them into the score.
- A documented kill switch that has never been pulled is the most common finding I see on agent deployments, and the one that takes an afternoon to close
- "The team owns it" is not an owner. The standard asks for roles and authorities assigned to people, and an auditor will ask that person to describe the agent
- A hold is cheaper to clear before go-live than after the first incident report
- Clear every hold before the row goes to the risk committee
- Schedule the halt exercise on the same cadence as the tier review
- If a Tier 4 agent is already live, treat the row as an open nonconformity under clause 10.2 and start the redesign
About This Tool
This matrix was built by Terence Kok, a certified ISO/IEC 42001 Lead Auditor and AIGP, who has unified AI governance across three national regulatory environments, Singapore's IMDA among them, under a single standard. The control set follows the Governed Agentic RAG reference architecture, whose machine-readable form is in the public AI governance toolkit.
See the full certification list →This matrix rates agents you have already decided to build. To test whether a task should be handed to an agent at all, run it through the TRACE Agent Evaluation first. To check whether the management system around the register would survive certification, use theISO 42001 AIMS Readiness Checklist. To test whether your board could answer for the agents in the register, use theBoard AI Oversight Checklist.
Taking the register through to a statement of applicability
The matrix tells you which tier each agent sits in. Agreeing the risk criteria, closing the holds and writing the treatment plan is work I do with clients directly.
Book a private session
I built this after seeing the same register three times in one quarter: a chatbot template with the word agent written over the top. The two-axis matrix is the fix I use in my own work. It is not clever; it just refuses to let a summariser and a payment agent share a row. Load the worked example first. The public-sector agent in Tier 4 is there on purpose, because that is where I most often find agents already running.
