Is this task ready for an AI agent?

Evaluate one candidate task against the five TRACE criteria (Traceability, Reversibility, Acceptance Criteria, Compliance, and Escalation) before any vendor selection or architecture decision begins.

A hand badging through a secure turnstile gate, representing an evaluation checkpoint before autonomous AI deployment

Listen to this briefing

The TRACE AI Pre-Flight Checklist

0:00
Task
TTraceability
RReversibility
AAcceptance
CCompliance
EEscalation
Result

Name the task you are evaluating.

TRACE evaluates one specific task at a time. Be precise: "process customer refunds" is too broad; "validate and approve refund requests under $500 in the CRM system" is evaluable.

Be specific enough that someone unfamiliar with your operations could understand what this agent would do.

T

Traceability of Inputs and Outputs

Every input consumed and output produced must be logged, timestamped, and attributed to specific data sources. Tasks drawing on unverified sources should be deprioritised until provenance controls are in place.

Can every input to this task (data, triggers, parameters) be logged, timestamped, and attributed to a specific source?

Can every output this task produces be recorded, attributed, and audited against its source inputs?

Are data provenance controls currently in place for the sources this task relies on?

R

Reversibility and Risk Tier

Tasks are classified by the consequence of an erroneous output: reversible with no harm, reversible with cost, or irreversible. Irreversible-consequence tasks require human-in-the-loop checkpoints before any agent autonomy is granted.

If this agent produces an erroneous output on this task, what is the worst-case consequence?

A

Acceptance Criteria Definition

Success criteria must be measurable and pre-agreed independent of the agent's output. A task lacking an external ground truth or validation method is not ready for deployment regardless of technical feasibility.

Does this task have measurable, pre-agreed success criteria that exist independently of what the agent outputs?

Is there an external ground truth or independent validation method that can confirm whether the agent's output is correct?

C

Compliance and Regulatory Mapping

Applicable regulatory regimes, sector codes, and governance policies must be identified before architecture decisions. This is where a task's exposure to regimes such as the EU AI Act, GDPR, and the NIST AI Risk Management Framework gets scoped. For government and regulated environments, this mapping is a prerequisite to scoping.

Have the regulatory regimes (for example the EU AI Act, GDPR, or NIST AI RMF), sector codes, and data protection requirements applicable to this task been identified?

Have applicable internal governance policies and approval processes been mapped before any agent architecture decisions?

E

Escalation and Override Pathway

A clear, tested mechanism must exist for a human operator to intervene, override, or halt the agent mid-task. Absence of this pathway disqualifies a task from autonomous execution regardless of model performance.

Does a clear, documented mechanism exist for a human operator to intervene in, override, or halt this agent mid-task?

Has this escalation and override pathway been tested in a realistic scenario?

T
R
A
C
E

T

Traceability

R

Reversibility

A

Acceptance Criteria

C

Compliance

E

Escalation

Read the full framework →

This is a self-assessment against the TRACE framework based on the answers you gave, not a certification, an accredited audit, or legal advice. The guidance above is a starting point for your own governance review, not a substitute for it. Your answers stay in your browser and are not sent to Terence Kok or reviewed by anyone.

What each criterion tests and why it comes before architecture.

TRACE is an original framework, not an implementation of an external standard. Its five letters loosely map onto the kind of instrumentation OpenTelemetry's GenAI semantic conventions now define for agent traces (orchestration, tool-call, model, and memory spans), but the mapping is conceptual, not a formal spec conformance claim.

A hand following a heavy metal chain link by link, representing a traceable audit trail of dataT

Traceability of Inputs and Outputs

Without full logging of what data entered the task and what output was produced, you cannot audit errors, satisfy regulators, or improve the system. Traceability is the precondition for accountability, not an implementation detail.

When this fails: Any erroneous output becomes unattributable. You cannot tell what caused it, cannot correct the root issue, and cannot demonstrate due diligence to auditors or affected parties.

A hand hovering between a lit and an unlit toggle switch, representing reversible versus irreversible riskR

Reversibility and Risk Tier

The risk classification determines how much oversight is required before granting autonomy. Low-reversibility tasks demand human review before any consequence propagates. Committing to autonomous execution on an irreversible task without a human checkpoint is not a risk tolerance decision. It is a governance gap.

When this fails: Irreversible errors occur without human review. A single mistaken output in a high-stakes domain can generate costs or harms that no amount of post-incident process can fully reverse.

A hand holding a precision caliper measuring a machined part, representing measurable acceptance criteriaA

Acceptance Criteria Definition

An agent that grades its own output is not evaluated. It is trusted. Acceptance criteria must exist independently of the agent, defined before deployment begins. Without an external ground truth, you have no basis for judging whether the agent is performing correctly, improving, or degrading.

When this fails: You cannot distinguish a performing agent from a malfunctioning one. The first indication of systematic errors is typically a downstream incident, not a monitoring alert.

A hand pressing a wax seal stamp onto a document, representing regulatory compliance approvalC

Compliance and Regulatory Mapping

Discovering applicable regulations after architecture decisions have been made forces redesign, delays, or deployment with known compliance gaps. The mapping must precede design, not follow it. For most agent tasks this means checking exposure to the EU AI Act, GDPR, and the NIST AI Risk Management Framework at minimum, alongside any sector code. For cross-jurisdiction deployments, identify the most restrictive regime first and design to it.

When this fails: The deployment violates a regulatory requirement that was not identified until after go-live. Remediation requires suspension, redesign, and regulatory engagement that could have been avoided with a pre-deployment compliance inventory.

A hand pulling an emergency-stop style lever on a wall panel, representing a human escalation pathwayE

Escalation and Override Pathway

A tested override mechanism is the non-negotiable minimum for any autonomous agent deployment. "Tested" is the operative word: a pathway that exists only on paper has unknown failure modes. The escalation pathway is what converts autonomous execution into accountable execution.

When this fails: When an agent behaves unexpectedly, no tested mechanism exists to halt it. The time spent constructing a manual stop while the system continues to run is the cost of this gap.

How to interpret the three outcomes.

Three rubber stamps beside green, amber, and red ink pads, representing pass, conditional, and fail outcomes
Deployment Candidate

All five criteria pass. This task is structurally appropriate for autonomous agent deployment. Proceed to architecture design and vendor selection, keeping the first deployment to no more than three to five tasks under close monitoring.

Conditional

One or more criteria are partially met. The task may proceed to scoping, but the amber criteria must be resolved before deployment begins. Identify whether the gap can be closed in 30 days; if not, treat the task as not yet ready.

Not Ready

Two or more criteria fail, or the Reversibility criterion classifies the task as irreversible without a human-in-the-loop checkpoint. The task should remain manual or assisted until the failing criteria are remediated. Deploying before remediation is not a calculated risk. It is an unmanaged one.

The full TRACE methodology is in the article.

The article behind this tool sets out the full TRACE methodology, a leadership checklist for scoping your first task cohort, and why the task inventory belongs in front of leadership before procurement begins.

Once a task passes TRACE and an agent is designed for it, rate that agent on theAgent Risk Assessment Matrix to get its oversight tier and the ISO/IEC 42001 control set for the register.

Read the framework →Consulting enquiry
Terence Kok
Before You Go

TRACE came out of watching teams jump straight to picking a vendor before anyone had asked whether the task itself was even fit for an agent. Reversibility is the criterion that trips people up most: it forces you to name the worst case before you've built anything. A conditional or not-ready result isn't a rejection. It's the cheapest information you'll get on this task all year.

Terence Kok