What decision does this method support?
Direct answer. Check AI-enabled work during operation against intended outcomes and agreed requirements, then use the evidence to decide whether the work should continue, expand, be restricted, or stop.
Executives face pressure to innovate faster and demonstrate a return on AI investment.
When agents can release payments, change records or make commitments to customers, weak controls can expose the organization to financial loss and harm the people it serves. Agents may repeat an error across many transactions before anyone intervenes. The organization remains accountable for the consequences. Boards and executives need assurance that expected business outcomes are being achieved, risks remain within agreed limits and controls work in practice.
Testing before agents go live is necessary, but it cannot show whether later work keeps meeting expectations. This paper sets out a method for continuous assurance: checking AI-enabled work during operation against intended outcomes and agreed requirements, and using the evidence to decide whether the work should continue, expand, be restricted or stop. It builds on existing management, risk and control responsibilities.
Leaders base those decisions on reports, prepared by management and, where required, by an independent reviewer. Reports should show the work and period covered, what was checked, what failed and what remains unchecked or unresolved. They should identify who assessed the work and the limitations of the conclusion. Management's own checks remain distinct from independent assurance.
The business case and budget should include the people needed to review and correct agent work, with the expertise, authority and time to challenge it and act on problems.
Start with one business process where agents are already in production. Ask the business process owner to answer the questions in Questions to ask management (C2) using recent results and supporting evidence. Use the review to decide whether to maintain, expand or restrict agent authority, and what needs to improve.
What does this brief claim, and what does it leave to you?
Direct answer. References to laws and standards show where the method can support an obligation or established practice. They do not establish compliance or conformance.
The method draws on regulation, standards and iTmethods' engineering and client experience. Where it draws on our own work, it describes what that work is designed to do.
This edition reflects regulations, standards and research available in September 2026. Check the current version of any instrument before relying on it.
Using laws and standards
References to laws and standards show where the method can support an obligation or established practice; they do not establish compliance or conformance. Mapping any of it to your own control framework is joint work with your compliance function and counsel, and whether evidence is sufficient for a purpose is a judgment for your own risk, compliance and audit experts.
The main paper covers the decisions and evidence leaders need. For an executive review, start with the executive summary, the five questions and Questions to ask management (C2). Appendix 1 provides control and operating guidance for delivery, risk and audit teams. Appendix 2 covers regulations and standards; the glossary and sources follow.
How do you apply existing controls to the work agents do?
Direct answer. Apply the controls the organization already uses, and extend them where agents change how that work is done.
Apply existing controls to the work agents do, and extend them where agents change how that work is done.
Running example: refunds and fee corrections
At a retail financial-services firm, AI agents handle about 6,000 refund and fee-correction requests a week. They triage each request, decide eligibility, issue refunds of up to $250 on their own, update the customer's record and reply. Larger refunds go to four approvers. The firm and its figures are fictional.
What must oversight cover before agents go live, and after?
Direct answer. Testing before deployment shows how the system behaved under the conditions tested. Oversight during operation examines the work performed since, including whether the outcome was achieved and whether the controls worked.
Management defines the intended outcomes, the risk limits that put the board's risk appetite into practice, and accountability for each AI-enabled process. The people responsible need evidence that the work meets those expectations and authority to respond when it does not.
| What needs attention | Management response |
|---|---|
| Outputs and chosen actions can vary from case to case | Test representative scenarios, and check whether the business result meets agreed requirements. |
| Behavior can change when models, instructions, tools, data or suppliers change | Identify material changes and reassess the affected controls and permitted scope. |
| Agents can take many consequential actions between reviews | Limit cumulative exposure, and define when systems must hold or restrict work. |
| An agent's self-assessment should not be the only check | Check the work against rules, source records or other evidence. |
| As agents take the routine cases, reviewers may receive more ambiguous, disputed or consequential cases | Plan and budget the people, skills and time to review, correct and escalate that work. |
Testing before deployment shows how the system behaved under the conditions tested. Oversight during operation examines the work performed since, including whether the outcome was achieved and whether the controls worked.
The plan-do-check-act cycle in Figure 2 connects these responsibilities before and after agents go live.
Plan
Set expectations
Define the intended outcomes, limits on agent authority, accountable owners and the resources and controls needed.
Do
Put the plan into practice
Build, test and operate the process within its approved scope, with the agreed controls and human oversight in place.
Check
Review the evidence
Assess business results, costs and control performance against expectations. Identify what failed, what was not checked and what remains unresolved.
Act
Decide what changes
Correct problems and decide whether to maintain, expand, restrict or stop the work. Update the controls, resources or permitted scope where needed.
Use findings to update the plan
Checks run during operation as well as at scheduled reviews. Findings may require immediate intervention or changes to the wider plan. Appendix 1.1 explains how this management cycle connects to the detailed workflow and the frameworks used in this paper.
Which control domains should you extend for agents?
Direct answer. For each AI-enabled process, establish who owns the outcome, what agents may do without approval, and which decisions remain with people.
Start with the controls the organization already uses. For each AI-enabled process, establish who owns the outcome, what agents may do without approval and which decisions remain with people.
The twelve control domains show where those controls may need extending, from planning and change management to oversight of live work. Use them to identify gaps in responsibility, control and evidence. The detailed reference is in Appendix 1.2.
Running example
The refund process has a business process owner, the head of customer operations, and a record of the three agents involved and the systems each can reach. Its specification defines the outcome: the right amount refunded, the customer's record corrected and an accurate reply.
The $250 limit on refunds an agent may issue alone is one expression of the firm's risk appetite. The process owner also weighs cumulative exposure, customer impact and how quickly errors can be found and corrected.
SETS DIRECTION
Governance and risk appetite
board and executives · how much an agent may do, at what consequence, before a person decides
THE LIFE OF ONE WORKFLOW (the workflow loop of Figure 8)
Planning and design
business process owner
Development
engineering
Verification and validation
risk, engineering
Deployment and change management
change board, IT
Operations and monitoring
operations, risk
Assurance
management checks; risk and compliance oversight; independent internal audit
Incident and problem management
incident owner, risk
what operation found, back to the plan
FOUNDATIONS EVERY STEP DRAWS ON
Training and competence
everyone who approves, reviews or owns
Security, identity and observability
IT, security
Data governance
data owners, compliance
Third-party risk management
procurement, vendor risk
- Overseeing agents at work
- feeds the next plan
How do you oversee agents while the work is running?
Direct answer. During operation, business leaders need evidence that the work remains within the agreed requirements, and a defined response when it does not.
During operation, business leaders need evidence that the work remains within the agreed requirements, and a defined response when it does not.
Why do people still have to be in the loop?
Direct answer. As routine cases are automated, reviewers may receive a higher proportion of ambiguous, disputed, or consequential cases. Plan capacity for that workload.
Agents change the work people must do. As routine cases are automated, reviewers may receive a higher proportion of ambiguous, disputed or consequential cases. Plan capacity for that workload, rather than assuming the old mix of cases still applies.
Plan and budget for the people who review, correct and escalate the work as part of the business case. Give reviewers the authority to refuse, the evidence and expertise to judge, and enough time to act. Training, breaks and rotation should reflect the demands of the work.
Set limits on how much work can await review and for how long. Agree what happens when those limits are reached: add capacity, narrow what agents may do, slow the flow or escalate to the business process owner.
Track review delays, overrides, rework and recurring problems alongside speed and cost. To assess whether oversight works in practice, ask reviewers to show a recent case they refused, changed or escalated, and why.
Running example
Four approvers can handle about 300 cases a day, with an agreed queue limit of 600. A billing error increases cases needing approval to 500 a day. As the queue reaches its limit, customers receive a holding reply and the process owner brings in the extra approvers the plan provides for peaks, within the budget in the business case. The process owner then checks whether the added capacity cleared the queue and whether effort and cost stayed within plan.
Appendix 1.3 sets out roles, workload, skills and tools, review limits and measures in detail.
How do you monitor outcomes and act when work goes wrong?
Direct answer. Define the intended outcome, operating limits, and the response to problems before the work begins. When outcomes deteriorate, authorized people decide whether the work should continue, be restricted, or stop.
Define the intended outcome, operating limits and response to problems before the work begins. Monitor the business results as well as the tasks completed.
When outcomes deteriorate, controls fail or material conditions change, use the agreed response. Systems enforce required holds; authorized people resolve exceptions and decide whether the work should continue, be restricted or stop.
Define outcomes
before it runs
what good means, written down
Monitor outcomes
results and tasks done
the process owner checks outcomes
Adjust
continue, restrict or stop
when outcomes or conditions change
a person decides: continue · restrict: a narrower scope, then establish again · stop → Fig. 10
a significant change reopens the decision
model, prompt, tool or process change; a vendor's notice; a new condition in operation
Match the response to the potential harm, how quickly it could spread, whether it can be reversed, applicable obligations and the consequences of interrupting the service. A problem in one part of a process may justify restricting that part rather than stopping everything.
If a required pre-action check fails, is inconclusive or returns no result, hold the action for review. If the checking service is unavailable, pause the work or use an approved fallback where the risk permits. Record the evidence gap; an outage is not a passed check.
Problems found after an action require correction and a decision about subsequent work. After a stop, reconcile unfinished work, correct affected records and address the cause before the authorized person approves restart. Preserve a route for affected people to challenge the outcome.
Running example
After a model update, the agent that replies to customers starts promising refunds outside the fee policy. Checks catch the problem. The business process owner restricts the agent's work: staff take over refund-eligibility questions, and the agent keeps handling other replies. Once the fix passes the agreed checks, the process owner returns those questions to the agent.
Appendix 1.5 lists the impact factors and response options, and explains fallback and restart.
Which controls have to hold in live operation?
Direct answer. A written instruction alone does not prevent an action that the available tools still permit. Authorization does not establish that the work was accurate, complete, or achieved its intended outcome.
Enforce limits through the systems agents use. Give each agent identity a named owner and only the access it needs, and test whether the access paths preserve the separation of duties the process requires. A written instruction alone does not prevent an action that the available tools still permit.
Approve and record material changes to models, instructions, tools, data and permissions. In the proposed design, where a consequential action requires approval, the approval identifies that action and the recorded scope it permits. Reassess it when relevant conditions change.
Then check the business result. Authorization does not establish that the work was accurate, complete or achieved its intended outcome. Appendix 1.4 shows how these controls apply to a refund requiring human approval.
What should assurance evidence show a leader?
Direct answer. Leaders need a supported assessment of whether agent work and its controls met agreed criteria within a defined scope and period, so they can decide how much authority agents should hold.
Why it matters for agents
Assurance gives leaders a supported assessment of whether agent work and its controls met agreed criteria within a defined scope and period, so they can decide with confidence how much authority agents should hold. It requires competent, objective reviewers, an agreed assessment process and systems that capture reliable evidence as the work happens.
Three features of agent-enabled work make this important:
- They change often. Models are updated, instructions and tools change and workflows are recomposed, sometimes by a supplier. Testing before an agent goes live shows how it behaved then, not how it behaves now.
- They can be wrong in convincing ways. Agents built on large language models produce likely answers rather than guaranteed ones. They can carry bias, which NIST's AI Risk Management Framework treats as a risk to manage, and in one study, five widely used AI assistants tended to favor a user's stated views over the correct answer.
- They cannot be relied on to check themselves. In one study, language models struggled to correct their own reasoning without outside feedback. Check the work against rules, source records or other evidence instead. A separate checking component helps, but does not by itself make the review independent.
Software delivery relies on testing and quality assurance before release. Agents need that discipline carried into live operation: checks on the work itself, as it runs and periodically, matched to the risk (Figure 5). Figure 6 gives examples.
Sycophancy: Sharma et al., Towards Understanding Sycophancy in Language Models, arXiv:2310.13548 (2023). Self-correction: Huang et al., Large Language Models Cannot Self-Correct Reasoning Yet, ICLR 2024, arXiv:2310.01798.
Build
Test before release
automated tests (evals), QA, red-teaming
Release
Operate
checks on the work as it runs
sampled review
outcome monitoring
Change
re-test and re-check after model, prompt, tool or supplier changes
↺ back to Operate
shaded: the testing discipline carried into live operation
| Check | What it does | Refund example |
|---|---|---|
| 1. Input check | Confirms the request and data are complete and valid before the agent acts | The request names a real transaction on the customer's account |
| 2. Rule and limit check | Tests the action against policy and authority limits | The refund is within $250 and the fee policy allows it |
| 3. Reconciliation | Compares the result with the system of record | The refund posted matches the approved amount and the right account |
| 4. Recalculation | Redoes the calculation separately | The amount is recomputed from the fee schedule |
| 5. Separate review | A separate rule set or model reviews the output | A second check reads each reply for promises the policy does not allow |
| 6. Sampled human review | A person reviews a sample, weighted to risk | A reviewer examines a weekly sample, weighted to new or unusual cases |
| 7. Outcome monitoring | Tracks results over time to spot drift | Reversals, repeat complaints and refund totals by week |
| 8. Re-check after change | Re-runs checks after a model, instruction, tool or supplier change | The reply checks re-run on recent cases after a model update |
Appendix 1 sets out the controls for each control domain.
Before relying on an assurance conclusion, ask what work and period it covers, what it was measured against, what was checked and what remains unresolved. Management's own checks support that assessment; independent assurance needs a reviewer who is independent of the work. Appendix 1.7 describes assurance more fully, including who provides it and how independence is judged.
What the evidence should show
Keep enough evidence to establish what work was done, what requirements applied, what was checked and what each result meant. Show the population in scope, the coverage achieved and any gaps. Preserve the history needed to understand earlier decisions under the organization's retention requirements.
Keep failed, inconclusive and unassessed work separate: they require different follow-up. Give unresolved issues an owner and a next action.
The checks also need scrutiny. Review their test results, false alarms, missed errors, known limitations and failure modes before relying on their findings.
Running example
Last week the process issued 4,800 refunds (Figure 7). A report showing only the 4,655 that passed would hide the rest. The process owner also needs to know the potential impact of unresolved cases and whether corrective actions are overdue.
4,800 refunds issued in the reporting week
4,655
passed the checks performed
no exception identified
3
failed
investigate
12
inconclusive
look again
130
not yet checked at the reporting cutoff
reconcile
For the records to keep and how to count coverage, see Appendix 1.6.
Which requirements apply, and what should you ask management?
Direct answer. Each organization designs its own controls for AI-enabled work, within the expectations that regulators and standards bodies set.
Each organization designs its own controls for AI-enabled work, within the expectations that regulators and standards bodies set. Part C sets out which of them apply to agent work, then gives leaders the questions to ask management about whether the controls exist and work.
Which requirements and guidance apply?
Direct answer. First determine, with risk, compliance, and legal teams, which requirements apply to the process. Recurring expectations are accountability, defined authority and limits, controlled changes, human oversight where required, monitoring, evidence, and a response when things go wrong.
In the frameworks reviewed here, the recurring operating expectations are similar: clear accountability, defined authority and limits, controlled changes, human oversight where required, monitoring of live operation, evidence of what happened, and a response when things go wrong.
First determine, with risk, compliance and legal teams, which requirements apply to the AI-enabled process, given the organization's role and jurisdiction, the system's intended use and classification, and application dates.
They come in four kinds. Binding legal and policy requirements, such as the EU AI Act and, for Canadian federal institutions, the Directive on Automated Decision-Making, apply by role, system and date. Supervisory guidance, such as OSFI's E-23, PRA SS1/23 and the US model-risk guidance, applies to regulated firms within its scope; SR 26-2 expressly leaves generative and agentic AI outside its scope. OSFI's July 2026 technology risk bulletin sets out sound practices institutions can consider. Voluntary frameworks, such as NIST AI RMF and ISO/IEC 42001, name functions and management-system requirements. Professional guidance, such as the IIA's, sets expectations for those who review the work.
The detailed table is in Appendix 2. It is a scoping aid, not a conformance assessment: each entry states what the regulation, standard or guidance contributes and what it leaves to you.
What should executives ask management to show?
Direct answer. Use these questions before agents are given more authority and at periodic reviews. Ask management to answer for a specific process, using recent work and supporting evidence.
Use these questions before agents are given more authority and at periodic reviews. Ask management to answer for a specific process, using recent work and supporting evidence.
| Executive question | Ask management to show |
|---|---|
| 1. Is the process delivering the intended outcome at an acceptable total cost? | Results and actual total cost against the business case, including technology and assurance costs, planned staffing, actual review and correction effort, and material variances. |
| 2. Who owns the result, and what may agents do without approval? | The business process owner, the agents and systems involved, enforced limits and decisions reserved for people. |
| 3. Are access and changes controlled? | Permissions in use and recent changes to models, instructions, tools or providers, with their impact assessed. |
| 4. Can people review exceptions competently and in time? | Capacity, training, queue age and examples of decisions reviewers refused, changed or escalated. |
| 5. When must we restrict or stop the work, and how do we recover? | Agreed triggers, decision authority, fallback and recovery arrangements, and routes for affected people to challenge outcomes. |
| 6. What is wrong, uncertain or unchecked, and what risk remains? | Coverage for a stated period, separate result categories, unresolved issues and their potential impact, owners and action dates, and any residual risk accepted by the business process owner within their authority. |
| 7. Who has challenged whether the controls and checks work? | The reviewer's scope, findings, limitations and independence from the work assessed. |
What should you do with one process already in production?
Direct answer. Choose one process where agents already do consequential work, and ask its business process owner to answer these questions.
Choose one process where agents already do consequential work, and ask its business process owner to answer these questions.
Decide whether the intended outcome is being achieved, whether review capacity matches demand and whether the current level of agent authority remains appropriate. Give each material gap an owner, an action and a date.
Use those findings, together with the business case, to decide what should continue, change or stop, and whether wider use is justified.
Checking agent work in live operation, and keeping the evidence that leaders and reviewers need, is the area iTmethods is focused on.
If you are running agents in live operation, we would welcome a conversation.
iTmethods, September 2026.