All field notesWorkflow Ownership

Put AI Above the Workflow, Not Inside Every Control

A practical operating model for letting AI interpret work while tested systems keep permissions, approvals, execution and auditability under control.

When I worked on ERP implementations, the safest part of the system was rarely the person clicking the button.

The controls sat underneath: approval limits, required fields, user roles, posting rules, transaction logs and exception queues. People could use judgment, but the system still decided what they were allowed to change and what evidence had to be kept.

We need the same design discipline for AI agents.

An agent is useful because it can interpret a messy request, pull together context and handle exceptions. That does not mean it should directly change a customer record, release a payment or update stock without a controlled execution layer.

My rule is simple: put AI above the workflow, not inside every control.

Let the agent decide what should happen. Let tested software decide how it happens.

A production example worth studying

Intuit has documented this split in an agentic disaster-recovery system built over its existing Ecosystem Wide Orchestrator Kit, or EWOK.[5]

EWOK already handled the deterministic work of moving services during a regional failover. Service owners declared recovery intent in structured configuration, and the system coordinated infrastructure actions across compute, databases, networking, caches and other workloads.[5]

The gap was judgment. An on-call engineer still had to identify the right recovery workflow, check whether an asset was ready and deal with policy exceptions such as a change-freeze window.

Intuit added an AI layer called EWOK Agent. Given a request such as "fail over this service in production", the agent resolves the asset, finds the available recovery workflows, selects one or asks the engineer when the choice is ambiguous, checks readiness and policy gates, triggers the existing orchestrator, then returns the execution ID and change record.[5]

The agent does not hold the production credentials or perform the low-level failover itself. Authentication is injected into the executor through request-scoped context, while the underlying system applies the same change-management, approval and audit rules used for any other authenticated caller.[5]

Intuit summarises the design boundary neatly: the model decides what to do; EWOK deterministically executes how.[5]

That is the useful lesson. It does not depend on AWS or on running thousands of microservices. It is an operating pattern that an SME can apply to ordinary finance, sales, service and operations work.

Split the job into four roles

A governed agent workflow has four distinct responsibilities.

1. The agent interprets

The agent reads the request and works out the intent.

It can collect context, compare options and handle language that would be difficult to express as a rigid form. If information is missing or two actions are possible, it should stop and ask rather than silently choose.

This is where probabilistic AI is useful. Real business requests are often incomplete, inconsistent or buried in email and documents.

2. The executor controls

The executor is conventional software: an ERP API, workflow engine, automation service or tightly scoped function.

It checks identity and permission. It validates the data. It enforces limits. It writes the approved change to the system of record. It returns a receipt.

The agent should call this executor through named, narrow tools. It should not receive a shared administrator password or an open connection that lets it improvise outside the approved job.

3. The human approves consequences

A human remains accountable for judgment calls with material consequences.

That may include approving a payment above a threshold, authorising a refund outside policy, accepting a stock write-off or deciding whether an exception justifies bypassing a normal control.

The point is not to make a person click "approve" after every harmless step. Put the gate where the consequence changes.

4. The audit layer proves what happened

Every run needs a record that answers basic questions:

  • What did the user request?
  • What context did the agent use?
  • Which action did it propose?
  • Which policy and approval gates were applied?
  • What changed in the system?
  • What was the result, and can it be reversed?

A green message in a chat window is not proof. The evidence should come from the system that performed the action.

What this looks like in an SME

Take supplier invoice approval.

An accounts employee forwards an invoice and writes: "Please process this for next Friday. It is for the warehouse repair we approved last month."

The agent can read the invoice, identify the supplier, match it to the purchase order and retrieve the relevant approval record. It can also notice that the invoice amount differs from the purchase order and explain the exception.

The agent then proposes one of three bounded outcomes: approve for posting, route for clarification or reject as a duplicate.

The executor handles the rest. It verifies the supplier ID, checks for an existing invoice number, enforces the approval limit, confirms the accounting period and writes the transaction only after the required approval is present.

If the amount exceeds the finance manager's limit, the system creates an approval task. The agent can prepare the explanation, but it cannot release the payment.

After posting, the executor returns the transaction ID, timestamp, approver and status. The agent reports those facts to the employee.

The same pattern works for customer refunds, CRM updates, stock adjustments and employee onboarding:

  • The agent understands the request and gathers context.
  • The executor applies permissions and business rules.
  • The human decides the consequential exception.
  • The audit record proves the result.

Where teams get this wrong

The common shortcut is to connect a capable model directly to a broad account and rely on instructions such as "do not make risky changes".

Instructions matter, but they are not access control. A model can misunderstand the request, use stale context or encounter a case its prompt never covered.

Another mistake is to build so many approval prompts that people click through without reading. That creates the appearance of human control without much judgment.

The better approach is to classify actions by consequence. Reading a product record may be low risk. Changing a price, issuing a refund or sending a customer message may require a stronger permission or human decision.

Deterministic execution does not make the system risk-free. The rules can be wrong. The integrations can fail. People can approve poor decisions. The design simply gives the business clear places to test, restrict, observe and recover.

Start with one bounded workflow

Do not begin with "an agent that can run operations". That job is too vague and the access is too broad.

Choose one workflow with a clear start and finish. Invoice approval is one. Updating a qualified lead in the CRM is another.

Write down:

  1. The trigger and expected outcome.
  2. The data the agent may read.
  3. The actions it may propose.
  4. The system that will execute each action.
  5. The approval threshold and stop conditions.
  6. The evidence the executor must return.
  7. The rollback or correction path.

Then connect the agent to an authenticated, auditable executor that exposes only those actions.

That is how AI moves from chat to execution without turning business systems into an experiment. The agent handles intent and ambiguity. The workflow keeps control.

Sources

[5] https://aws.amazon.com/blogs/machine-learning/how-intuit-built-an-agentic-disaster-recovery-assistant-with-amazon-bedrock — How Intuit built an agentic disaster recovery assistant with Amazon Bedrock

Continue the work

Turn a capable model into dependable execution.

Nexius Labs helps SMEs design the context, tools, permissions, approval gates, and evidence trails around useful Digital Coworkers.