Your finance agent has a ten-step job.
It reads an approved sales order, checks the customer record, creates an invoice, posts the receivable, updates the CRM, drafts the customer email, attaches the invoice, sends the message, records the activity, and closes the task.
Then it crashes after step six.
The agent screen shows an error. The invoice already exists. The ERP may already contain the receivable. The CRM has changed. The email has not been sent. Nobody knows whether the attachment was created correctly.
The tempting response is simple: restart the workflow.
That can make the situation worse. A full restart may create a second invoice, post another receivable, overwrite the CRM field, or send two customer messages. The agent did not fail before doing anything. It failed after doing part of the work.
This is the reliability problem operators need to design for.
Recovery is part of the business workflow
Reuters reported on 14 September that Temporal raised $550 million at a $12.55 billion valuation. The company makes software that helps applications, including AI agents, recover from failure instead of requiring engineers to build all the recovery logic themselves.[1]
The funding is not the main point for an SME. The useful signal is what investors and builders are treating as valuable: long-running software must survive interruptions without losing track of what happened.
An AI agent that works across ERP, CRM, finance, email, and service systems is not one action. It is a chain of actions with different consequences. Some steps only read data. Others create records, trigger transactions, or communicate with customers.
A chat response can be regenerated. A partially executed business process cannot be safely regenerated without checking its state.
Every production agent therefore needs a failure contract before it gets a real workflow.
The seven-part failure contract
The contract does not need to be a long policy document. It needs to answer seven operational questions clearly.
1. Where is the checkpoint?
The agent must record progress after each meaningful step, not only when the whole task finishes.
For the invoice example, the checkpoint should show:
- customer record verified;
- invoice created, including its invoice number;
- receivable posted, including its transaction reference;
- CRM updated, including the changed fields;
- email drafted but not sent;
- workflow paused before the next action.
A checkpoint is more than “step six completed.” It should carry the identifiers needed to inspect the systems that changed. Without those references, recovery becomes guesswork.
2. What is the idempotency key?
Give every workflow run a unique business key and pass it through every system that can store it.
For example:
invoice-followup-SO-48217-2026-09-15
Before creating a record or sending a message, the agent checks whether that key already exists for the intended action. If it does, the agent should inspect the existing result rather than create another one.
This is how the workflow distinguishes “try again” from “do it twice.”
A timestamp alone is not enough. The key should connect the action to the business object: the order, invoice, case, account, or approval request being processed.
3. Which retries are safe?
Not every failed step should be retried in the same way.
A temporary read failure may justify two or three automatic retries. Creating a financial record or sending a customer email needs a state check before any retry. A rejected payment, invalid tax code, or permission error should usually go to an exception queue instead of looping.
Unlimited retries are not resilience. They are repeated uncertainty.
Set a retry cap for each action type. Record the reason for every retry. When the cap is reached, stop cleanly and escalate with the current state attached.
4. How are duplicates prevented?
Duplicate prevention must exist at the point where the consequence occurs.
The orchestration log may say an invoice was not created because the agent timed out before receiving confirmation. The ERP may still have created it. The recovery step must query the ERP using the business key, order number, amount, customer, and expected date before issuing another create request.
Apply the same discipline to CRM tasks, support tickets, purchase orders, messages, and approvals.
Do not trust the absence of a success response as proof that the action failed. Verify the target system.
5. Who owns rollback?
Rollback is a business decision as much as a technical one.
If the agent creates an invoice and updates the CRM but fails before notifying the customer, the right response may be to continue from the email step. It may not be appropriate to delete or reverse the valid accounting entries.
If the agent used the wrong customer or amount, reversal may be required. Finance should define that path before deployment: who can cancel the invoice, which audit note is required, and whether a credit entry must be created instead of deleting history.
The recovery runbook should say which actions are reversible, which are compensating actions, and which need human approval.
6. Where does unresolved work go?
A failed agent run must not disappear into a technical log.
Route it to a human-visible exception queue with:
- the workflow and business object;
- the current checkpoint;
- systems already changed;
- the failed step and error;
- retries attempted;
- the recommended next action;
- the person or team who must decide.
The queue is where the digital coworker hands judgment back to an operator. It should be monitored like an operational work queue, not treated as a developer backlog.
7. What proves completion?
“Agent finished” is a technical status. It is not proof that the business outcome exists.
For the invoice workflow, completion evidence should confirm that:
- the ERP contains one valid invoice;
- the receivable points to that invoice;
- the CRM contains the intended update;
- one customer email was sent to the correct recipient;
- the activity record links to the message and invoice;
- no unresolved exception remains.
Each affected system needs a read-back check. The agent can assemble those checks, but the workflow should not report success until the evidence agrees.
Design the restart before the first run
Most agent demonstrations show the happy path. The agent receives a task, uses several tools, and returns a completed result.
Production work is defined by the unhappy path: expired sessions, slow APIs, duplicate events, permission changes, malformed records, partial writes, and ambiguous responses.
An SME does not need a complicated reliability platform on day one. It does need explicit recovery behaviour.
Take one real workflow and interrupt it deliberately after every step. Check what changed. Try the same action twice. Disconnect a system after it accepts a request but before it returns confirmation. Confirm that the exception appears where an operator can see it. Then practise recovery.
This test usually exposes more than another round of prompt tuning.
A capable agent can execute a ten-step process. A dependable digital coworker can tell you exactly which six steps succeeded, avoid repeating them, recover from the seventh, and prove the final business state.
Before granting production access, write the failure contract.
Sources
[1] https://www.reuters.com/business/temporals-valuation-spikes-126-billion-lightspeed-led-funding-round-2026-09-14 — Temporal’s valuation spikes to $12.6 billion in Lightspeed-led funding round