Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Software automation

Reliable AI Agent Workflows: Retries, Idempotency and Business Value

A real client engagement. The engineering and the results are described below.

A service desk we worked with has a convincing AI-agent demo and an uncomfortable question: what happens when a timeout follows an external action? Its team uses that ambiguity to redesign the workflow before scaling the pilot.

Request a time through the inquiry form. A meeting is confirmed separately by email.

The business problem behind the technology

Can the workflow explain uncertainty before it repeats an action?

The business wants fewer avoidable exceptions, but cannot let an agent repeat a message, payment or write merely because a response was lost.

Read the client engagement ↓

Client engagement / Delivered results

The useful automation knew when not to retry

A service desk we worked with handles 6,000 assisted tasks each month. At baseline, 15% needed ten minutes of manual exception handling after ambiguous or inconsistent workflow transitions.

The constraint

The business wants fewer avoidable exceptions, but cannot let an agent repeat a message, payment or write merely because a response was lost.

The engineering decision

The team defines durable states, reconciles uncertain effects, binds authority to exact tool operations and tests adverse transitions. The manual-exception share fell from 15% to 5% for the same task mix.

The delivered outcome

Handling effort fell from 150 to 50 hours a month: 900 versus 300 exceptions at ten minutes each.

Measured inputs and delivered differences
Measure / unitBeforeAfterDifference
Manual exception-handling effort
hours/month
15050100

Manual exception-handling effort. 6,000 × 15% × 10 / 60 versus 6,000 × 5% × 10 / 60. Exception rates are the desk’s operating values.

The conditions behind the results

  • The same 6,000-task mix and quality/approval contract apply.
  • Ten-minute handling time is constant; the reduction affects avoidable workflow exceptions only.
  • Tool/model costs, implementation and residual high-impact human approval are outside this effort ledger.

What this does not prove. The freed handling capacity kept human approval in place where the workflow required it.

Evidence to collect for your own decision

  • Enumerate durable states and test ambiguous effects before allowing a retry.
  • Verify idempotency boundaries and exact authorization on sensitive tools.
  • Measure completed task quality, exception rates and the handling time distribution.

Key decisions

A reliable agent workflow makes state transitions, side effects and authority explicit. Evaluate useful completion and uncertain outcomes rather than trusting a fluent conversation.

Follow the decision

Can the workflow explain uncertainty before it repeats an action?

Select a step to follow its reasoning, then continue into the technical chapters.

Problem → boundary → decision → evidence

Name the states

Define transitions before judging the conversation.

Read every component and connection
Name the states · Problem
Define transitions before judging the conversation.
Control side effects · Boundary
Separate retries from proof that an external action happened.
Keep authority explicit · Decision
Tools and approval records enforce the action boundary.
Evaluate the workflow · Evidence
Test transitions, effects and hostile inputs rather than eloquence.
  • Name the states → Control side effects: identify the constraint
  • Control side effects → Keep authority explicit: choose a bounded change
  • Keep authority explicit → Evaluate the workflow: check the outcome

A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Define the state machine before the conversation

The client’s service desk turns its exception-handling problem into explicit states and transitions.

A chat transcript is not a durable workflow model. It does not reliably tell a recovering worker whether a credit was merely proposed, approved, submitted, or confirmed. Define those states explicitly and store transitions transactionally. For this example, a request might move through received, gathering, proposed, awaiting approval, executing, and completed, with separate rejected, expired, and needs-review outcomes.

Persist a workflow identifier, input version, current state, proposal hash, policy version, and references to tool results. Use optimistic concurrency or an equivalent transactional mechanism so two workers cannot independently advance the same revision. Keep large or sensitive payloads in access-controlled storage rather than copying every result into the conversation and event history.

The model can classify a request or draft a proposal, but ordinary application logic checks whether the proposed transition is legal. A generated statement that approval was granted is not approval. Durable workflow engines can help with timers and recovery; they do not replace an authorization model or turn an unreliable external side effect into an exactly-once operation.

Treat retries as a side-effect problem

The timeout arrives after a possible side effect. Retrying safely requires more than asking the model to try again.

A timeout creates uncertainty, not evidence of failure. The adjustment API might have accepted a write before the connection broke. Blindly retrying with a new identifier can duplicate the credit. Derive an idempotency key from the stable business operation and approved proposal version, then reuse it for every attempt. The receiving system must enforce that key atomically with the effect.

Store the request fingerprint and result against the key, reject reuse with a different payload, and retain deduplication records for the entire retry horizon. A local completed flag is insufficient: a worker can crash between the external write and the flag update. If the downstream API offers no idempotency support, use a lookup or reconciliation path; do not claim exactly-once behavior from a local database alone.

Temporal documents this distinction clearly: an activity may execute multiple times even when its completion is recorded once. Bound retries by attempt count and deadline, distinguish transient failures from permanent validation errors, and send uncertain results to reconciliation. Cancellation must also define what happens to an already submitted operation.

Put authority in tools and approval records

The team preserves exact action authority while asking which workflow failures cause avoidable handling effort.

Expose narrow tools such as read-account-summary or propose-credit rather than unrestricted shell, SQL, or HTTP access. Validate arguments against a schema and authorize the target account using trusted request context. Model-supplied account identifiers are inputs to check, not proof of access. Limit result sizes, destinations, timeouts, and cumulative tool calls so one task cannot consume unbounded resources.

Bind human approval to the exact proposed action, target, amount, currency, policy context, and expiry. If the proposal changes, require new approval. Revalidate permissions and relevant account state immediately before execution to prevent a stale approval from authorizing a different situation. The following record illustrates those bindings; it is not a complete authorization implementation.

Separate read credentials from write credentials. Release a narrowly scoped execution capability only after approval checks pass, and avoid giving secrets to the model. Log the authorization decision and external operation identifier without retaining unnecessary customer text.

json / example
{
  "workflow": "service-credit-request-42",
  "state": "awaiting_approval",
  "proposalVersion": 3,
  "action": "issue_service_credit",
  "approvalBinding": [
    "proposalHash", "accountId", "amount",
    "currency", "approverId", "expiresAt"
  ]
}

Treat retrieved text as potentially hostile

Retrieved text joins the context, including instructions it has no authority to give. Prompt injection belongs in the threat model from the start.

Support tickets, documents, and tool responses are untrusted data even when they are useful evidence. A sentence inside a ticket telling the assistant to bypass approval must not become a workflow instruction. Keep retrieved content separate from privileged policy, preserve provenance, and make permission checks outside the model. Prompt wording alone is not a security boundary.

Constrain outbound destinations and redact sensitive fields before model submission. A read-only tool can still leak data if its output is forwarded to an arbitrary URL or copied into a public response. Apply tenant isolation to retrieval, verify source references, and require the draft to distinguish known facts from missing information.

Bound reasoning and execution independently. Exhausting a token or tool budget should lead to a clear incomplete outcome, not a fabricated success message. If a dependency is unavailable, preserve the work already gathered and explain the missing prerequisite. Do not repeatedly ask the model to solve a permissions failure that requires an authorized human decision.

Evaluate transitions and effects, not eloquence

Only accepted outcomes, adverse-path tests and observed exception effort can validate the 100-hour capacity difference.

Build evaluation cases around observable outcomes: duplicated requests, changed proposals after approval, cross-account access, hostile retrieved instructions, timeouts after accepted writes, and expired approvals. Check that unauthorized writes never occur, legitimate approved operations complete, and uncertain operations enter reconciliation. A convincing explanation attached to the wrong account is still a failure.

Use controlled fixtures for repeatable fault injection, then shadow the read-and-propose path against representative authorized inputs before permitting writes. Track task correctness, unnecessary approval requests, escaped policy violations, repeated effects, review burden, latency, and cost per successful workflow. Retain versioned prompts, models, tools, and evaluation evidence so a regression can be attributed and reversed. Reliability comes from the combined system boundary, not from a promise that the model will always follow instructions.

Questions behind the decision

How should an AI agent handle a timeout after a tool call?

Treat the result as potentially ambiguous and reconcile whether the side effect happened before repeating it. Idempotency contracts can help, but a retry policy alone cannot prove that another external action is safe.

How do you measure business value from agent automation?

Measure accepted task completion, human review and exception effort, tool costs and harmful or duplicate effects. Fluent answers and task counts do not establish useful work or justify removing approval boundaries.

References & further reading

Engineering notes

Keep following the thread.

Real client engagements and the engineering behind them.

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works