Why AI Agents Fail in Production (and How Durable Execution Fixes It)
The demo always works. An agent takes a request, reasons through a plan, calls four or five tools, and returns a tidy answer. Then it ships, real traffic arrives, and a new kind of bug report starts piling up: runs that stopped halfway, customers charged twice, support tickets opened three times for the same issue, and a lot of “I don’t know what happened” from the on-call engineer.
None of this is a model problem. Swapping in a better LLM won’t fix it. Production agents fail for the same reason distributed systems have always failed: long-running work, spread across unreliable networks and processes, with no durable record of where it got to. This article covers the most common failure modes, why the usual workarounds don’t hold up, and how durable execution, a pattern borrowed from workflow engines, gives agents the reliability they’re missing.
An agent is a long-running distributed transaction
Look at what a typical agent actually does:
- It calls an LLM, which may take anywhere from 500 milliseconds to a minute and may be rate-limited.
- It calls external tools: a CRM, a payment API, a ticketing system, a vector database, an MCP server.
- It loops, re-plans and calls more tools based on what came back.
- It may wait on a human approval step that takes hours or days.
That’s a long-running, multi-step, stateful process touching several external systems, which is exactly the kind of work distributed-systems engineers learned not to run as a single in-memory function call. Most agent frameworks still run it that way by default, though. The agent’s progress lives in the memory of one Python process, and when that process goes away, so does the progress.
Five ways agents break in production
1. Process crashes and restarts
Pods get evicted. Nodes get recycled. Someone deploys a new version in the middle of the afternoon. On Kubernetes, restarts are a normal part of operation, not an edge case. If an agent is on step seven of twelve when its container is killed, a naive implementation has two options: start over from step one, or give up. Starting over is expensive (every LLM call costs money) and often unsafe, because the first six steps may have had side effects.
2. Duplicate side effects
This is the most dangerous failure mode because it’s usually silent. An agent calls a “create refund” tool, the tool succeeds, and then the process crashes before the agent records that it succeeded. On retry, the agent calls the tool again. Now the customer has two refunds. Tool calls that write to the outside world need exactly-once semantics from the agent’s point of view, which means the result of each completed step has to be persisted before moving on.
3. Transient failures that look permanent
LLM providers return 429s and 503s. Third-party APIs time out. A vector database fails over. Most of these errors go away if you wait and retry with backoff, but many agent loops treat any exception as fatal, and the whole run fails because of a two-second blip.
4. Long waits and human-in-the-loop steps
“Wait for a manager to approve this purchase” can’t be implemented as await asyncio.sleep(86400). The process won’t live that long, and even if it did, you’d be tying up resources for no reason. Agents that need approvals, scheduled follow-ups or external callbacks need a way to suspend, release their compute, and resume later with their full context intact.
5. No audit trail
When an agent does something unexpected, like approving a large discount or emailing the wrong customer, the first question is “what exactly happened, and in what order?” Scattered logs rarely answer it. For regulated industries, “we think the model did this” isn’t an acceptable answer to an auditor.
Why the usual fixes don’t hold up
Teams usually reach for one of three workarounds.
Framework checkpointers. Some frameworks can persist state to a database after each step. That helps, but persistence isn’t recovery. If the process dies, the checkpoint sits in Postgres and nothing notices the run has stalled. Someone still has to build the supervisor that detects dead runs and resumes them, and make that supervisor itself reliable.
Queues and retries everywhere. Wrapping every tool call in a message queue with retry logic works in theory, but it spreads orchestration logic across dozens of handlers, makes the control flow hard to follow, and still doesn’t solve idempotency on its own.
Hand-written state machines. Some teams model agent progress as rows in a database table with a status column. That’s a workflow engine built by hand, without the testing, tooling or years of edge-case fixes.
Durable execution: the missing layer
Durable execution is a programming model in which the runtime, not your code, is responsible for making sure a multi-step process runs to completion. You write the orchestration as ordinary code, and the engine records every step’s input and output in an append-only history. If the process crashes, the engine replays that history on a healthy worker: completed steps return their recorded results instantly instead of re-executing, and execution continues from the exact point of failure.
For agents, the mapping is natural:
- Each LLM call becomes a durable activity. If it succeeds, its response is recorded and never paid for twice.
- Each tool call becomes a durable activity with its own retry policy. A recorded success is never repeated, which prevents duplicate side effects.
- Waits for human approval or external events become durable timers or event subscriptions. The agent consumes no compute while it waits.
- The execution history is a complete, ordered record of everything the agent did, which is exactly what you need for debugging and audits.
The open-source Dapr Workflow engine is one widely used implementation. Workflows are written as code in Python, .NET, Java, Go or JavaScript, and state is stored in a pluggable state store. A minimal example looks like this:
import dapr.ext.workflow as wf
wfr = wf.WorkflowRuntime()
def support_agent(ctx: wf.DaprWorkflowContext, ticket: dict):
plan = yield ctx.call_activity(plan_with_llm, input=ticket)
for step in plan[“steps”]:
yield ctx.call_activity(run_tool, input=step)
return (yield ctx.call_activity(summarize, input=ticket))
Each yield is a checkpoint. If the worker dies after the second tool call, a new worker replays the history, skips the completed calls using their recorded results, and runs the third. Platforms such as Diagrid build on this same engine to run agent workloads durably at scale.
One rule to know: keep orchestration deterministic
Durable execution relies on replay, so the orchestrating code has to be deterministic: given the same history, it must make the same decisions. In practice that means:
- Put all I/O (LLM calls, HTTP requests, database writes) inside activities, never directly in the workflow function.
- Don’t read the system clock or generate random values in the workflow body; use the workflow context’s replay-safe equivalents.
- Keep activities idempotent where you can, for example by passing an idempotency key to payment or ticketing APIs.
LLM output isn’t deterministic, but that doesn’t matter here. Because the LLM call runs inside an activity, its result is recorded once and replayed from history. The workflow only sees the recorded response.
You don’t have to rewrite your agent
A reasonable objection: “We already built our agent in LangGraph (or CrewAI, or the OpenAI Agents SDK). We’re not rewriting it as a workflow.”
You don’t need to. A growing number of integrations wrap existing agent frameworks in a durable execution engine so each node or tool call automatically becomes a durable activity. Diagrid, the company founded by Dapr’s creators, offers durable execution for AI agents through Catalyst, which plugs into frameworks such as LangGraph, CrewAI, Google ADK, AWS Strands, the OpenAI Agents SDK and Microsoft Agent Framework with a few lines of code. It also adds cryptographically signed execution history for teams that need to prove what an agent did.
Whatever tooling you pick, look for these properties:
- Automatic recovery: crashed runs resume without an external supervisor.
- Per-step retries with configurable backoff, timeouts and limits.
- Durable waits for timers, human approvals and external events.
- A queryable execution history you can inspect, replay and export.
- Deployment flexibility, so agent state can stay in your own cloud or data center if your compliance team requires it.
- Framework neutrality, because the agent framework you pick today probably won’t be the only one you run in two years.
A production-readiness checklist for agents
Before an agent goes to production, ask these questions:
- If the process is killed at any step, does the run resume automatically, and from where?
- Can any tool with side effects be executed twice for the same logical request?
- What happens when the LLM provider returns a 429 for thirty seconds?
- Can the agent wait for a human approval for three days without holding a process open?
- Can you reconstruct exactly which tools the agent called, with which inputs, in which order?
- Could you show that record to an auditor and prove it hasn’t been altered?
If the honest answer to most of these is “no” or “not sure,” the model isn’t what’s standing between you and a reliable production agent. The missing piece is the execution layer underneath it.
Conclusion
Agentic AI has brought one of distributed computing’s oldest lessons back to the front: any process that takes more than a few seconds and touches more than one system will eventually be interrupted. Durable execution doesn’t try to prevent failure. It makes failure recoverable by design. Treating LLM calls and tool calls as durable, recorded steps turns fragile demos into agents that finish their work, never repeat a dangerous side effect, and leave a clear record of what they did.