Your Agent Crashed Halfway Through A Payment Run: Durable Execution, Idempotency And Sagas For Financial Workflows
Agent frameworks spent two years rediscovering a problem distributed systems solved a decade ago, and 2026 is the year the industry admitted it: a long-running agent is a workflow, and workflows need durable execution. The failure mode is specific and expensive - a tool call succeeds, the process dies before the result is recorded, the workflow resumes and calls it again. Without an idempotency key that means a duplicate payment, a duplicate order, a duplicate ticket. Agents make it harder still, because they are probabilistic: a retry does not reproduce the previous decision, so the usual cache-lookup answer does not apply. Here is how we build financial agent workflows that survive crashes, deploys and their own non-determinism.
AlchmAI Engineering15 min read
Exactly once
What durable execution engines guarantee for activity completion, via deterministic replay of a logged workflow history
Probabilistic
Why agents break the usual retry assumptions - the same prompt can produce a different decision, so a retry is not a repeat
Every side effect
Needs its own idempotency story; a tool call that succeeds before a crash will otherwise run twice on resume
Compensate
Not roll back - financial side effects that have left your system can only be reversed by a second, deliberate action
There is a specific incident shape that every team running long-lived agents eventually experiences, and it is worth describing precisely because the lesson is in the details. An agent is working through a batch - reconciling breaks, releasing payments, submitting orders, updating records. It calls a tool. The tool succeeds. Before the framework records that success, the process dies: a deploy, a pod eviction, an OOM kill, a network partition, or simply a bug. The workflow resumes from its last checkpoint, decides the tool call has not happened, and calls it again.
In a documentation assistant that produces a duplicate log line. In a financial workflow it produces a duplicate payment, a duplicate order, or a duplicate instruction to a counterparty, and you find out from reconciliation the next morning. The industry spent 2026 arriving at the conclusion that agent workflows are rediscovering durable execution, and the framing is correct: durable execution provides the automatic state persistence, retries and workflow resumption that make agents production-ready, ensuring workflows run to completion even across server restarts, deployments and outages. Engines like Temporal guarantee that each activity is effectively run once even if the worker retries or crashes, which the documentation describes as critical for payments, emails and other act-only-once operations.
Rule One: Decide Once, Record The Decision, Then Act
The structural fix is to split every agentic step into two activities with a durable boundary between them. The first is non-deterministic and has no side effects: it calls the model and produces a decision object. The second is deterministic and has side effects: it executes the decision. The engine persists the decision between them, so a crash during execution resumes with the original decision rather than asking the model again.
from datetime import timedelta
from temporalio import workflow, activity
@workflow.defn
class PaymentReleaseWorkflow:
@workflow.run
async def run(self, batch_id: str) -> ReleaseSummary:
items = await workflow.execute_activity(
load_batch, batch_id, start_to_close_timeout=timedelta(minutes=2)
)
released, compensated = [], []
for item in items:
# 1. REASON. Non-deterministic, no side effects. Retried freely -
# a different answer here costs nothing because nothing has
# happened yet. The result is persisted by the engine.
decision = await workflow.execute_activity(
assess_payment,
item,
start_to_close_timeout=timedelta(seconds=90),
retry_policy=RetryPolicy(maximum_attempts=3),
)
if decision.action != "release":
continue
# 2. GATE. Deterministic policy, outside the model entirely.
verdict = await workflow.execute_activity(
pretrade_gate, decision, start_to_close_timeout=timedelta(seconds=5)
)
if not verdict.allowed:
continue
# 3. ACT. Deterministic, idempotent, side-effecting. On resume
# after a crash this replays with the SAME decision object -
# the model is never consulted again for this item.
try:
receipt = await workflow.execute_activity(
submit_payment,
SubmitRequest(
decision=decision,
# Derived from workflow identity + item, so it is
# stable across every retry and every resume.
idempotency_key=f"{workflow.info().workflow_id}:{item.id}",
),
start_to_close_timeout=timedelta(seconds=30),
retry_policy=RetryPolicy(maximum_attempts=5),
)
released.append(receipt)
except ActivityError:
# 4. COMPENSATE. Not a rollback - a deliberate reversing action.
await workflow.execute_activity(
cancel_payment, item.id, start_to_close_timeout=timedelta(seconds=30)
)
compensated.append(item.id)
return ReleaseSummary(released=released, compensated=compensated)Rule Two: Every Side-Effecting Tool Needs An Idempotency Story
The engine guarantees your activity runs once from its point of view. It cannot guarantee the downstream system behaves, and in financial infrastructure the downstream systems vary enormously in what they offer. Four cases, in descending order of comfort:
- 01Native idempotency keys. Most modern payment APIs accept one and will return the original result for a repeated key. Use it, store the key with the decision, and you are genuinely safe.
- 02Client-supplied unique identifiers. FIX order submission with a client order ID, or any API that rejects a duplicate identifier. Equally safe in practice, with the wrinkle that you must handle the duplicate-rejection error as a success rather than a failure.
- 03Read-your-write verification. No idempotency support, but you can query whether the effect exists. Before acting, check; after acting, confirm. This is a race and you must treat it as one - use a lease or an application-level lock keyed on the same stable identifier.
- 04Nothing at all. A legacy system that accepts an instruction and offers no way to ask whether it already did. Here the only honest answer is an outbox in your own database, written transactionally with the decision, and a single-threaded dispatcher that marks entries sent. You cannot make the remote system safe; you can make sure you never send twice.
@activity.defn
async def submit_payment(req: SubmitRequest) -> Receipt:
"""Idempotent by construction. Safe to call any number of times with the
same request; exactly one payment leaves the building."""
key = req.idempotency_key
# Outbox written in the same transaction as the decision record. If the
# process dies after commit but before dispatch, the dispatcher picks it
# up; if it dies after dispatch but before marking sent, the provider's
# idempotency key protects us on the retry.
async with db.transaction() as tx:
existing = await tx.fetchrow(
"SELECT status, receipt FROM payment_outbox WHERE idempotency_key = $1",
key,
)
if existing and existing["status"] == "sent":
return Receipt.parse_raw(existing["receipt"]) # already done
if not existing:
await tx.execute(
"INSERT INTO payment_outbox (idempotency_key, payload, status) "
"VALUES ($1, $2, 'pending')",
key, req.json(),
)
receipt = await provider.pay(req.decision, idempotency_key=key)
await db.execute(
"UPDATE payment_outbox SET status='sent', receipt=$2, sent_at=now() "
"WHERE idempotency_key = $1",
key, receipt.json(),
)
return receiptRule Three: Compensate, Do Not Roll Back
Engineers arriving from a database background reach instinctively for transactional rollback, and it does not exist here. Once a payment instruction has reached a counterparty, no amount of local state management retracts it. The saga pattern is the correct model: every forward action has a defined compensating action, compensation is itself a durable activity that can be retried, and compensation may fail in ways that require a human.
- Write the compensating action at the same time as the forward action, not later. A forward step whose compensation is unspecified is a step you cannot safely include in a workflow, and discovering that during an incident is expensive.
- Accept that compensation is not always symmetric. Cancelling a submitted order may fail because it already filled. The compensation for a fill is a hedging trade or an error account entry - a business decision that must be made in advance and encoded, because nobody makes it well at 14:32.
- Make compensation idempotent too. It runs under exactly the same crash conditions as the forward path, and a duplicate compensation is its own incident.
- Escalate rather than loop. When compensation exhausts its retries, the workflow should park in a state that pages a human with full context, not retry indefinitely. A workflow stuck in a compensation loop is an outage that looks like normal operation on every dashboard.
Rule Four: Determinism Constraints Are Not Optional
Durable execution works by replaying a workflow's logged history deterministically after failure. That imposes constraints on workflow code that catch every team at least once, and the failure is confusing because it appears long after the offending line was written - typically on the first resume after a deploy.
- No wall-clock reads, no random numbers, no UUID generation, no direct I/O in workflow code. Anything whose value could differ between the original run and the replay must live in an activity, whose result is recorded.
- No iteration over unordered collections. A set or a dictionary that iterates in a different order on replay produces a different history and fails the determinism check.
- Versioning matters when you change a running workflow. A workflow started under yesterday's code may resume under today's; the engines provide versioning primitives for precisely this, and using them is the difference between a safe deploy and a mass failure of in-flight work.
- Model calls are never deterministic and therefore always belong in activities. This is the same rule as above, arrived at from a different direction, and it is the one that keeps the probabilistic-retry problem contained.
“The discipline durable execution imposes is the discipline financial workflows needed anyway. Most teams discover that the constraints are not a tax on the agent - they are the specification for a system they could actually operate.”
Where This Fits Next To The Agent Framework
A reasonable question: agent frameworks have their own orchestration, state and retry handling, so why introduce a second layer? Our answer after building both ways is that they operate at different levels and the boundary is clean once you draw it. The agent framework owns a turn: assembling context, choosing a tool, parsing a result. The durable engine owns the process: what has happened, what must happen next, what must be undone, and what survives a restart.
- Use the agent framework's loop for anything inside a single reasoning step. It is good at that and rebuilding it is wasted effort.
- Use the durable engine for anything that spans steps, spans minutes, or touches the outside world. Framework-level retry state that lives in memory is not state, it is an optimistic assumption.
- Keep the model call inside an activity so the framework's non-determinism is contained by the engine's replay boundary.
- Resist encoding business sequence in the prompt. If the order of operations matters - and in financial workflows it always does - it belongs in workflow code where it is reviewable, testable and enforced, not in an instruction the model may reorder.
Testing It Honestly
The tests that have caught real defects for us, in order of value per hour invested:
- 01Kill the worker mid-activity, repeatedly, at randomised points, and assert the external effect count. Most engines provide a test environment that makes this straightforward. Run it a few hundred times; the failures cluster around the boundaries you did not think about.
- 02Duplicate every activity invocation deliberately in a test harness and assert that the downstream state is identical. This is the idempotency property stated as an executable assertion rather than an intention.
- 03Force compensation paths. They are the least-exercised code in the system and the code most likely to run during an actual incident.
- 04Replay production histories against new workflow code before deploying. This is the determinism check that prevents a version change from breaking in-flight work, and it costs almost nothing to run in CI.
- 05Assert that no model call appears outside an activity. A static check, trivially automated, and it enforces the rule that makes everything else work.
The Bottom Line
Long-running agents are workflows, and financial workflows have needed durable execution since long before anyone put a model in one. The 2026 realisation is simply that the two problems are the same problem: a process that spans minutes, touches external systems and must survive crashes, deploys and partitions. What agents add is a genuinely new wrinkle - they are probabilistic, so a retry is not a repeat, and the answer is to never put a model call and a side effect in the same retryable unit. Split reason from act with a durable boundary between them, derive idempotency keys from workflow identity so they survive replay, give every forward action a compensating action written at the same time, respect the determinism constraints because replay depends on them, and test by killing workers rather than by hoping. Do that and a crash mid-payment-run becomes a non-event rather than a morning of reconciliation. That is the workflow automation architecture we build for financial clients in London, and it is the layer that decides whether an agent is a demo or a system.
References & Further Reading
- Temporal - durable execution documentation. docs.temporal.io
- Inngest - Durable execution: the key to harnessing AI agents in production. inngest.com/blog/durable-execution-key-to-harnessing-ai-agents
- Olmec Dynamics - Temporal and the 2026 shift to durable agentic workflows. olmecdynamics.com/news/temporal-durable-execution-agentic-workflows-2026
- Spheron - AI agent workflow orchestration: Temporal, Inngest and Restate for durable multi-step pipelines (2026). spheron.network/blog/ai-agent-workflow-orchestration-temporal-inngest-restate-gpu-cloud
- Atomix: timely, transactional tool use for reliable agentic workflows (arXiv). arxiv.org/pdf/2602.14849
- Verified detection and prevention of concurrency anomalies in multi-agent LLM systems (arXiv). arxiv.org/pdf/2606.17182
- Garcia-Molina & Salem - Sagas (the original paper). cs.cornell.edu/andru/cs711/2002fa/reading/sagas.pdf
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information