By Mert Dönmezler12 min read

Orchestration Patterns for Multi-Agent Workflows

The reliability of a multi-agent workflow is set by its weakest recovery path. How to allocate autonomy step by step, and how to design retries, compensation, idempotency keys and verification around that.

  • ai-agents
  • orchestration
  • distributed-systems
  • research
  • saga-pattern
  • reliability
  • agentic-ai

Most writing about multi-agent workflows treats autonomy as something to maximise, with deterministic orchestration as the cautious fallback. We see the two as ends of one decision about who holds control, and that decision gets made again at every step. How reliable an agentic workflow is depends on its weakest recovery path much more than on its best reasoning trace. Below we go through where to place each step on that spectrum, how to compute the compounding reliability of a chain before you build it, and how to write compensating actions for side effects that cannot be undone. After that come idempotency keys that survive replay, and a verification stage that amounts to more than the model reviewing itself.

Autonomy allocation in multi-agent workflow design

Every step in a system that calls a language model has a control locus. Something decides what happens next, and that something is either code written in advance or a model deciding at run time. If you frame the design question as "workflow or agent", a decision that belongs to each step turns into one architectural stance for the whole system. The more useful question is how many steps really need judgement at run time, and what each unit of judgement costs in variance, latency and recovery complexity.

Fully scripted control

Control flow, branching, retries and error handling are all written in code, and no model chooses the next action. Variance is near zero, and failures are easy to diagnose because the execution graph is known before the run. The limitation is that scripted control only handles the cases its author listed, and listing them gets more expensive the messier the input domain is.

Scripted control with judgement at the leaves

Control flow is still deterministic code, but leaf operations are handed to a model: classify this document, extract these fields, decide whether two addresses refer to the same venue. The model returns a typed, validated value and the orchestrator routes it. Variance stays inside the leaf, and so does recovery. A bad classification is a wrong value in a known slot, and it has not set off a sequence of side effects that you now have to reconstruct.

Fully autonomous agent loops

The model holds the control flow and alternates between reasoning and acting until it decides the task is done (Yao et al., 2022). This builds on the observation that making intermediate reasoning explicit changes how models behave on multi-step tasks (Wei et al., 2022). It is the only point on the spectrum that can handle truly open-ended tasks. It is also the only one where you cannot know in advance which actions a failed run took, so in principle there is no upper bound on its cost or on what compensation it needs.

AllocationControl locusOutput varianceCost predictabilityDominant failure mode
Fully scriptedCodeNoneHighUnhandled input class
Judgement at the leavesCode, model-valued leavesBounded per leafModerateWrong value in a known slot
Fully autonomousModelUnboundedLowUnknown side-effect sequence

Scripted control with judgement at the leaves should be the default. Spend autonomy where the input domain really cannot be enumerated. The reason is in the table itself: output variance, the third column, decides whether the failure mode in the fifth column can be recovered from.

Compounding reliability across multi-step agent chains

The most common design error in agentic workflows is to read per-step accuracy as end-to-end accuracy. If a chain of n steps must all succeed, and each succeeds independently with probability p, the chain succeeds with probability pn. At p = 0.95 and n = 10 that is roughly 0.599, and at n = 20 roughly 0.359. Working backwards, to reach 99% end-to-end across ten dependent steps each step has to succeed at about 0.991/10 ≈ 0.9990, which means cutting per-step error roughly fiftyfold. These figures follow from the independence assumption and are not measurements. Independence is also the optimistic case: one ambiguous input that trips several steps in the same way makes the real distribution worse than the product.

Recovery closes most of that gap. If a step can be retried and its failures are roughly independent across attempts, three attempts at p = 0.95 cut step failure from 0.05 to about 0.000125. In the same illustrative model, the ten-step chain goes from roughly 0.599 to roughly 0.9988, and all of that gain comes from the recovery path. A single step that cannot be retried, on the other hand, puts its own failure probability back onto the whole chain, which is why the weakest recovery path is the part of the design to look at first. Which failures can be retried at all depends on how you classify them. A failure taxonomy for robotic process automation carries over directly, because the categories that resist retry in scripted automation resist it in agentic workflows for the same reasons.

Parallelism has a similar limit. Fanning work out to several agents cannot bring latency below the serial part of the job (the merge, the reconciliation, the final verification). That is Amdahl's argument (Amdahl, 1967) applied to a different resource. With four agents and a serial fraction of 0.3, the ceiling is 1/(0.3 + 0.7/4) ≈ 2.1, not 4. When fan-out is worth its cost, the reason is usually coverage, and only rarely latency.

Agent workflows as long-running transactions

A workflow that reads a record, calls a model, writes a row, calls an external API and sends a notification is a distributed transaction in every respect, except that it holds no locks and offers no atomicity. Gray and Reuter's treatment of transaction processing (Gray & Reuter, 1993) explains why ACID properties cannot be held across independent systems for as long as a model call takes. Nobody can keep locks open across services for minutes, so the guarantee has to be rebuilt in the application.

The saga pattern (Garcia-Molina & Salem, 1987) is the usual way to rebuild it. A saga is a sequence of local transactions T1…Tn, and each one has a compensating action Ci that undoes it semantically. If Tk fails, the orchestrator runs Ck-1…C1 in reverse order. The undo works at the level of meaning. The inverse of "charged the customer" is "issued a refund", so both events stay in the ledger and the system never returns to a state where neither happened.

run_saga(steps, ctx):
  done = []
  for step in steps:
    try:
      step.forward(ctx)          # idempotent, keyed
      done.append(step)
    except:
      for s in reverse(done):
        s.compensate(ctx)        # idempotent, keyed, never silent
      raise

Two things follow from this. First, some actions have no compensator: a sent email, a transferred sum, a deleted file. For those, the structural answer is ordering. Every reversible step and every verification stage goes before the irreversible one, so that by the time the workflow reaches the point of no return it knows as much as it can. Second, compensators run rarely, which makes them the least tested paths in the system. Until a compensator has actually been exercised, you have no evidence that it works.

Idempotency keys and replay in agent workflows

Retries and compensations are only safe if the operations are idempotent, and in practice idempotency means a key. Kleppmann's account of delivery semantics (Kleppmann, 2017) puts it plainly. Exactly-once delivery is not generally achievable, but effectively-once processing is, as long as each operation carries a deduplication key that the receiver stores and checks. Data engineering has the same requirement, and data contracts and idempotency are what keep a pipeline reliable when it is re-run.

Deriving the key is where agentic systems differ from ordinary pipelines. The instinct is to hash the operation's arguments. In an agent workflow, though, those arguments are often model outputs, and model outputs are not stable across runs. Hashing them gives a new key on every retry, and the retry becomes a duplicate side effect. The key has to come from deterministic upstream state:

key = hash(workflow_run_id, step_name, attempt_group,
           canonical(deterministic_inputs))

Model-produced values go into a decision log and stay out of the key. That keeps apart two operations that often get mixed up. Replay rebuilds state by re-reading the log and re-applying the recorded decisions deterministically. It calls no model and produces no new side effects, so it is the right tool for debugging and audit. Re-execution calls the model again and may take a different path. Use it to recover from a transient fault, and never to reproduce an incident.

Tool boundaries and the principle of least authority

An agent's blast radius is exactly the union of the authority of its tools. The principle that governs this is fifty years old: every program should operate with the least set of privileges necessary to complete its job (Saltzer & Schroeder, 1975). Applied to tool design, it rules out the tools that are easiest to build. A general execute_sql tool carries the authority of the database role it runs as. A general http_request tool carries the authority of whatever network position it sits in. Both are convenient because nothing limits them.

The alternative is narrow, verb-shaped tools with typed arguments, separate read and write surfaces, per-tool scopes and explicit caps, such as a maximum row count, a maximum spend, or a dry-run mode that returns the intended effect without applying it. Nygard's stability patterns (Nygard, 2007) fill in the rest of the boundary: timeouts on every call, circuit breakers on every dependency, and bulkheads so that one saturated tool does not eat capacity meant for unrelated work.

BrewX, the venue operating system we have in development, is the clearest orchestration surface in our own work. Its data model is a 46-table schema, and no client touches those tables directly. All access goes through more than 80 remote procedures. Four agents work on top of that surface, each with its own remit (promotion and hype, churn prediction, fraud detection, atmosphere scoring), and each is scoped to the procedures that remit needs. The indirection gives us two things. Authority can be enumerated, because the tool inventory is a list of procedure names and not a query language. And validation, authorisation and idempotency sit in one place where they can be enforced, instead of being rebuilt by every caller.

Conway's observation that a system's structure mirrors the communication structure of the organisation that designed it (Conway, 1968) holds here as well. Teams tend to split agents along their own reporting lines and the API surface they inherited, and not along the boundaries of the task. Coordination overhead then shows up at those seams.

Independent verification stages for agent output

Asking a model to check its own output is the weakest verification available, because the checker shares the generator's context, priors and blind spots. The errors are correlated, and correlated checks do not multiply reliability. A verification stage is only worth having if it brings in real independence, and there are three places to get it.

Start with deterministic checks, and use them up before you ask a model anything: schema validation, invariant assertions, recomputed arithmetic, cross-footing a total against its components, confirming that a cited record exists. They are cheap, they have no variance, and they catch a large share of real defects. Next comes adversarial framing. Instead of asking the model to "review this", you ask it to build the strongest case that the output is wrong, and that turns up different defects from a neutral review. The third source is a different lens: a different model, a different prompt frame, or a checker that gets the source evidence but not the candidate answer. That checker produces its own result for you to compare, where a reviewer would only be judging someone else's. When the final check is a person and not a program, the failure mode turns around, as human-in-the-loop design under automation bias describes: reviewers come to ratify machine output instead of testing it.

In our own studio workflow we use the adversarial version on research findings before they leave the building. A second pass argues against the finding from the same sources, and a finding that survives is treated as provisionally sound, but not as proven.

Observability for non-deterministic systems

Conventional logging assumes the execution path can be reproduced, and in an agentic workflow it cannot. For each step you need to capture the deterministic inputs, the resolved context fingerprint, the model identifier and version, the tool schema version, the raw output, the parsed value and the outcome. Without the version triple of prompt, model and tool schema, you cannot attribute a regression, because any of the three may have changed underneath the workflow.

Single-run assertions are the wrong unit of measurement. The useful metrics are distributional: per-step success rate, retry rate, compensation rate, escalation rate, cost and latency per completed task (rather than per call), and how often a step's output changes when its deterministic inputs are identical. That last one is the practical read on variance, and it is the one most often missing. How to turn those distributions into an ongoing judgement of whether the system is getting better or worse is covered in evaluating language-model systems in production.

Practical decision rules for agent orchestration

Allocate autonomy step by step, and default to deterministic control with model judgement at the leaves. Compute pn before you build. If the number is unacceptable, put the effort into the recovery path before you touch the prompt. List the irreversible steps first and order the workflow so they come last, after verification, with compensators written and exercised for every reversible step before them. Derive idempotency keys from deterministic upstream state, never from model output, and keep replay strictly separate from re-execution. Give every tool the narrowest verb that does the job, and read the tool inventory as the full list of things the system could do wrong.

Limitations

This is an architectural argument, and we have not tested it empirically. The reliability arithmetic assumes independence, and real workflows break that assumption in both directions. Correlated failures make chains worse than pn. Tasks with partial credit, where a degraded output is still useful, make the all-or-nothing model too harsh. We make no claim about the accuracy of any particular model, about which framework to use, or about how much these patterns improve results. Those questions need first-party measurement on a specific task distribution. The first-party figures describe systems we have built. They are not benchmark results, and BrewX is still in development, not running at scale. Finally, these patterns come from distributed transaction processing, where failures are mostly independent and detectable. Model failures often are neither, and we have not shown how well compensation-based designs hold up when inputs are systematically wrong and not just randomly wrong.

References

  • Amdahl, G. (1967). Validity of the single processor approach to achieving large scale computing capabilities.
  • Conway, M. (1968). How do committees invent?
  • Garcia-Molina, H., & Salem, K. (1987). Sagas.
  • Gray, J., & Reuter, A. (1993). Transaction processing: concepts and techniques.
  • Kleppmann, M. (2017). Designing data-intensive applications.
  • Nygard, M. (2007). Release it! Design and deploy production-ready software.
  • Saltzer, J., & Schroeder, M. (1975). The protection of information in computer systems.
  • Wei, J., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models.
  • Yao, S., et al. (2022). ReAct: synergizing reasoning and acting in language models.