Multi-Agent Orchestration: Patterns and What They Cost

Every ranking guide on multi-agent orchestration explains the patterns and stops there. None of them asks what it costs to have a model re-derive the same routing decision on every single run. Here are the patterns, and the point where coordination should stop being a reasoning p

Rahul Ramakrishnan
multi-agent-orchestration-hero.png

Key takeaways

  • Multi-agent orchestration coordinates several specialized AI agents across one workflow, deciding which agent runs, in what order, and with what authority.
  • Four pattern families cover most real designs: sequential, concurrent, handoff, and hierarchical supervisor.
  • The bigger design decision is where coordination state lives. A context window forgets. A durable store resumes, and doubles as the audit record.
  • When a model decides coordination, you pay for that decision on every run and can get a different answer on Tuesday than you got on Monday.
  • Coordination logic that has stopped changing is code that has not been written yet.

What multi-agent orchestration actually is

Multi-agent orchestration is the coordination of several specialized AI agents across a single multi-step workflow, deciding which agent runs, in what order, with which tools, and with what authority to act. The orchestration is the coordination layer itself, distinct from the agents it directs.

Two distinctions keep the term precise.

It is narrower than AI orchestration more broadly, which covers models, retrieval, tools, and prompts in any arrangement. Multi-agent orchestration is specifically about agents as the coordinated units.

It differs from workflow automation on one axis: who decides the next step. In workflow automation someone wrote the routing rule in advance and the runtime follows it. In multi-agent orchestration a model can make that call at runtime, reading context and choosing where the work goes. Nearly every production system is a mix of the two, and the real engineering is deciding which steps sit on which side.

That mix is worth naming as two layers. The agent layer reasons, plans, and picks. The app layer is deterministic code where work executes and state lives. Every pattern below is a different answer to how much coordination belongs in each.

Why this matters now

Single agents run out of room on long multi-step work. Context fills up, tool inventories grow past the point where the model chooses well, and one prompt ends up carrying four jobs. More agents is the obvious fix, so teams reach for it.

The primary documentation is noticeably more cautious than the discourse around it. Microsoft's Azure Architecture Center places agent design on a complexity ladder running from a direct model call, to a single agent with tools, to multiagent orchestration, and instructs readers to "use the lowest level of complexity that reliably meets your requirements." It records that the multiagent level "adds coordination overhead, latency, and failure modes." Microsoft's Copilot Studio guidance is blunter: "start with one agent. Then only split into multiple agents when you clearly see a need for modularity or a boundary a single agent shouldn't cross."

That is the honest reason the topic is live. Not that multi-agent systems are better, but that single agents are hitting a ceiling and the industry is working out what to do next.

The coordination patterns worth knowing

Sequential and pipeline coordination

Sequential orchestration chains agents in a predefined linear order, each one processing the output of the last. Azure also files it under pipeline, prompt chaining, and linear delegation.

It fits progressive refinement, the draft, review, polish shape, and stages that genuinely cannot be parallelized. Its failure mode is error propagation. A weak output at stage one becomes the input to stage two, and by stage four you are risk-scoring a contract built on the wrong template. Azure names this directly, advising against the pattern when "early stages might fail or produce low-quality output" with no way to stop later steps consuming the bad result.

Concurrent and parallel coordination

Concurrent orchestration runs several agents simultaneously against the same input, then aggregates their independent results. Fan-out and fan-in, scatter-gather, and map-reduce all describe the same shape.

It fits problems where diverse perspectives on one input beat a single pass, such as four analyst agents scoring the same stock. The failure modes are aggregation and contention. Azure warns against it when agents "can't reliably coordinate changes to shared state" or when there is "no clear conflict resolution strategy" for contradictory results. Reconciling five disagreeing outputs is itself work, and it is frequently another model call.

Handoff and triage routing

Handoff orchestration lets each agent assess the task and decide whether to complete it or transfer control to a better-suited agent. Routing, triage, dispatch, and delegation are the same pattern under other names.

It fits cases where the right specialist is not knowable from the initial input. The failure mode is bouncing, where control loops between agents that each decide someone else should own the ticket. Azure's own advice is the sentence this article is built around: avoid handoff orchestration when "task routing is deterministic and rule-based," or when the right agent "is identifiable from the initial input." In that case, it says, use deterministic routing.

Hierarchical supervisor and worker coordination

Hierarchical orchestration puts a supervisor agent above a roster of workers, delegating subtasks and synthesizing what comes back. It is the pattern most teams picture when they say multi-agent.

Two constraints deserve naming. First, depth is not free, and at least one major platform refuses it outright. Anthropic's managed agents documentation states that "the coordinator can only delegate to one level of agents; referencing an agent that has its own multiagent.agents roster fails the create or update request with a validation error," with a ceiling of 20 unique roster agents. If your design assumes supervisors of supervisors, check whether your platform allows it before the design review.

Second, delegation crosses permission boundaries. Copilot Studio's guidance flags that a connected agent "might have access to things the parent agent doesn't," so a parent forbidden from deleting records can call a child that is allowed to. That is a privilege escalation path wearing an architecture diagram, and it is a good argument for control at the point of action rather than at the top of the tree.

Choosing between them

| Pattern | How coordination is decided | Where state lives | Best when | Main failure mode | | --- | --- | --- | --- | --- | | Sequential | Fixed order set at design time | Passed forward in context, stage to stage | Stages must build on each other | Errors propagate downstream unchecked | | Concurrent | Fan-out decided up front, aggregation at the end | Split across parallel contexts, merged at join | Independent perspectives on one input | Conflicting results with no resolution rule | | Handoff / triage | The model, per run, per turn | Transferred with control, often partially | The right specialist is unknown up front | Bouncing and loops between agents | | Hierarchical supervisor | The supervisor model, per run | Supervisor context, plus isolated worker threads | Cross-domain work with clear subtasks | Supervisor becomes a cost and failure bottleneck | | Code-orchestrated pipeline | Deterministic code, written once | A durable store outside any context window | Routing rules have stopped changing | Rigid when requirements genuinely shift |

Where the state lives, and why that is the real design decision

Take a support triage flow. A supervisor agent classifies an inbound ticket and routes it to a billing agent, a technical agent, or an escalation path. Simple enough to draw, and it is the most common multi-agent system in production anywhere.

Now ask where the ticket state lives. In the default design, it lives in context windows. The supervisor holds the classification, the worker holds the conversation, and the platform holds the threads for the life of the session. Anthropic's model is more durable than most, with persistent context-isolated session threads, so a coordinator can follow up with an agent and that agent still remembers its earlier turns.

Three things break when coordination state exists only in context. A failure mid-run leaves nothing to resume from, so the work restarts. There is no continuity between runs, so ticket 400 from the same customer arrives as a stranger. And every run pays to refill the context that the last run already assembled.

Dataiku's agent orchestration guide handles this well, describing short-term, long-term, and episodic memory tiers, and separately calling for an "immutable audit trail" of actions, tool calls, and data access. Those are treated as two concerns. They are one. A durable coordination store, recording which agent received which ticket, under whose credentials, and what it did, is the memory and the audit log at the same time. Build it for resumption and you get agent observability for free, or build it for the auditor and your workflow resumes after a crash.

What orchestration costs on run one thousand

Coordination decided by a model is a recurring bill. The triage supervisor classifies ticket one, and then it classifies ticket one thousand, spending tokens both times on a decision whose rules stopped moving somewhere around ticket two hundred.

Nobody in the ranking set quantifies this. The OpenAI Agents SDK gets closest, splitting the whole topic into orchestrating via LLM and orchestrating via code, and stating that "orchestrating via code makes tasks more deterministic and predictable" on "speed, cost and performance." It offers no number, no benchmark, and no rule for where the boundary sits. Dataiku prescribes monitoring "token usage per agent" and enforcing per-agent budget caps, which measures the bill without questioning why the line item recurs. IBM's page on the topic does not use the word token at all.

So the argument here is structural rather than benchmarked, and I am not going to manufacture a figure to fill the gap. The structure is this. Model-decided coordination costs in proportion to usage, because each run re-derives the decision. Code-decided coordination costs once to establish and then flattens, because the decision is already written down. Front-loaded and then flat, against climbing with volume. If you want the tactical layer underneath this, caching and model selection are covered in what actually reduces LLM cost.

Cost is only half of it. Re-derivation is also why the same ticket can take a different path this week than it took last week.

| Dimension | Model decides | Code decides | | --- | --- | --- | | Cost as usage grows | Climbs with every run | Front-loaded, then flat | | Behaviour on repeat runs | Can vary for identical input | Identical for identical input | | Auditability | Reconstructed from transcripts and traces | Read the routing rule and the log | | Flexibility when requirements change | Adapts immediately, no deploy needed | Needs a change to the code |

That last row favours the model, honestly and unambiguously. If your routing genuinely changes week to week, a model deciding it is the right call and hard-coding the rule is the wrong one.

When coordination should stop being a reasoning problem

Three questions per coordination step, answerable from your own logs.

First, over the last few hundred runs, how often did this step reach a different decision for the same shape of input? If the answer is close to zero, the step is a lookup dressed as reasoning.

Second, can you write the rule down? If you can state it in a sentence a colleague would sign off on, it is a rule, and a rule belongs in code.

Third, does the decision need information that only exists at runtime and cannot be enumerated in advance? If yes, keep it with the model. An ambiguous ticket matching no rule is exactly what judgment is for.

In the triage flow, that split is concrete. The routing table, the ticket state, and the handoff record move into code. The genuinely ambiguous ticket goes to the model, which is also where the next rule comes from. That is the shape of building an agentic workflow that survives its own success.

Common misconceptions

More agents does not mean more capability. Copilot Studio warns that assigning two subagents the same knowledge source produces one useful answer and one redundant search.

A supervisor agent is not a scheduler. A scheduler is deterministic and costs nothing per decision. A supervisor is an inference call with an opinion.

Stateful is not a long context window. A long context is a bigger buffer that still empties. State is a record that outlives the run.

Parallel agents do not automatically finish faster. Azure caps group chat at "three or fewer agents" to keep control manageable, and Copilot Studio notes the "slightly longer execution time due to context switching" that separate agents add. Fan-out buys wall-clock time only when the aggregation step is cheaper than the work it saved.

What we are building at Major in response

The position: coordination that has stopped changing should stop being a model call. Most multi-agent systems spend their largest recurring cost re-deriving decisions that were settled weeks ago, and that same re-derivation is why identical inputs can take different paths on different days.

Major is the enterprise platform where agents build the software they run on. When an agent works out how to handle a repeatable part of a task, it builds an app for that part and then runs the app instead of reasoning through the step again. The triage routing table becomes code. The ticket state and the handoff record live in the app's managed database rather than in a context window that ends when the session does. Because the work runs as a deployed app, each action carries scoped credentials and writes its own audit record, so what each agent did is inspectable and attributable rather than reconstructed after the fact from a transcript.

The model is still there, and it still does the hard part. Judgment calls go to it, and it is what notices the next piece of coordination that has settled. Only the settled part moves into code. Reason once, run forever.

This argument does not cover agent evaluation, which is a harder problem and deserves its own piece. If you are working out which of your coordination steps have settled and which still need judgment, and what enterprise-grade actually means for the ones you move into code, see how Major turns settled coordination into governed apps.

Related articles

Frequently asked questions

What is multi-agent orchestration?
Multi-agent orchestration is the coordination of several specialized AI agents across a single multi-step workflow, deciding which agent runs, in what order, and with what authority to act. Four pattern families cover most designs: sequential pipelines, concurrent fan-out, handoff routing between specialists, and a hierarchical supervisor delegating to workers.
What is the difference between multi-agent orchestration and workflow automation?
The difference is who decides the next step. Workflow automation follows a routing rule someone wrote in advance. Multi-agent orchestration lets a model make that call at runtime, reading context and choosing where work goes. Most production systems combine both, and the real engineering decision is which steps sit on which side of that line.
When should you use multiple agents instead of one?
Use multiple agents when subtasks are genuinely parallelisable, when agents need different tool permissions or security boundaries, or when contexts must stay isolated to keep output quality up. Microsoft's Copilot Studio guidance says to start with one agent and split only when a boundary appears that a single agent should not cross. Most teams should follow that advice.
Where does state live in a multi-agent system?
By default it lives in context windows, held by each agent for the life of the session. That breaks in three ways: a mid-run failure leaves nothing to resume from, there is no continuity between runs, and every run pays to refill context the last run already assembled. A durable store outside the agents fixes all three, and records who did what.
Does multi-agent orchestration increase cost?
Usually yes, through two mechanisms. Coordination itself consumes inference, since a supervisor or routing agent is a model call. And that call repeats on every run, re-deriving decisions whose rules stopped changing long ago. Move settled coordination into code and cost is front-loaded then flat. Leave it with the model and cost climbs with volume.