All posts
AI Agents

AI Agent Orchestration: Coordinator, Workers, Verifier

AI agent orchestration as I run it: a coordinator splits the job, one worker and one independent verifier per item, a final check, then my approval.

October 3, 2026·7 min read·by Olexander Cheberko
Table of contentstap to expand

AI agent orchestration is the part of a multi-agent setup that splits a job into pieces, hands each piece to an agent and decides what happens to every result. The version I run for this blog's articles and for research has five parts: a coordinator that splits the work, one worker agent per item, an independent verifier per item, one final check across all items, and my approval before anything ships. Everything about Claude Code below was checked against its documentation on October 3, 2026.

What does each part do?

PartIts jobIn one run on October 2, 2026
CoordinatorSplits the job into items, writes each item's brief, sets the orderThe main Claude Code session, which wrote the run as a workflow script
WorkersDo one item each, from the brief alone13 writers, one per new manual, and 5 editors tightening existing drafts
VerifiersCheck one item against outside sources, then fix or cut13 fact-checkers, one per manual, re-reading the vendor's help pages
Final checkLooks across all items at once1 agent rendering all 20 pages in a real browser at desktop and phone width
MeApprove what shipsEvery commit and every push

That is 32 agents for 20 pages.

Why a script and not an agent deciding as it goes?

Claude Code offers both. With subagents, Claude is the orchestrator: it decides turn by turn what to spawn, and every result comes back into its context window. A dynamic workflow moves the plan into a script that holds the loop, the branching and the intermediate results, so Claude's context holds only the final answer, and the script can be read before it runs and saved to run again.

Anthropic's orchestrator-workers pattern has a central model break the task down as it goes, which suits tasks whose subtasks you cannot predict. Content batches are the opposite: the list of articles is known before the run starts. So the split is fixed up front, and the coordinator's real work is a precise brief per item: the sources to use, the facts it may claim, the links it may make.

Why does every item get its own verifier?

A worker that rechecks its own draft does it with the same context that produced the mistake. A Claude Code subagent starts with a fresh, isolated context window and does not see the conversation history, so the fact-checker reads the finished draft cold and re-reads the vendor's help pages itself.

What the fact-checkers caught:

  • A claim about Google Ads "Calls from ads" that missed Google's newer AI rating of recorded calls. Google's help now says that when call recording is on, AI analyzes the recording to determine the lead quality.
  • An outreach hook an agent drafted from a site audit: consent "denied" on a site with no banner. It was true only because the check ran from Bulgaria. Wix's help says that on a site without a cookie banner, no data is sent to Google Analytics when the visitor comes from a country that requires consent, such as those covered by GDPR. A US visitor is not in that group; more in testing consent mode from Europe.
  • A sentence that invented frequency: "the case I see most".

Lauren Tan, an engineer at Cursor, describes the same separation in a recorded talk: when she evaluates a skill, a coordinator agent writes a rubric and spawns sub-agents that run the skill, and a judge agent on a different model can cross-reference the score so the first model's bias does not go unchecked. How I set up the verifier is in subagents that verify work; how to test a verifier's instructions is in evals for skills.

When should items run in a pipeline, and when should you wait for all?

In a Claude Code workflow script, pipeline() runs one agent per item in a list and parallel() runs a set of agent tasks at once and waits for all of them. The script-writing reference Claude Code ships as the /workflow-authoring skill adds that a pipeline has no barrier between stages: one item can be in its second stage while another is still in its first.

My rule follows from that: a step that needs only its own item goes in the pipeline, and a step that compares items waits for all.

  • Pipeline. A manual's fact-check needs only that manual, so the October 2 run chained each writer to its own fact-checker, and a slow article could hold up only its own check. The five editing passes ran alongside.
  • Wait for all. The render check was one agent loading all 20 pages. Some questions have no answer until every item exists, such as whether an FAQ question repeats one in another article.

Simplified, the script had this shape:

const manuals = await pipeline(
  NEW_MANUALS,
  (t) => agent(writePrompt(t), { label: `write:${t.slug}` }),
  (w, t) => agent(verifyPrompt(t), { label: `verify:${t.slug}` }),
)
const render = await agent(renderPrompt(ALL_20_PAGES), { label: 'render:all' })

What does orchestration cost?

Tokens first. Claude Code's docs say a workflow can use meaningfully more tokens than working through the same task in conversation, and they flag a run that schedules more than 25 agents with a "Large workflow" warning. My 32-agent run was past that line. The docs' advice is to run a small slice first. Tan makes the same point in the talk from the trust side: do not jump to hundreds of agents while you do not trust the output of one, because you will just waste a lot of tokens.

Then attention. Without a verifier per item, I would reread every claim myself. In Tan's words, when the agent has no way to verify its work, "you are the verifier": the bottleneck that keeps anything from running in parallel. With verifiers, my attention moves to the briefs before the run and to the results and the commit after it.

Verifiers also miss what they are not asked to check. A fact-checker answers whether a claim matches the docs, not whether the page is worth reading. I rejected early drafts as padded, and the fix was not another agent. It was a fixed manual format (numbered checks with Where, Do, You should see, If not), written into every manual's brief and kept as a memory rule. Rules that need no judgment went to a linter that fails a draft on them: em dashes, links to removed pages, over-long titles, banned phrasing. More in hard checks for agents.

Where does the human approval go?

At the end of a run, before anything ships. Claude Code workflows take no user input mid-run; for a sign-off between stages, the docs say to run each stage as its own workflow. In my setup agents never commit or push on their own: I approve every commit and every push.

Research runs are stricter. When agents audited the public pages of practices I researched for outreach, they used headless Chrome with network capture (gtm.js, the gtag config, conversion requests), read-only, and never submitted a form. One finding had to be withdrawn after the EU-location issue above.

When is one agent enough?

When the job is one item, or when nothing in the result can be checked against an outside source. Anthropic's advice is to find the simplest solution possible and add complexity only when needed, since agentic systems trade latency and cost for better task performance. Verifiers pay off when there are many items of one kind, each with a source to check against.

Where this meets measurement

An orchestrated run is only as trustworthy as the source each verifier checks against, and in marketing that source is usually the tracking itself. If the conversions your Google Ads and GA4 reports count are the part you cannot verify yet, that is the work I do in marketing analytics.

Tags

ai-agent-orchestrationmulti-agent-workflowclaude-codesubagentsverificationhuman-approval

Frequently asked questions

Is AI agent orchestration the same as a multi-agent workflow?

Close enough in practice. A multi-agent workflow is the set of agents and steps; orchestration is the part that decides how the job is split, in what order the pieces run and what happens to each result. In Claude Code that part can be Claude deciding turn by turn with subagents, or a workflow script that holds the plan and can be read and rerun.

How many agents should a multi-agent workflow use?

The number of items times the roles each item needs, and no more. My pattern is one worker and one verifier per item plus one final check, so 13 new articles and 5 edited drafts took 32 agents. Claude Code's default size guideline aims for fewer than 10 agents per dynamic workflow (fewer than 5 on a Pro plan) and flags runs past 25 agents with a "Large workflow" warning, so start with a small slice.

Should agents in a workflow run in parallel or one after another?

Both, decided step by step. A step that needs only its own item, like fact-checking one article, runs as soon as that item's previous step is done, without waiting for the other items. A step that compares items, like a render check of every page or a search for FAQ questions repeated across articles, waits until all items are finished.

Should a human approve AI agent output before it goes live?

In my setup, always. Agents write, check and render, but nothing is committed or pushed until I approve it, every commit and every push. Claude Code workflows take no user input mid-run, so if you want a sign-off between stages, run each stage as its own workflow and approve in between.

Related posts