1. An agent starts small
Take accounts payableAccounts payableThe function that receives, checks and pays supplier invoices.See the glossary . A first agent reads invoices, matches them to purchase ordersPurchase orderThe document formalising a purchase from a supplier: items, quantities, prices. The invoice must match it.See the glossary , flags discrepancies and proposes an approval. Four tools, a clear scope, an easy-to-check result: a good start.
Then the requests come: handle disputes, chase suppliers, check payments, apply the purchasing policy, produce the reporting. Each addition is reasonable. Their sum is not.
agent: supplier-invoice-assistant
tools:
# the starting scope
- read-invoice
- match-purchase-order
- flag-discrepancy
- propose-approval
# added over the months
- open-dispute
- chase-supplier
- check-payments
- schedule-payment
- search-purchasing-policy
- update-supplier
- generate-reporting
# ...2. When one agent does too much
- It becomes fragile
- Instructions grow longer to limit errors, every small change means retesting everything, and nobody dares touch it any more.
- It makes mistakes
- It calls the wrong tool, passes the wrong parameters, and answers the same question differently.
- It costs more
- It needs the most powerful models, its instructions grow with every call and it repeats steps. According to Anthropic, 58 tools alone take about 55,000 tokensTokenTwo meanings. For a model: a piece of a word, the unit that measures processed text and therefore cost. In security: a temporary key proving an access right.See the glossary before the first question.
3. Split into specialised agents
The answer is the same as in software: split. One agent for invoices, one for disputes, one for payments. Each has few tools, short instructions and rights limited to its domain.
- Each agent is simpler to write, test and evolve.
- Each agent is more reliable: fewer possible choices, fewer chances to go wrong.
- Independent tasks can run in parallel, which shortens lead times.
It is not free. Anthropic measures that a multi-agent system uses about 15 times more tokens than a simple chat. Their advice: find the simplest solution possible, and only add complexity when it is needed.
4. A main agent that plans and delegates
Splitting into independent agents has a flaw: each sees only part of the context, and their decisions can conflict. Hence the approach spreading since 2025, popularised by LangChain’s Deep Agents and inspired by Claude Code, Manus and Deep Research: a main agent keeps control of the context and relies on four ingredients.
- Detailed instructions
- A long, precise system promptSystem promptThe standing instructions given to an agent: its role, its rules, how it works.See the glossary describing the business, the expected steps and the rules to follow.
- A plan
- A task list the agent keeps up to date: it forces structure and shows where the agent stands.
- A notes space
- A file system where the agent stores intermediate results, instead of keeping everything in its working memory.
- Subagents
- Precise tasks handed to subagents with isolated context; they return a result, the main agent decides.
The principle: one head that decides, hands that execute. You keep the benefits of splitting without losing the thread of the context.
5. Coordination patterns
A main agent that delegates is a form of supervisor. But depending on the process, other ways of making agents work together remain relevant. FrameworksFrameworkA development toolkit that provides an application’s structure; for agents: LangGraph, Strands, CrewAI…See the glossary and vendors use different names, but the same families appear everywhere.
- WorkflowFixed steps, parallel when possible
- SupervisorOne agent delegates to specialists
- HierarchicalPlanner, supervisors, agents
- GraphBranches and loops
- SwarmAgents hand off to each other
| Pattern | Principle | When to use it | Watch out for |
|---|---|---|---|
| WorkflowWorkflowA sequence of steps defined in advance. Some steps can be handed to an AI model, but the order stays fixed.See the glossary | Fixed steps defined in advance; independent tasks run in parallel. | Known, repetitive process: matching, checks, closing. | Rigid when facing unforeseen cases. |
| Supervisor | One agent receives the request and calls specialised agents as tools. | Varied requests within one domain. | The supervisor must know each agent well; it is a single checkpoint. |
| Hierarchical | A planner agent splits the request and hands each part to a domain supervisor. | Processes spanning several departments. | Delays and costs add up at each level. |
| Graph | Agents linked by explicit transitions; the model picks the branch, with loops allowed. | Processes with branches: check, correct, re-check. | Bound the loops so they cannot run forever. |
| Swarm | Agents hand off to each other on their own, with shared memory. | Open research, exploring an ill-defined problem. | The least predictable: traceabilityAudit logThe record of who did what, when and with which data. It lets you check and explain every action.See the glossary is essential. |
In real life you combine them: a workflow for the backbone of the process, a supervisor at the step that needs judgement, and human approval before any write.
6. Middleware and hooks: where governance plugs in
An agent runs in a loop: it queries the model, calls a tool, reads the result, starts again. Recent frameworks let you step in at each point of that loop without touching the agent’s core: that is middleware and hooks. This is where controls belong.
- Request
- Before the agentrights, context
- ↻ Agent loop, repeated until the answer
- Before the modelmasking, summary
- Model callaround: cap, fallback
- Around the toolhuman approval, rights
- Tool call
- After the agentlog, evaluation
- Answer
| When | What we plug in |
|---|---|
| Before the agent | Check the request and the user’s rights, load the right context. |
| Before each model call | Mask personal data, summarise an overlong context, pick the useful tools. |
| Around the model call | Call cap, retry, fallback to another model on error. |
| Around each tool call | Human approval before a write, rights check, call cap, block or fix parameters. |
| After the agent | Log the run, evaluate the answer, trigger a review. |
The pattern is now the same at every major player, under different names:
| Framework | Mechanism |
|---|---|
| LangChain and Deep Agents | Middleware (before and after the agent or model, around model and tool calls), including ready-made ones: human approval, personal data, caps, model fallback, context summarisation |
| Claude Agent SDK (Anthropic) | Hooks: before and after each tool (to allow, deny, ask or modify), on promptPromptThe written instructions given to the AI model to steer its answer.See the glossary submit, when a subagent stops |
| OpenAI Agents SDK | Input, output and tool guardrailsGuardrailsAutomatic checks that block risky use: sensitive data leaks, forbidden content, out-of-scope actions.See the glossary that stop the run, plus lifecycle hooks |
| Strands Agents (AWS) | Hooks, for example before a tool call to cancel it or rewrite its parameters |
| Google ADK | Callbacks before and after the agent, the model and each tool |
Governance becomes code: one library of approved middleware, reused by every agent, rather than rules rewritten in each prompt.
7. What we build
- Autonomous agents
- Agents that chain several steps of a process: read, analyse, prepare, update, stopping at approval points.
- Talk to my data
- Ask questions in everyday language about company data (Snowflake, Databricks, business applications) and get quantified, sourced answers, within each person’s rights.
- Decision support
- Reasoned recommendations, with their sources and confidence level, ready for an expert to approve.
- Runtime isolation
- Each agent runs in its own environment, with its own rights: an error or misuse stays contained within its scope.
- Business connectors
- ERPERPEnterprise resource planning software: finance, purchasing, inventory, production.See the glossary , CRMCRMCustomer relationship management software: contacts, opportunities, interaction history (Salesforce, for example).See the glossary , ITSMITSMIT service management: tickets, incidents, requests and changes (ServiceNow, for example).See the glossary , finance tools: agents act inside existing applications, through MCP serversMCP serverThe small service that exposes a tool (an application, a document base) to agents through the MCP protocol.See the glossary governed by the AI Platform.
8. Guardrails specific to business agents
- Rights limited to each agent’s domain, never an account that “can do everything”.
- Human approval before any write into a system of record.
- Approved, versionedVersioningKeeping every change to a file with its author and date, so you can compare, review and roll back.See the glossary reference sources: policies, price lists, procedures.
- A full trace of every run: which data, which tools, which decision.
- EvaluationEvaluationTesting an agent on a set of known cases to measure the quality of its answers before updating it.See the glossary on real cases before every update of an agent or a model.
- A cost cap per agent and per process.
- The ability to suspend an agent in a single operation.
9. Where to start
- Pick a frequent, measurable process where errors are easy to spot.
- Start with a single agent, few tools, read-only.
- Plug in the control middleware from day one: human approval, personal data, caps, logging.
- When the agent grows, move to a main agent that plans and delegates to subagents.
- Choose the simplest coordination pattern that works, often a workflow.
- Add writes into systems only after human approval and quality measurement.
Further reading
- Anthropic — Building effective agents (opens in a new tab)
- Anthropic — How we built our multi-agent research system (opens in a new tab)
- Anthropic — advanced tool use (tool search) (opens in a new tab)
- Microsoft — AI agent orchestration patterns (opens in a new tab)
- Strands Agents — multi-agent patterns (opens in a new tab)
- Strands Agents — agents as tools (opens in a new tab)
- Strands Agents — swarm (opens in a new tab)
- Strands Agents — graph (opens in a new tab)
- Strands Agents — workflow (opens in a new tab)
- A2A — specification (opens in a new tab)
- LangChain — Deep Agents (opens in a new tab)
- LangChain — Deep Agents documentation (opens in a new tab)
- LangChain — Deep Agents harness (opens in a new tab)
- LangChain — LangChain and LangGraph 1.0 (opens in a new tab)
- LangChain — built-in middleware (opens in a new tab)
- LangChain — custom middleware (hooks) (opens in a new tab)
- Claude Agent SDK — hooks (opens in a new tab)
- OpenAI Agents SDK — guardrails (opens in a new tab)
- OpenAI Agents SDK — lifecycle hooks (opens in a new tab)
- Strands Agents — hooks (opens in a new tab)
- Google ADK — callbacks (opens in a new tab)
- Cognition — Don’t build multi-agents (opens in a new tab)
- LangChain — How and when to build multi-agent systems (opens in a new tab)