Skip to main content

Command Palette

Search for a command to run...

Multi-Agent Orchestration: Architecture Patterns We Use at Buda

Updated
•14 min read•View as Markdown
Multi-Agent Orchestration: Architecture Patterns We Use at Buda

Most multi-agent demos fall apart the moment you give them a second task. The first one runs beautifully: a planner agent dispatches a researcher, the researcher hands a draft to a writer, the writer returns a polished result, and the conference-talk applause rolls in. Then you run it again with a different input and discover that "orchestration" was really just a prompt chain with a for loop around it. There is no isolation between runs, no durable place for one agent to leave knowledge for another, and no scheduler deciding what should actually execute. The graph diagram on the slide was aspirational.

We have spent a while building the orchestration layer behind Buda, a Drive-based agent platform for teams, and the lessons that actually moved reliability were almost never about the LLM. They were about the boring distributed-systems questions: where does state live, what is the unit of isolation, who decides what runs next, and how do two agents exchange work without corrupting each other's context. This post walks through the patterns we converged on. None of them are exotic. The value is in how they compose.

The mental model: agents as employees, not functions

The first architectural decision is how you conceptualize an agent. If you treat an agent as a function — input prompt, output text — every collaboration becomes argument-passing, and the only shared state is whatever you can cram into a context window. That model breaks down fast, because the interesting work in real organizations is not stateless. It accumulates.

We model the system around a company analogy, and it is more than a marketing metaphor — it dictates the boundaries in the code:

  • A Space is the company. It owns members, permissions, billing, and shared storage. Isolation between Spaces is hard: different Spaces are different tenants, billed separately, with no implicit data flow between them.

  • An Agent is an employee. It has a stable identity, instructions (its system prompt and operating constraints), its own private file storage, a set of skills, and any number of concurrent conversations.

  • A Drive is the employee's filing cabinet — long-term, durable memory that survives across every conversation.

  • A Session is a meeting room: an isolated, short-term context for one task or one conversation stream.

  • A Skill is a reusable SOP — a method the agent can invoke, not another agent.

The reason this matters for orchestration is that it gives you two clearly different storage tiers with different lifecycles, and a clean unit of concurrency. When people say their multi-agent system is "hard to debug," it is usually because they have collapsed these into one blob. Everything is context, context is ephemeral, and so nothing is auditable.

Sessions: isolation as the default, not an afterthought

The single most important pattern is treating the Session as the unit of isolation. One Agent can have many Sessions, and by default they do not see each other.

Consider a support agent connected to WhatsApp. Two customers message it within the same minute. If their turns share a context window, you get cross-contamination: customer B's order number bleeds into customer A's reply, or worse, private data leaks across the boundary. The fix is not a clever prompt — it is structural. Each inbound user gets a dedicated Session, and the channel layer is responsible for routing turns to the correct one.

We encode the boundaries explicitly per channel, because the natural grain differs:

Channel surface Session boundary
WhatsApp DM One Session per phone number
WhatsApp group One Session per group, shared by members
Telegram DM One Session per user
Discord One Session per channel

The general rule: a Session maps to a conversation stream, and a conversation stream is whatever the outside world treats as a single coherent thread. Get this mapping wrong and no amount of model quality saves you.

There is a second, subtler reason Sessions matter for orchestration: they bound the blast radius of context pollution. Long-running agents accumulate cruft — half-finished reasoning, abandoned tool calls, stale instructions. If everything lives in one ever-growing context, quality degrades and you cannot tell why. By scoping context to a Session, you can reset a task without wiping the agent's knowledge. The agent forgets the meeting; it does not forget the filing cabinet. That separation between "what happened in this conversation" and "what the agent durably knows" is the line most naive designs never draw.

Drive: the shared substrate that makes hand-offs durable

Here is the question that exposes whether a multi-agent design is real: when agent A finishes and agent B takes over, what exactly gets handed off?

The tempting answer is "the conversation history" or "a summary in the next prompt." Both are fragile. Conversation history is large, noisy, and ephemeral. A summary is lossy and lives nowhere you can inspect later. Neither survives a restart, and neither lets a third agent — or a human — audit what was actually produced.

Our answer is that durable hand-offs go through files. An agent's Drive is its long-term memory: SOPs, policies, research notes, contracts, and — critically — the artifacts it generates. When a research agent finishes, it does not just emit text into a Session; it writes a structured brief to Drive. The writer agent reads that brief as a file. The output is a file. The reviewer reads the file. State is concrete, named, and persistent.

This is the core architectural bet, and it is why "Drive-based" is in the platform's name rather than being a feature footnote. Chat history is not durable knowledge. The moment a piece of context matters beyond the current Session, it should be written to Drive, where it has a path, survives restarts, and is visible to whatever needs it next. If you have built RAG systems, this will feel familiar — except instead of a vector index that agents query indirectly, the file system is the shared workspace, and agents read and write it the same way a developer would.

Drive also gives you a natural permission surface. A customer-facing support agent should read from Drive but never mutate the canonical knowledge base, so you separate public files from internal ones and lock the agent to read-only on the source material. (A real caveat we ship with: the workspace tools can still create files in the agent's own scratch space, so read-only has to be enforced through both instructions and tool permissions, not assumed.) The point is that because state lives in a file system, your access model is one you already know how to reason about, instead of an opaque "what's in the context right now."

Teams: hand-offs as explicit topology

Once isolation (Sessions) and durable state (Drive) exist, multi-agent collaboration becomes tractable. We model a multi-agent workflow as a Team: a group of Agents with distinct roles and Skills that hand work off to each other.

The design discipline that makes Teams work is resisting the urge to build one omniscient agent. A single agent with twenty tools and a 4,000-word system prompt is a worse system than five focused agents, for the same reasons a monolith is worse than well-bounded services: the prompt becomes a contended resource, tool selection degrades as the option set grows, and you cannot reason about failure in isolation. So we split by role boundary. Create a new Agent when the identity, knowledge base, tone, routing, permissions, or skill set genuinely needs to be separate — not before.

A concrete shape we see often:

  1. An intake agent classifies an incoming request and writes a normalized task spec to shared Space storage.

  2. A specialist agent (research, drafting, code) picks up the spec, does the work in its own isolated Session, and writes artifacts to Drive.

  3. A review agent reads those artifacts, checks them against policy files, and either approves or writes feedback.

  4. A human, or the intake agent, closes the loop.

Notice what the hand-off is: not a function call passing a giant context object, but a file appearing in a shared location. That is deliberate. File-based hand-offs are inspectable (you can open the artifact), replayable (re-run the writer against the same brief), and resilient (a crash mid-pipeline leaves the completed stages intact on disk). It is the same instinct as preferring durable message queues and idempotent consumers over tightly-coupled synchronous RPC chains. The agents are loosely coupled through a shared substrate rather than through direct, stateful calls.

This is also where teams of agents differ from a single agent with many skills. A Skill is a method one agent invokes inline — generate a slide deck, rewrite social copy, parse a spreadsheet. A Team is a set of independent runtime identities that coordinate. Knowing which abstraction you reach for is half the battle: repeated methods become Skills; separate responsibilities become Agents in a Team.

The Organizer: a scheduling layer, not a megaprompt

The pattern that most distinguishes a production orchestration system from a demo is having a real scheduler. In the demo, "what runs next" is hardcoded in control flow or, worse, decided by asking the model to pick. Neither survives contact with concurrency, retries, scheduled jobs, or long-running asynchronous work.

Buda separates two architectural layers, and the separation is the whole point:

  • Claw Computer is the compute layer: isolated, durable, scalable runtime environments where agents actually execute. Think of it as where the sandbox — terminal, browser, file system, Git — lives for each agent.

  • Buda Organizer is the scheduling layer: it decides what runs, when it runs, and how tasks are orchestrated across agents.

Keeping scheduling out of the agent's reasoning loop is what makes the system an agent runtime rather than a model wrapper. The model decides how to do a task; the Organizer decides whether and when a task should run at all. That distinction lets you support things a pure prompt-chain cannot:

  • Scheduled work. A monitoring agent that wakes on a cron-like trigger, checks a data source, and only escalates on a signal. The model is not "always thinking" — the scheduler invokes it.

  • Asynchronous long-running tasks. Video generation, a large scrape, a multi-minute build. The job runs in the background on the compute layer while the user keeps interacting; the scheduler tracks completion and routes the result back.

  • Concurrency control. Many Sessions, many agents, finite compute. Something has to decide ordering and resource allocation, and you do not want that something to be a language model improvising.

If you are designing your own system, the takeaway is to make the scheduler a first-class component with its own state, distinct from agent reasoning. The failure mode of conflating them is that your orchestration logic becomes non-deterministic and untestable, because it is buried inside generations. Pull it out. The agent should be a worker the scheduler dispatches, not the dispatcher itself.

Compute isolation: why cloud-native changes the design

A design choice that ripples through everything else is where agents run. There is a healthy ecosystem of self-hosted, local-hardware agent setups — projects in the spirit of OpenClaw, where a developer runs the whole stack on their own machine, maybe a Mac Mini under the desk. That approach has real strengths: total control, local data, and a hacker-friendly feedback loop where you own every layer. For a single power user, it can be ideal, and we respect the model.

But it pushes against multi-agent orchestration in specific ways. One machine is one failure domain and one concurrency ceiling. Isolating ten agents' sandboxes from each other on local hardware is your problem to solve. Scheduling across them, surviving a reboot mid-task, and giving five teammates shared access all become infrastructure projects in their own right.

Running on a cloud compute layer inverts those constraints. Each agent gets its own isolated, durable runtime; the scheduler allocates across a pool rather than fighting over one box; and isolation between agents is enforced by the platform instead of by you. The trade-off is real and worth stating plainly: you give up some control and local-data ownership in exchange for isolation, durability, and team-scale concurrency that you would otherwise have to build and operate yourself. For a team that wants orchestration to be a property of the platform rather than a side project, an AI agent platform that provisions per-agent cloud sandboxes removes an entire category of undifferentiated infrastructure work. That is the axis to weigh: single-user control versus team-scale, hands-off concurrency.

Channels: entry points, not memory

A pattern worth calling out because it is so commonly misdesigned: the boundary between how a user reaches an agent and where knowledge lives.

An agent can be reachable through web, Slack, WhatsApp, Telegram, Discord, Feishu, WeCom, Microsoft Teams, or a REST/OpenAPI integration. We call these Channels, and the architectural rule is that a Channel is an entry point, not a memory layer. Its job is narrow: receive an inbound message, map the sender to the correct Session, deliver the reply back to the right surface, and keep the connection alive. That is it.

The temptation is to let channel-specific state creep into your knowledge model — to treat "what this Slack thread said" as durable context. Resist it. Durable knowledge belongs in Drive; stable behavior belongs in agent instructions; the Channel is plumbing. Keep that boundary clean and you can add or swap surfaces without touching your orchestration logic. Blur it and every new integration becomes a special case in your core.

Putting the layers together

Step back and the architecture is a set of cleanly separated concerns, each answering one of the questions a demo never has to:

Concern Component The question it answers
Tenant / org boundary Space Who owns this, and who pays?
Runtime identity Agent What role does this worker play?
Durable state Drive What survives across tasks and restarts?
Isolation Session What is the unit of concurrent, non-leaking context?
Reusable method Skill What inline capability does an agent invoke?
Collaboration topology Team How do responsibilities hand off?
Compute Claw Computer Where does work actually execute, isolated and durable?
Scheduling Organizer What runs, when, and in what order?
Entry point Channel How does the outside world reach an agent?

The reason to draw these lines explicitly is that each becomes an independent axis you can reason about, test, and scale. You can change the scheduling policy without rewriting agent prompts. You can add a Channel without touching state. You can reset a Session without losing Drive. You can scale compute without redesigning the topology. That decomposability is the entire payoff — it is what turns a fragile prompt chain into a system you can operate.

Lessons that generalize

If you are building your own multi-agent system on a different stack, the platform names do not matter, but a few principles transfer cleanly:

  • Separate ephemeral context from durable knowledge, physically. Conversation history and long-term memory should not share a storage tier. The clearest version of this is a real file system as the shared substrate, where artifacts have paths and survive restarts.

  • Make isolation the default. Scope context to a task or conversation stream. Sharing should be an explicit decision through a shared store, never an accidental side effect of one big context window.

  • Hand off through artifacts, not arguments. File-based hand-offs are inspectable, replayable, and crash-resilient in a way that passing context objects between calls never will be.

  • Pull scheduling out of the model. "What runs next" is a systems decision with its own state, not something to delegate to a generation. Treat the agent as a dispatched worker.

  • Split by responsibility, not by capability. Many small, well-bounded agents that coordinate beat one omniscient agent with a sprawling prompt — the same reason bounded services beat a monolith.

None of this requires a frontier model or a clever prompting trick. It requires treating a multi-agent system as what it actually is: a small distributed system with language models as the workers. The orchestration patterns that hold up are the ones we have trusted in distributed systems for years — isolation, durable state, loose coupling through a shared substrate, and a scheduler that is not also a worker. The LLM is the interesting part of the story, but it is rarely the part that determines whether your system survives its second task.

2 views

agent

Part 1 of 3

Welcome to our AI Agent Testing Lab. Here, we dive deep into the rapidly evolving world of autonomous artificial intelligence. From intelligent coding assistants to complex, multi-step task runners, we rigorously test, review, and benchmark the latest AI agent products on the market. Join us as we explore their real-world capabilities, analyze their limitations, and discover how these autonomous systems are redefining human-AI collaboration.

Up next

How We Built Persistent Memory Into an AI Agent Workspace

Every team that ships an agent eventually hits the same wall: the agent forgets. Not in the dramatic sci-fi sense, but in the mundane, frustrating way where it re-asks a question it answered last Tues

More from this blog

A

Ai Tools List

20 posts