/lab iterated
Multi-agent LLM experiments
Notes from spinning up multi-agent systems with CrewAI and LangChain — what they're good for and where they fall over.
Hypothesis and baseline
The hypothesis was straightforward: specialist agents with distinct roles should outperform one well-instructed model on research-and-synthesis tasks. I tested that assumption with three shapes:
- A single prompt with a structured output contract.
- One tool-using agent with retrieval and deterministic utilities.
- A multi-agent flow with researcher, synthesizer, and critic roles.
The important control was the single-agent baseline. Without it, added messages and role labels can look like progress even when they only increase token use and latency.
Each variant received the same task packet, source set, output schema, and evaluation rubric. The rubric checked groundedness, task completion, contradictions, citation coverage, invalid tool calls, and whether the run terminated within a fixed step budget. The examples here are illustrative; they do not expose private prompts or evaluation data.
System architecture
Tool-oriented agent experiment
- Task contract Goal, permitted tools, output schema, and stop conditions Product boundary
- Orchestrator Chooses the next bounded action and tracks remaining budget Control boundary
- Tools Retrieval, lookup, calculation, or code execution with typed results Capability boundary
- Evidence store Keeps source identity, provenance, and intermediate facts Grounding boundary
- Evaluator Checks completion, unsupported claims, and contract validity Quality boundary
Termination is an engineering feature
“Continue until done” is not a stop condition. A robust loop has explicit terminal states: completed with valid output, blocked on missing evidence, failed after a non-retryable tool error, or stopped by budget.
type RunState = {
step: number;
maxSteps: number;
evidence: Evidence[];
repeatedActions: Map<string, number>;
};
// Illustrative controller, independent of a specific framework.
function decideNext(state: RunState, proposal: ModelProposal): NextAction {
if (proposal.final && validates(proposal.final, state.evidence)) return { type: "complete" };
if (state.step >= state.maxSteps) return { type: "stop", reason: "step_budget" };
if (!isAllowedTool(proposal.tool)) return { type: "stop", reason: "invalid_tool" };
if ((state.repeatedActions.get(fingerprint(proposal)) ?? 0) >= 2) {
return { type: "stop", reason: "repeated_action" };
}
return { type: "execute", call: proposal.tool };
}
The controller owns retries and limits. The model can propose an action, but it cannot silently expand its permissions, invent a tool, or decide that an infinite loop is productive.
Request sequence
One bounded tool turn
- 01 Orchestratorbuilds current context
Includes the task, compact evidence, remaining budget, and allowed tool schemas.
- 02 Reasonerproposes one action
Returns either a typed tool call, a clarification, or a candidate final answer.
- 03 Policyvalidates proposal
Rejects disallowed tools, malformed arguments, and repeated unproductive actions.
- 04 Tool adapterexecutes and normalizes
Records result, latency, error class, and provenance without prompt-formatted ambiguity.
- 05 Evaluatorchecks terminal contract
Completes, continues, or stops with an explicit reason.
What multi-agent roles actually changed
Separate agents helped when they had genuinely different context or capability: for example, one stage gathered evidence while another reviewed a compact claim-to-source map. Roles were less useful when every agent saw the same context and only rewrote the previous answer.
Tradeoff matrix
Agent topology by task shape
| Option | Strengths | Costs | Decision |
|---|---|---|---|
| Structured single callChosen | Lowest coordination cost and easiest evaluation | Cannot interact with changing external state | Baseline for every task |
| Single tool agent | Can gather evidence and recover from bounded tool errors | Needs termination, policy, and trace design | Best default for tool work |
| Multi-agent pipeline | Separates context and independent review | More latency, variance, and failure surfaces | Use only for measurable role separation |
Evaluation observations
The single structured call won many tasks that required synthesis but no new information. Tool use created the meaningful step change: a calculator removed arithmetic ambiguity, retrieval supplied current evidence, and code execution could test a claim. Adding more conversational agents without different capabilities often added tokens and contradictory intermediate conclusions.
Representative failure cases included:
- Tool loops: the agent repeated a search with semantically equivalent queries.
- Evidence drift: a synthesizer cited a source that supported a neighboring claim, not the stated one.
- Premature completion: a valid-looking schema hid missing required evidence.
- Context flooding: passing every tool result reduced attention to the strongest evidence.
- Critic theater: a critic produced broad objections but no actionable correction contract.
The fix was usually control-plane work: normalize tool results, store evidence separately from prose, cap retries by error class, validate output deterministically, and give review stages exact defects to find.
Bounded proof
Experiment conclusions
- Baseline
- Always required
- Leverage
- Comes from tools
- Safety
- Controller-owned
A multi-agent flow needs to beat a strong single-call contract, not an intentionally weak prompt.
Typed external capabilities changed outcomes more than role-playing alone.
Budgets, permissions, retries, and termination live outside model discretion.
Practical conclusion
I would start with a strong reasoner, a small set of typed specialist tools, an evidence model, and deterministic completion checks. I would add another agent only when it owns a distinct context boundary, capability, or independent evaluation that demonstrably improves the rubric. Framework choice affects ergonomics; the durable architecture is the contract around the model.