/lab iterated

Multi-agent LLM experiments

Notes from spinning up multi-agent systems with CrewAI and LangChain — what they're good for and where they fall over.


Hypothesis and baseline

The hypothesis was straightforward: specialist agents with distinct roles should outperform one well-instructed model on research-and-synthesis tasks. I tested that assumption with three shapes:

  1. A single prompt with a structured output contract.
  2. One tool-using agent with retrieval and deterministic utilities.
  3. A multi-agent flow with researcher, synthesizer, and critic roles.

The important control was the single-agent baseline. Without it, added messages and role labels can look like progress even when they only increase token use and latency.

Each variant received the same task packet, source set, output schema, and evaluation rubric. The rubric checked groundedness, task completion, contradictions, citation coverage, invalid tool calls, and whether the run terminated within a fixed step budget. The examples here are illustrative; they do not expose private prompts or evaluation data.

System architecture

Tool-oriented agent experiment

  1. Task contract Goal, permitted tools, output schema, and stop conditions Product boundary
  2. Orchestrator Chooses the next bounded action and tracks remaining budget Control boundary
  3. Tools Retrieval, lookup, calculation, or code execution with typed results Capability boundary
  4. Evidence store Keeps source identity, provenance, and intermediate facts Grounding boundary
  5. Evaluator Checks completion, unsupported claims, and contract validity Quality boundary
The useful abstraction was a controlled reasoner around specialist tools, with evidence and policy checked outside the model.

Termination is an engineering feature

“Continue until done” is not a stop condition. A robust loop has explicit terminal states: completed with valid output, blocked on missing evidence, failed after a non-retryable tool error, or stopped by budget.

type RunState = {
  step: number;
  maxSteps: number;
  evidence: Evidence[];
  repeatedActions: Map<string, number>;
};

// Illustrative controller, independent of a specific framework.
function decideNext(state: RunState, proposal: ModelProposal): NextAction {
  if (proposal.final && validates(proposal.final, state.evidence)) return { type: "complete" };
  if (state.step >= state.maxSteps) return { type: "stop", reason: "step_budget" };
  if (!isAllowedTool(proposal.tool)) return { type: "stop", reason: "invalid_tool" };
  if ((state.repeatedActions.get(fingerprint(proposal)) ?? 0) >= 2) {
    return { type: "stop", reason: "repeated_action" };
  }
  return { type: "execute", call: proposal.tool };
}

The controller owns retries and limits. The model can propose an action, but it cannot silently expand its permissions, invent a tool, or decide that an infinite loop is productive.

Request sequence

One bounded tool turn

  1. 01
    Orchestratorbuilds current context

    Includes the task, compact evidence, remaining budget, and allowed tool schemas.

  2. 02
    Reasonerproposes one action

    Returns either a typed tool call, a clarification, or a candidate final answer.

  3. 03
    Policyvalidates proposal

    Rejects disallowed tools, malformed arguments, and repeated unproductive actions.

  4. 04
    Tool adapterexecutes and normalizes

    Records result, latency, error class, and provenance without prompt-formatted ambiguity.

  5. 05
    Evaluatorchecks terminal contract

    Completes, continues, or stops with an explicit reason.

Every turn produces inspectable state. The final answer is accepted only after deterministic contract checks.

What multi-agent roles actually changed

Separate agents helped when they had genuinely different context or capability: for example, one stage gathered evidence while another reviewed a compact claim-to-source map. Roles were less useful when every agent saw the same context and only rewrote the previous answer.

Tradeoff matrix

Agent topology by task shape

OptionStrengthsCostsDecision
Structured single callChosen Lowest coordination cost and easiest evaluationCannot interact with changing external stateBaseline for every task
Single tool agent Can gather evidence and recover from bounded tool errorsNeeds termination, policy, and trace designBest default for tool work
Multi-agent pipeline Separates context and independent reviewMore latency, variance, and failure surfacesUse only for measurable role separation
The simplest topology that passes the evaluation is usually the best production starting point.

Evaluation observations

The single structured call won many tasks that required synthesis but no new information. Tool use created the meaningful step change: a calculator removed arithmetic ambiguity, retrieval supplied current evidence, and code execution could test a claim. Adding more conversational agents without different capabilities often added tokens and contradictory intermediate conclusions.

Representative failure cases included:

  • Tool loops: the agent repeated a search with semantically equivalent queries.
  • Evidence drift: a synthesizer cited a source that supported a neighboring claim, not the stated one.
  • Premature completion: a valid-looking schema hid missing required evidence.
  • Context flooding: passing every tool result reduced attention to the strongest evidence.
  • Critic theater: a critic produced broad objections but no actionable correction contract.

The fix was usually control-plane work: normalize tool results, store evidence separately from prose, cap retries by error class, validate output deterministically, and give review stages exact defects to find.

Bounded proof

Experiment conclusions

Baseline
Always required

A multi-agent flow needs to beat a strong single-call contract, not an intentionally weak prompt.

Leverage
Comes from tools

Typed external capabilities changed outcomes more than role-playing alone.

Safety
Controller-owned

Budgets, permissions, retries, and termination live outside model discretion.

These conclusions compare architecture behavior; they do not claim a universal benchmark advantage.

Practical conclusion

I would start with a strong reasoner, a small set of typed specialist tools, an evidence model, and deterministic completion checks. I would add another agent only when it owns a distinct context boundary, capability, or independent evaluation that demonstrably improves the rubric. Framework choice affects ergonomics; the durable architecture is the contract around the model.