In earlier pieces about multi-agent architecture, I mostly discussed boundaries, context, and collaboration patterns. This time I built a small Agent Architecture Lab and ran different architectures against the same task set.
The experiment was small, but the conclusion was clear: complexity is not a quality amplifier. It amplifies cost, context pressure, and failure surface. Multi-agent design is only worth considering when the task truly needs decomposition, review, or multiple positions.
How the Experiment Worked
The tasks covered four types: generating a platform proposal, expanding it into detailed design, reviewing that design, and choosing the most ambiguous trade-off from the review report to make a technical judgment.
I compared four architectures: single, planner_executor, planner_executor_reviewer, and debate. The observations included latency, token usage, estimated cost, truncation, final answer usefulness, and whether the intermediate process had review value.
This was not a rigorous large-sample study, but it was enough to expose typical engineering problems.
Single Agent Is Still a Strong Baseline
In the experiment, single had the lowest cost and latency, and the best stability. None of the four tasks were truncated, and its output was not weak on simple proposals or clear technical decisions.
That means Single Agent is not obsolete. It is the baseline every complex architecture must beat. When the task boundary is clear, the input is sufficient, and the output shape is explicit, adding more Agents often just adds process cost.
Planner Helps, but Output Must Be Controlled
planner_executor worked well for review tasks because review naturally requires multiple dimensions. A Planner can first separate the checklist, then an Executor can handle each part, which improves coverage.
But Planner also creates output expansion. In the detailed-design task, it pushed the Executor to cover more sections and eventually hit finish_reason=length. A Planner should not only make the model think more; it should manage section budget, length constraints, and output boundaries.
Reviewer Value Has to Be Visible
planner_executor_reviewer had the highest cost and latency, but it did not produce a stable quality jump in this run. Even with a Reviewer, the detailed-design task was still truncated.
Reviewer value is not the extra model call itself. It is the evidence chain: what the draft said, what the review found, and what changed in the final answer. If that is not structured and visible, the Reviewer becomes expensive ceremony.
Debate Only Fits Real Trade-Offs
debate made the most sense for technical trade-offs. Snapshot Mode vs Chain Mode is a good example: one side emphasizes fairness, isolation, and reproducibility; the other emphasizes continuity, state transfer, and realistic workflow.
But Debate does not fit every task. Without real disagreement, it often degrades into one Agent proposing, another adding risks, and a Judge summarizing. Cost goes up without necessarily adding true adversarial thinking.
The Cost of Complexity Is Concrete
In this experiment, planner_executor cost about 2.42 times as much as single, debate about 3.90 times, and planner_executor_reviewer about 6.17 times.
That is only model-call cost. Real systems also need result storage, intermediate-process UI, retries, timeouts, visualization, human review effort, and a more complex debugging path.
Final Thoughts
The value of multi-agent architecture is not making the model suddenly stronger. It is organizing uncertainty inside complex tasks. Planner organizes task structure, Reviewer organizes quality feedback, Debate organizes conflicting positions, and Judge organizes final merging.
But every structure has cost and failure modes. The real engineering question is not whether we can add Agents, but whether the system becomes more controllable, more reliable, and more worth its cost after we add them.