1
Return

Capable language models can outgrow the benefits of collaboration

delete2026-07-24
delete0
delete
OA
AI
Y
Yubin Kim *
K
Ken Gu
C
Chanwoo Park
C
Chunjong Park
S
Samuel Schmidgall
A
A. Ali Heydari
Y
Yao Yan
Z
Zhihan Zhang
Y
Yuchen Zhuang
刘勇 cover
刘勇 (Liu Y)
M
Mark Malhotra
P
Paul Pu Liang
H
Hae Won Park
Y
Yuzhe Yang
X
Xuhai Xu
Y
Yilun Du
S
Shwetak Patel
T
Tim Althoff
D
Daniel McDuff *
X
Xin Liu *
DOI:10.1038/s42256-026-01268-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Agents, language model-based systems that can reason, plan and act with tools to accomplish tasks, are widely deployed, yet it remains unclear when multi-agent coordination outperforms a strong single agent. Here we conduct a controlled experiment that holds task prompts, tools and compute budgets constant while varying only coordination structure and model capability. Across 260 configurations spanning six benchmarks, five architectures and three LLM families, we derive a predictive model using empirical coordination metrics. Across benchmarks, single-agent baseline performance emerges as the most robust predictor of whether coordination improves or decreases performance. In particular, we identify an empirical capability-saturation threshold beyond which additional agents are unlikely to improve performance. This threshold correctly predicts the effect of multi-agent coordination on performance in 94% of validation configurations on SWE-bench Verified and Terminal-Bench. We therefore interpret this threshold as a practical selection rule rather than a universal scaling principle. A second effect, baseline-scaled error amplification, survives cluster-robust inference (Probust = 0.030) and supports the failure-mode taxonomy. The fitted model achieves cross-validated R2 = 0.373 (0.413 with a task-grounded capability metric) and selects the best architecture in 87% of held-out configurations. These results provide a quantitative framework for within-domain architecture selection and for estimating when multi-agent coordination is likely to improve performance or add overhead. A controlled study of large language model agents across 260 configurations shows when multi-agent collaboration helps or hurts performance, and introduces a predictive model that selects the best architecture in 87% of held-out within-domain configurations.

Journal

Nature Machine Intelligence cover
Nature Machine Intelligence
IF:
23.9
Papers:
1.3K
Citations:
1.5W

Organization

G
google research
Scholars:
187
Papers: 30
Citations: 0
G
google deepmind
Scholars:
220
Papers: 59
Citations: 44
M
massachusetts institute of technology
Scholars:
3.3K
Papers: 1.2K
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers