1975 年,Fred Brooks 在《人月神话》里写下一句被引用了五十年的话:向一个已经延期的软件项目加人,只会让它更晚。 不是新人不聪明,而是人一多,分工、交接、对齐、合并的成本会把新增的人手吃掉。
五十年后,我们以为终于找到了例外:Agent 不用入职培训、不用开会,复制一个只要一秒。于是 Claude Code、Codex、Kimi Code 都上了「动态并发」——主代理干到一半,自己决定拉起一群子代理并行开工。
结果 Brooks 又赢了一次。
这次有对照实验:Han Li、HanHaoNing Li、Ziqian Jiang、Yiling Lou《When Sub-Agents Work in Parallel: The Promises and Pitfalls of Dynamic Concurrency in Long-Horizon Coding Tasks》(arXiv:2610.10263) 把 Codex(GPT-5.4)、Claude Code(Claude Opus 5)、Kimi Code(Kimi K3)三个前沿编码 Agent 放到 354 个任务上,每个任务跑两遍:一遍开并发,一遍关掉,模型、提示词、环境、时间预算全都一样。总共 2,124 次执行、112 亿 token,按标准 API 价格约 20,780 美元。任务从 SWE-bench Verified 那种修一个 issue 的小活,一直到从零重建整个仓库、跨多个开发单元持续迭代的大活(LoopsBench,参考实现平均 2.68 万行)。

多花的钱,没有换来时间
先看账单。开并发后,token 用量是顺序执行的 1.41–3.31 倍(Claude Code 1.41 倍,Codex 3.31 倍)。按「每解决一个任务」折算更难看:Codex 是三家里唯一总解决数略有增加的,但开并发后每解决一个任务要花 2,412 万 token,顺序执行只要 769 万。
再看时间——这才是大家开并发的理由。15 组「工具 × 基准」对比里,有 14 组开了并发反而平均更慢。 Claude Code 在 SWE-bench Verified 上,开并发每题 27.8 分钟,关掉是 17.5 分钟,而且解出来的题还更少。只看两种模式都做成的 366 个任务,并发版更快的只有 102 个(27.9%),快 20% 以上的只有 46 个(12.6%)。
你以为你买的是「八个人同时干,一个小时变八分之一」。你实际买到的,是八份账单和一个要等所有人交作业的主代理。
小活上,它在帮倒忙
在 SWE-bench Verified 这类边界清楚的小任务上:
- Claude Code:任务通过率 83% → 59%(−24 个点);
- Kimi Code:77% → 59%(−18 个点);
- Codex:71% → 71%,不变。
把长短任务合起来做配对检验,Claude Code 和 Kimi Code 的下降在统计上显著,Codex 的小幅提升不显著。
并发不是没用。在最长的 LoopsBench 上,Claude Code 从 7.1% 涨到 21.4%(+14.3),Codex 从 21.4% 到 25.0%,Kimi Code 持平。四个长任务基准按代码规模排开,并发带来的平均测试通过率差从 RepoZero 的 −8.29 个点,到 NL2Repo 的 −0.04、ProgramBench 的 +1.24,再到 LoopsBench 的 +4.31——活越大,并发越划算;活不够大,协调成本先把收益吃掉。 这就是 Brooks 定律的 Agent 版。
失败清单,读起来像项目复盘会
作者人工标注了全部 1,062 条开并发的执行轨迹,其中 650 条出了并发特有的问题,一共 804 处。拿最大的几类和你见过的项目事故对一下:
- 共享状态与合并(33.2%):两个代理改同一个文件、互相覆盖交付物。仅「并发写」一项就占全部问题的 26.5%,在 Codex 的问题里占 62.9%。两个子代理各自删掉再重建
/tmp/t1当测试数据,其中一个不知情地接着用了对方的文件;还有一个子代理直接替换了系统里共享的 git 可执行文件。 - 执行管理(29.0%):子代理没干完就被掐掉(15.9%)、结果来得太晚赶不上合并、失败了没人接手。
- 任务编排(28.5%):人派太多,把预算花光(19.2%);或者活看上去在并行,关键路径其实还是串行的。
三个具体的故事:
- Kimi Code 重建命令行工具 seqtk:40 分钟的任务,第 17.5 分钟一次性把 22 个子命令派给 20 个子代理,11 秒内全部开工。主代理假设子代理超时是 两小时——比整个任务预算还长——然后干等整批交差,再也没调用过任何工具。到点时 18 个子代理被强行终止,负责
seq命令的那个一行实现都没写出来。最终 440 个测试过了 92 个。 - Claude Code 做一个 Rust 移植任务:主代理已经把能通过全部 39 个官方测试的版本构建好、放进了输出目录。临近截止,它开始阻塞等待一个还没干完的子代理,一等十分钟,任务在给出最终答复前被终止。
- Kimi Code 重建日志高亮工具:17 个子代理同时调研各种行为,没有一个被分配去写实现。主代理等齐所有报告才开始动手,到点时还在做实现规划。785 个测试,过了 0 个。
这里面没有一条是「模型不够聪明」。全是管理事故:没人认领、工期估错、到点不收工、两个人改同一份文件、交接文档写到 8 万字符没人看得进去。Brooks 当年算的是人与人之间的沟通线路;Agent 把会议换成了共享文件夹,沟通成本就以合并冲突的形式回来了。

并行真正赚钱的三种情况
论文也找出了并发确实划算的 131 组对照,归成三类:
- 互补分工:边界清楚的几块活分给不同子代理,主代理自己写核心再统一集成。一个 NL2Repo 任务两种模式都全过,但并发把耗时从 734 秒压到 468 秒。
- 多路尝试再择优:主代理和子代理各自实现一版,取长补短。一个 ProgramBench 任务,并发版过了 292/391 个测试,顺序版连构建都没过,0/391。
- 独立验证:子代理去核实一个关键假设,及时纠正主代理。一个安全实验任务里,子代理查到 ECDSA 需要的 P-256 曲线阶,主代理据此改对算法,顺序版则没做出来。
三种情况有一个共同点:主代理真的在当工头——拆得开、收得回、会验收。
边界
只测了三个「工具 + 自家模型」的组合(2026 年 7 月的配置),五个基准都有时间预算;长任务太贵,没法大量重复跑;失败分类是人工标注(Cohen's κ = 0.77)。编排做得更好的下一代工具,数字一定会变。但方向很难变:并发的收益取决于主代理会不会管人,而现在它大多数时候不会。
下次想开「多代理并行」之前
第一,默认关,按任务开。 修 bug、改一个函数、加一个接口这种活,顺序跑。只有在任务大到能拆成几块互不依赖的部分时才开。论文里各家的开关分别是 Codex 的 multi_agent / multi_agent_v2、Claude Code 的 Dynamic Workflow、Kimi Code 的 swarm_mode;具体用法以你所用版本的文档为准。
第二,开之前,先替它把分工表写好。 在提示词里写清:每个目录或文件归谁、模块之间的接口长什么样、谁负责最后集成和跑全量测试、子代理的超时必须短于总预算。能隔离就隔离——给每个子代理一个独立的 git worktree:
git worktree add ../wt-parser -b agent/parser
git worktree add ../wt-cli -b agent/cli
各写各的分支,最后由主代理(或你)合并,而不是所有人往同一个目录里写。
第三,拿你自己上周的一个真实任务做一次 A/B。 同一个任务,开并发、关并发各跑一遍,记下墙钟时间、token、测试通过数。如果在你自己的活上并发不更快,答案就有了。
判断很简单:子代理不是免费的人手,是需要管理的下属。并行 Agent 缺的不是人手,是工头——在主代理学会当工头之前,这个工头得由你来当。
In 1975, Fred Brooks wrote the sentence software people have quoted for fifty years: adding manpower to a late software project makes it later. Not because the new people are dumb, but because every extra person adds splitting, handoff, alignment and merge costs that eat the extra hands.
Fifty years on, we thought we had finally found the exception. Agents need no onboarding and no meetings, and a copy costs one second. So Claude Code, Codex and Kimi Code all shipped "dynamic concurrency": halfway through a task, the main agent decides on its own to spin up a crew of sub-agents working in parallel.
Brooks won again.
This time there is a controlled experiment: Han Li, HanHaoNing Li, Ziqian Jiang and Yiling Lou, When Sub-Agents Work in Parallel: The Promises and Pitfalls of Dynamic Concurrency in Long-Horizon Coding Tasks (arXiv:2610.10263) ran three frontier coding agents — Codex (GPT-5.4), Claude Code (Claude Opus 5) and Kimi Code (Kimi K3) — on 354 tasks, each task twice: once with concurrency on, once with it off, with the same model, prompt, environment and time budget. That is 2,124 executions, 11.2 billion tokens, about $20,780 at standard API rates. Tasks range from fixing a single SWE-bench Verified issue to rebuilding whole repositories from scratch and sustained multi-unit development (LoopsBench, mean reference size 26.8K lines).

The extra money didn't buy time
Start with the bill. With concurrency on, token use is 1.41–3.31× sequential (Claude Code 1.41×, Codex 3.31×). Per solved task it looks worse: Codex is the only one of the three whose solved count went up slightly, yet under concurrency it spends 24.12 million tokens per solved task versus 7.69 million sequentially.
Now time — the whole reason people turn concurrency on. In 14 of 15 agent × benchmark combinations, mean runtime went up with concurrency. Claude Code on SWE-bench Verified takes 27.8 minutes per task with concurrency and 17.5 minutes without, while solving fewer tasks. Looking only at the 366 cases both modes solved, the concurrent run was faster in 102 (27.9%), and at least 20% faster in only 46 (12.6%).
You thought you were buying "eight workers at once, an hour becomes eight minutes". What you got was eight bills and a main agent waiting for everyone to hand in their homework.
On small jobs it actively hurts
On bounded tasks like SWE-bench Verified:
- Claude Code: task pass rate 83% → 59% (−24 points);
- Kimi Code: 77% → 59% (−18 points);
- Codex: 71% → 71%, unchanged.
Pooling short and long tasks in a paired test, the declines for Claude Code and Kimi Code are statistically significant; Codex's small gain is not.
Concurrency isn't useless. On the longest benchmark, LoopsBench, Claude Code goes from 7.1% to 21.4% (+14.3), Codex from 21.4% to 25.0%, Kimi Code flat. Order the four long-horizon benchmarks by implementation size, and the average test-pass difference from concurrency climbs from −8.29 points on RepoZero to −0.04 on NL2Repo, +1.24 on ProgramBench and +4.31 on LoopsBench. The bigger the job, the more concurrency pays; when the job isn't big enough, coordination eats the gain first. That is Brooks's law, agent edition.
The failure list reads like a project post-mortem
The authors hand-annotated all 1,062 concurrent trajectories; 650 showed concurrency-specific failures, 804 instances in total. Match the biggest categories against incidents you have lived through:
- Shared state and merge (33.2%): two agents edit the same file and overwrite each other's deliverables. "Concurrent writes" alone is 26.5% of all instances and 62.9% of Codex's. Two sub-agents each deleted and recreated
/tmp/t1as a fixture, and one carried on unknowingly with the other's file; another sub-agent replaced the shared system git executable. - Execution governance (29.0%): sub-agents killed before finishing (15.9%), results arriving too late to merge, failed work nobody takes over.
- Task orchestration (28.5%): too many workers burning the budget (19.2%), or work that looks parallel while the critical path stays serial.
Three concrete stories:
- Kimi Code rebuilding the command-line tool seqtk: a 40-minute task. At minute 17.5 it hands 22 subcommands to 20 sub-agents in one batch, all started within 11 seconds. The main agent assumes a sub-agent timeout of two hours — longer than the whole task budget — then waits for the full batch and never makes another tool call. At the deadline 18 sub-agents are aborted; the one assigned to the
seqcommand never wrote its implementation. 92 of 440 tests pass. - Claude Code on a Rust port task: the main agent had already built the port, checked it and put it in the output directory — it passes all 39 official tests. Near the deadline it starts a ten-minute blocking wait on an unfinished sub-agent, and the run is terminated before a final response.
- Kimi Code rebuilding a log highlighter: 17 sub-agents investigate behaviors in parallel; none is assigned to implementation. The main agent waits for every report before writing code and is still planning the implementation when time runs out. 0 of 785 tests pass.
Not one of these is "the model wasn't smart enough". They are all management incidents: no owner, a wrong schedule, nobody calling time, two people editing one file, an 80K-character handoff brief nobody can absorb. Brooks counted communication paths between people. Agents swapped meetings for a shared folder, and the communication cost came back as merge conflicts.

The three cases where parallel actually pays
The paper also examined 131 pairs where concurrency did better, and they fall into three groups:
- Complementary work: cleanly separated pieces go to different sub-agents while the main agent writes the core and integrates. In one NL2Repo task both modes passed everything, but concurrency cut wall time from 734 to 468 seconds.
- Alternative attempts, then selection: the main agent and a sub-agent each build a version and combine the best of both. In one ProgramBench task the concurrent run passed 292/391 tests; the sequential run failed to build, 0/391.
- Independent validation: a sub-agent checks a key assumption and corrects the main agent in time. In a security-lab task, a sub-agent found the P-256 curve-order requirement for ECDSA, the main agent fixed its algorithm, and the sequential run failed.
All three have one thing in common: the main agent actually acts as foreman — it splits cleanly, collects the results, and checks the work.
Boundaries
Only three agent + native-model pairs (July 2026 configurations), five benchmarks, all with time budgets; long-horizon runs are too expensive to repeat many times; the failure taxonomy is hand-annotated (Cohen's κ = 0.77). A next generation with better orchestration will move the numbers. The direction is harder to move: what concurrency is worth depends on whether the main agent can manage, and right now it mostly can't.
Before you turn on multi-agent parallelism next time
First, default off, turn on per task. Bug fixes, one-function changes and single endpoints: run sequentially. Turn it on only when the task is big enough to split into pieces that don't depend on each other. In the paper the switches were Codex's multi_agent / multi_agent_v2, Claude Code's Dynamic Workflow and Kimi Code's swarm_mode; check your version's docs for exact usage.
Second, write the org chart before you turn it on. Put it in the prompt: who owns each directory or file, what the interfaces between modules look like, who does the final integration and full test run, and that sub-agent timeouts must be shorter than the total budget. Isolate where you can — give each sub-agent its own git worktree:
git worktree add ../wt-parser -b agent/parser
git worktree add ../wt-cli -b agent/cli
Each writes its own branch and the main agent (or you) merges, instead of everyone writing into one directory.
Third, A/B it on one real task from your last week. Same task, concurrency on and off, one run each; record wall-clock time, tokens and tests passed. If parallel isn't faster on your own work, you have your answer.
The judgment is simple: sub-agents are not free hands; they are reports who need managing. Parallel agents aren't short of hands, they're short of a foreman — and until the main agent learns to be one, the foreman has to be you.