DeLM

Decentralized Multi-Agent Systems with Shared Context

Parallel AgentsShared ContextTask Queue
Yuzhen MaoJerry GuAadi ChauhanQizheng ZhangHangoo KangAzalia Mirhoseini

Stanford University

Overview

TL;DR

  • Problem. Multi-agent systems run agents in parallel to tackle long-horizon tasks, yet existing designs waste much of this parallelism in bubbles: agent time spent waiting on others or redoing a peer's work.
  • Idea. DeLM replaces the main agent with a shared context and a task queue. Agents claim tasks asynchronously, publish findings as soon as they are usable, and build on or correct one another's progress, with every peer's status visible to all.
  • Results. On long-horizon tasks from Terminal-Bench 4.0 and DeepSWE v1.1, and on SWE-bench Verified, DeLM is both more accurate and faster than Codex, Claude Code, their native subagents, and AOrchestra: up to 17.5 points more accurate than the strongest baseline and up to 2.49× faster than the harness it builds on.

Existing multi-agent systems waste much of their parallelism in bubbles: agent time spent waiting on others or redoing a peer's work. DeLM squeezes out most of these bubbles by having agents coordinate asynchronously through a shared context and a task queue, with no main agent in the loop.

Accuracy against speedup on 10 selected long-horizon tasks from Terminal-Bench 4.0, headed up to 2.49× faster and 11.7 points more accurate. With Codex and GPT-6-Astra, DeLM with 2 agents is 1.54× faster and with 4 agents 2.05× faster, both above Codex, native subagents, and AOrchestra in accuracy. With Claude Code and Claude Opus 5.5, DeLM with 2 agents is 2.02× faster and with 4 agents 2.49× faster, both above Claude Code, native subagents, and AOrchestra.
Accuracy against speedup on 10 selected long-horizon tasks from DeepSWE v1.1, headed up to 1.57× faster and 17.5 points more accurate. With Codex and GPT-6-Astra, DeLM with 2 agents is 1.25× faster and with 4 agents 1.34× faster. With Claude Code and Claude Opus 5.5, DeLM with 2 agents is 1.50× faster and with 4 agents 1.57× faster, well above Claude Code, AOrchestra, and native subagents in accuracy.
DeLM is both faster and more accurate than existing multi-agent coding systems. Back-end models: GPT-6-Astra (with Codex) and Claude Opus 5.5 (with Claude Code), both at xhigh reasoning effort. Averaged over the two models, DeLM (n = 2) is 1.78× faster and 12.5 points more accurate than Codex/Claude Code (dashed line, 1×) on Terminal-Bench 4.0, and 1.38× faster and 13.3 points more accurate on DeepSWE v1.1. Each benchmark uses 10 long-horizon tasks, selected by how long Codex with GPT-6-Astra takes on them: between 15 minutes and 1 hour on Terminal-Bench 4.0, and at least 10 minutes on DeepSWE v1.1.

Motivation: workflow bubbles in Multi-Agent Systems

Multi-agent systems (MAS) offer a natural way to scale large language model reasoning at test time: instead of solving a complex task in a single trajectory, they decompose it into subtasks, dispatch agents in parallel, and aggregate their progress. Borrowing a term from pipeline parallelism, we call agent time that does not advance the solution a bubble: time an agent spends waiting on others, or redoing work a peer has already done. Bubbles waste compute and stretch wall-clock time, and each of the three dominant families of MAS creates them in its own way.

Independent agents

Redundant work

Their bubbles are redundant work: agents share nothing while they run, so each one rediscovers the faults, fixes, and dead ends its peers have already found.

Peer-communicating agents

Barrier waits

Their bubbles are barrier waits: each synchronous round lasts as long as its slowest agent, so agents that finish early sit idle.

Centralized orchestration

Relay waits

Their bubbles are relay waits: the main agent idles while delegated work runs, and sub-agents cannot build on each other's progress until the main agent relays it.

Execution timelines for the retro-console-soc task with four agents. Claude Code with native subagents takes 162.7 minutes with 44% of agent time waiting; its main agent's bar is mostly yellow waiting time. DeLM takes 29.3 minutes with 2% of agent time waiting, 5.5 times faster.
Centralized orchestration creates bubbles; DeLM squeezes them out. Execution timelines on Terminal-Bench 4.0's retro-console-soc task with Claude Opus 5.5 and four agents. With Claude Code's centralized native subagents, the main agent spends 73% of its time waiting on delegated work (yellow), and 44% of all agent time is spent waiting.

Replace the main agent with a shared context and a task queue

Left: a centralized system where one main agent reads and writes a global context, hands subtasks and subcontexts to sub-agents, and collects their returned results. Right: DeLM, where parallel agents read and write a shared context and a task queue directly.
Centralized vs. decentralized multi-agent systems. Centralized MAS relies on a main agent to assign subcontexts, spawn sub-agents, and integrate their results, so progress is shared mainly through the main agent and sub-agents remain unaware of one another's work. In contrast, DeLM decentralizes coordination through parallel agents, a shared context, and a task queue: agents asynchronously claim ready tasks, publish findings incrementally to a shared context, and build on or correct each other's progress without a main agent in the loop.

We propose Decentralized Language Models (DeLM), a coordination layer that squeezes out these bubbles by having agents coordinate asynchronously through a shared context and a task queue, with no main agent in the loop.

  • No barrier or relay waits. Agents claim tasks from the task queue asynchronously, so no agent waits for a round to end or for a main agent to assign work.
  • No waiting on unfinished work. Agents publish compact findings to the shared context as soon as they are usable, so peers build on intermediate progress instead of waiting for a relay.
  • No redundant work. The task queue and shared context show what each agent is working on and which approaches have failed, so agents avoid redoing a peer's work.
  • Updates reach everyone. Improvements and fixes are appended as new entries rather than relayed through a main agent, so every agent can pick up a peer's latest work without another round of delegation.
Collaboration and coordination mechanisms in the compared multi-agent systems. A checkmark means the system provides the mechanism while a cross means that it does not.
General collaboration capabilitiesExplicit coordination mechanisms
System Parallel subtask execution Cross-agent work reuse Feedback & correction Intermediate-results sharing Peer-status awareness Stateful shared context
Isolated agents✕✕✕✕✕✕
AOrchestra✕✓✓✕✕✕
Codex✓✓✓✓✕✕
Claude Code✓✓✓✓✕✕
DeLM✓✓✓✓✓✓

Decentralized Language Models (DeLM)

Given a task D, DeLM runs n agents concurrently, each with its own persistent session and workspace. Agents coordinate through two shared structures, a shared context C and a task queue T, and each agent produces its own complete solution.

Overview diagram of DeLM. A task queue holds subtasks T1 to T3. Parallel agents claim subtasks from the queue and read the shared context. Each agent's update is compressed into a gist G and appended to the shared context. Agents add more subtasks to the queue when needed, and otherwise finalize an answer.
Overview of DeLM. Parallel agents asynchronously claim subtasks {Ti}, read the shared context, and implement {Ti} locally. Each agent's update is compressed and appended to the shared context as a gist Gi, making reusable progress visible to all agents.
  1. Initialize the task queue
  2. Execute subtasks in parallel
  3. Summarize and publish intermediate results
  4. Update the task queue as work progresses
  5. Finalize each agent's solution

These operations proceed asynchronously rather than in lockstep: while one agent publishes a result, others keep working and can reuse it as soon as it appears.

Shared Context

An append-only log of what agents learn. Each entry has a declared type: FACT records a finding supported by observations, FAIL records an approach or hypothesis that was contradicted, and DONE summarizes a finished subtask. Entries stay brief and attach files by reference, so the log stays under 15K tokens even with four agents.

Task Queue

Divides the work among agents, with no planner in charge. Any agent can add tasks or claim open ones at any time, and the first claim wins. Every task shows whether it is open, claimed, or done, and by whom, so each agent knows who is responsible for what and how much remains.

Experiments

We evaluate DeLM on four agentic coding benchmarks: Terminal-Bench 4.0, where agents share discovered environment state and failed attempts in live terminals; DeepSWE v1.1, where agents reuse each other's debugging progress on long-horizon software engineering tasks; SWE-bench Verified, where agents explore alternative root causes and fixes for real GitHub issues; and ProgramBench, where agents rebuild entire programs from scratch. We compare against Codex and Claude Code, with and without native subagents, AOrchestra, and, on SWE-bench Verified, mini-SWE-agent.

Because DeLM targets long-horizon tasks where coordination matters, we evaluate on a subset of long-running tasks from Terminal-Bench 4.0 and DeepSWE v1.1. We select tasks based on the wall-clock latency of Codex with GPT-6-Astra: for Terminal-Bench 4.0, we select 10 tasks with latency between 15 minutes and 1 hour; for DeepSWE v1.1, we select 10 tasks with latency of at least 10 minutes.

Show the 10 selected tasks from each benchmark
Terminal-Bench 4.0Codex Latency
ks-solver-cpp21.00 min
telecom-entity-resolution24.23 min
vf2-speedup-networkx23.42 min
biped-contact-dynamics21.56 min
cumulative-layout-shift39.34 min
retro-console-soc27.28 min
rs-archive-clone21.92 min
wdm-design39.27 min
lake-temp-glm32.40 min
payments-pipeline-fix18.67 min
Average26.91 min
DeepSWE v1.1Codex Latency
oxvg-structural-selector-preservation19.71 min
wasmi-trap-coredumps14.55 min
happy-dom-abort-pending-body-reads12.52 min
scriggo-method-declarations17.35 min
dynamodb-toolbox-lazy-recursive-schemas22.19 min
dynamodb-toolbox-conditional-attribute16.78 min
opa-template-string-reconstruction13.99 min
boa-hierarchical-evaluation-cancellation13.28 min
numba-stencil-boundary-modes15.91 min
pebble-durability-wait-apis19.38 min
Average16.57 min

Selected tasks from Terminal-Bench 4.0 and DeepSWE v1.1. Baseline latency is the average execution time of the Codex baseline using GPT-6-Astra at xhigh reasoning effort.

Terminal-Bench 4.0

Comparison on 10 long-horizon tasks from Terminal-Bench 4.0 across two base models. Best mean values within each model are bolded.
MethodAvg. Acc (%)Avg. LatencySpeedupCost/TaskCost/Submission
GPT-6-Astra
Codex71.67 ±11.5526.91 ±0.91 min1.00×$9.01 ±0.71$9.01 ±0.71
Codex (subagent)71.67 ±11.6920.04 ±1.94 min1.34×$25.80 ±2.21$25.80 ±2.21
AOrchestra70.00 ±5.0025.78 ±1.15 min1.04×$11.94 ±1.02$11.94 ±1.02
DeLM (n = 2)85.00 ±13.2317.47 ±1.82 min1.54×$13.65 ±1.49$6.83 ±0.74
DeLM (n = 4)81.67 ±12.3313.16 ±1.43 min2.05×$23.90 ±2.27$5.97 ±0.57
Claude Opus 5.5
Claude Code81.67 ±2.89130.75 ±6.27 min1.00×$20.31 ±0.23$20.31 ±0.23
Claude Code (subagent)80.00 ±5.00138.58 ±10.17 min0.94×$30.88 ±3.17$30.88 ±3.17
AOrchestra76.67 ±5.77157.21 ±9.03 min0.83×$24.22 ±1.05$24.22 ±1.05
DeLM (n = 2)93.33 ±3.3364.77 ±5.03 min2.02×$17.91 ±1.17$8.96 ±0.59
DeLM (n = 4)89.17 ±5.2052.51 ±4.43 min2.49×$29.98 ±0.91$7.50 ±0.23

DeepSWE v1.1

Comparison on 10 long-horizon tasks from DeepSWE v1.1 across two base models. Best mean values within each model are bolded.
MethodAvg. Acc (%)Avg. LatencySpeedupCost/TaskCost/Submission
GPT-6-Astra
Codex88.33 ±2.8916.57 ±0.39 min1.00×$9.40 ±0.18$9.40 ±0.18
Codex (subagent)90.00 ±0.0015.76 ±2.05 min1.05×$24.97 ±3.15$24.97 ±3.15
AOrchestra83.33 ±7.6418.31 ±1.08 min0.91×$10.74 ±0.97$10.74 ±0.97
DeLM (n = 2)98.33 ±2.8913.29 ±0.98 min1.25×$16.83 ±1.65$8.42 ±0.82
DeLM (n = 4)90.00 ±10.0012.37 ±0.66 min1.34×$30.87 ±1.09$7.72 ±0.27
Claude Opus 5.5
Claude Code71.67 ±2.8949.58 ±3.11 min1.00×$12.01 ±0.60$12.01 ±0.60
Claude Code (subagent)50.00 ±10.0057.46 ±2.16 min0.86×$19.83 ±0.99$19.83 ±0.99
AOrchestra73.33 ±5.7747.32 ±2.86 min1.05×$12.11 ±0.52$12.11 ±0.52
DeLM (n = 2)88.33 ±2.8932.98 ±2.47 min1.50×$12.15 ±0.57$6.07 ±0.28
DeLM (n = 4)90.83 ±8.7831.61 ±0.96 min1.57×$24.49 ±0.67$6.12 ±0.17

With Claude Opus 5.5, the gains are larger: DeLM (n = 4) reaches 90.83%, 17.5 points above the strongest baseline, AOrchestra (73.33%), and runs 1.57× faster than Claude Code.

SWE-bench Verified

Comparison on SWE-bench Verified with Gemini 3 Flash. Best mean values within each model are bolded.
MethodAvg. Acc (%)Avg. LatencySpeedupCost/TaskCost/Submission
Gemini 3 Flash
Claude Code49.32 ±1.98129.13 ±5.12 s1.00×$1.00a ±0.00$1.00a ±0.00
mini-SWE-agent54.73 ±2.3791.78 ±3.53 s1.41×$0.26 ±0.07$0.26 ±0.07
AOrchestra55.26 ±2.0389.94 ±3.07 s1.44×$0.24 ±0.04$0.24 ±0.04
DeLM (n = 2)63.72 ±2.2971.86 ±2.88 s1.80×$0.24 ±0.05$0.12 ±0.03
DeLM (n = 4)66.08 ±1.8661.01 ±2.63 s2.12×$0.47 ±0.07$0.12 ±0.02

a The Claude Code CLI sends cache_control blocks in the Anthropic API format, so cache reuse only works if the upstream provider honors those blocks. For Gemini-3-Flash the real cost without cache reuse is therefore around $1 per task.

With n = 4, DeLM reaches 66.08% accuracy, 10.8 points higher than the strongest baseline, AOrchestra (55.26%).

ProgramBench

Hidden-test pass rate over 120 minutes on the ctags and pandoc ProgramBench tasks. On ctags, Claude Code ends at 34.32%, DeLM with two agents at 41.05%, and DeLM with four agents at 53.28%. On pandoc they end at 30.09%, 38.26%, and 50.00%.
DeLM makes more progress within the same time budget. Test pass rate over time on two ProgramBench tasks with Claude Opus 5.5 under a 120-minute budget. DeLM uses Claude Code as its per-agent harness. From 60 minutes onward, the curves stay ordered (n = 4 above n = 2 above Claude Code), so DeLM reaches any given pass rate sooner and ends the budget higher.

Trace Analysis: How Does Decentralized Coordination Help?

Parallel subtask execution shortens the critical path

Execution timelines on the retro-console-soc task with GPT-6-Astra. A single agent builds CPU, graphics, and audio in sequence and finishes at 27.01 minutes. Two DeLM agents split graphics and system from CPU and audio and finish around 21 minutes. Four DeLM agents build system, graphics, CPU, and audio concurrently and finish around 7.5 minutes.
More agents build more components in parallel. Execution timelines on Terminal-Bench 4.0's retro-console-soc task with GPT-6-Astra.

A single agent builds the CPU, graphics, and audio in sequence and finishes in 27.0 minutes. With n = 2, one agent builds the graphics and system components while its peer builds the CPU and then the audio, finishing in 20.9 minutes (1.29×). With n = 4, separate agents build the system, graphics, CPU, and audio concurrently, and the task finishes in 7.5 minutes (3.6×).

Peer-status awareness avoids duplicate exploration

DeLM exposes every agent's claims, status, and stated scope through the task queue and shared context, so an agent can notice an overlap while its search is still running and redirect. In one four-agent run, Agents 1 and 3 independently began tuning the same port widths and positions. Agent 1 then saw Agent 3's claim through the shared context:

[agent-3/FACT] [...] I will tune only w_in,w_long,w_short,y_long,y_short [...]
[agent-1/FACT] [...] applying the cavity artifact revealed agent-3 had just claimed
    the same five-parameter port tuning. [...]
[agent-1/EXEC] kill -TERM 770
[agent-1/EXEC] python binary_refine.py [...]

The agents resolve the overlap themselves, before either search finishes, without a main agent having to detect the duplication and re-plan.

Stateful shared context enables reuse and correction

Because updates and corrections are appended as new entries that reference revised files, rather than overwriting old ones, an agent can re-import a peer's updated work at any time. Across 720 agent trajectories on Terminal-Bench 4.0 and DeepSWE v1.1, agents imported shared files 5,237 times, and 80.3% of trajectories later imported a revised version of a file they had already imported. Because agents publish short summaries and attach files by reference, the shared context averages under 15K tokens at task completion.

Final shared-context size (thousands of tokens) and the cost of reading it as a percentage of total task cost.
Terminal-Bench 4.0DeepSWE v1.1
ModelAgentsContext (k)Cost (%)Context (k)Cost (%)
GPT-6-Astran = 27.31 ±0.429.17 ±0.387.54 ±0.217.71 ±0.18
n = 412.88 ±0.9717.68 ±0.2513.03 ±0.9614.56 ±0.37
Claude Opus 5.5n = 26.62 ±0.341.92 ±0.114.67 ±0.111.69 ±0.06
n = 414.48 ±0.473.66 ±0.2910.43 ±0.533.04 ±0.23

BibTeX

citation.bib
@misc{mao2026delm,
  title         = {Decentralized Multi-Agent Systems with Shared Context},
  author        = {Yuzhen Mao and Jerry Gu and Aadi Chauhan and
                   Qizheng Zhang and Hangoo Kang and Azalia Mirhoseini},
  year          = {2026},
  eprint        = {2606.10662},
  archivePrefix = {arXiv},
  primaryClass  = {cs.MA},
  url           = {https://arxiv.org/abs/2606.10662}
}