A structural representation that makes emergent LLM-agent collaboration explicit - tasks, environment, and agents on
one record - without prescribing agent reasoning, behavior, or communication order. Infuse turns
that record into a better collaboration policy.
Paper coming soon Code coming soon
Record → roll up → infuse. Execution writes the collaboration onto a scroll with three lanes:
tasks (T), agent activations and messages (D), and the environment
(Z). Rolling the scroll back up merges each agent's activations into a node and tallies messages
into weighted edges - the recovered topology and rules. Infuse steeps that record and pours out an updated policy
Γ′ for the next episode.
Collaboration Structure Emerges During Execution
A collaboration policy - an executable graph, or code and instructions - rarely controls everything. Agents deviate from
workflows, orchestrators decompose work on the fly, and in open-ended settings agents organize collaboration themselves.
Flexibility helps performance, but the structure that emerges is hard to track, and therefore hard to improve.
1
Represent
TEA couples an Agent Interaction Graph (asynchronous activations and events) to a Task Activity Graph (immutable task versions and the actions that create them) and the shared environment. Task changes are attributed to activations; evidence attaches to exact task versions.
2
Recover
Merge each agent's activations into a node and count messages per edge to recover the observed topology. Activation timing and message paths then expose sequential, parallel, repeated, conditional, and join rules - including the ones the policy never specified.
3
Infuse
Query the record to check the recovered policy against task progress and environmental outcomes, then propose an updated policy Γ′. Starting from an empty policy, one Infuse update improves task performance in both benchmarks.
43.0% vs 19.5%
LoC reduction in code refactoring with TEA's task graph + communication, against communication alone.
0.77 vs 0.03
Silo-Bench success with the shared task graph, against neither task graph nor communication.
0.63 → 0.70
Average pass rate across six repositories after one Infuse update.
0.77 → 0.83
Silo-Bench success after one Infuse update.
Read more: the paper's overview figure and contributions
TEA and Infuse. (1) A collaboration policy Γ may fully, partially, or
minimally specify coordination. (2) TEA combines views of Task evolution,
Environmental outcomes, and Agent interactions; solid edges are prescribed by the
policy, dashed edges emerge from agents. (3) Aggregation and trace analysis recover and discover the topology and
candidate execution rules. (4) Infuse queries these records to assess coordination against task progress and
environmental outcomes, and (5) proposes an updated policy Γ′ to execute.
Contributions. (1) The TEA representation, which grounds asynchronous agent interactions and agent-created task evolution in the shared environment. (2) A formal definition of collaboration policies, a proof that TEA can represent a large class of agent interactions, and policy recovery and discovery through topology reconstruction, inferred execution rules, and cross-layer queries. (3) Representation experiments showing that TEA records diverse interaction and task structures and correctly recovers their topology and policy properties. (4) Infuse, which learns from open-ended collaboration by recovering a policy from TEA traces, assessing the traces against task progress and environmental outcomes, and updating the policy.
Task versions and the actions that create them. A version carries a persistent identity, a goal and evaluation rule, and a state. Versions are immutable, so branching and merging work keeps its full history.
D
Agent Interaction Graph AIG
Events and activations. An activation is one reasoning cycle of an agent: its interval, received events, and ordered tool calls with results. Events carry any payload and can wake other agents.
Z
Shared Environment Env
The world agents observe and change. Environment calls return fresh events into the AIG, so task declarations and evidence stay tied to externally verifiable effects.
The three layers, aligned over time. Task rows hold versions (orange circles) and actions (squares).
Wide rectangles are activations, inner squares are tool calls, circles are events. The environment row records state
transitions. Every tool call is recorded in its issuing activation; task and environment calls additionally update the
task graph and the environment, tying the layers together.
Read more: activations, tools, transitions, and what TEA can represent
One activation. Agent i's k-th activation
hi,k = ([t:t′], Eh, Uh) records its input events,
its duration, and the ordered tool calls ur, each paired with the events
Er it returned.
Besides task and environment tools, TEA provides sendh(e, J), which delivers an event to agents
J (a new message or a forwarded event), and attachh(φ), which attaches
evidence φ = (h, x, q) to an exact task version. Evidence stays with that version and
does not carry over to later revisions.
Tool call
AIG
TAG
Env.
Evidence
task tool u ∈ UT
D′
T′
Z
Φ
environment tool u ∈ UZ
D′
T
Z′
Φ
attachh(φ)
D′
T
Z
Φ′
sendh(e, J)
D′
T
Z
Φ
TEA transitions: highlighted cells are updated by the call.
Six task tools
OPEN(δ, σ)∅ → q
Create a fresh task identity with an initial version.
EDITδ(q, δ′)q → q′
Revise the goal or evaluation rule while preserving the reported state.
UPDATEσ(q, σ′)q → q′
Change the reported state while preserving the task specification.
SPLIT(q, {(kj, δj, σj)})q → {q1, …, qm}
Decompose one version into versions with distinct task identities.
JOIN({qj}, k′, δ′, σ′){q1, …, qm} → q′
Integrate versions from distinct identities into a single successor version.
CLOSE(k)propagate closure
Closure propagates backward through recorded ancestry; a split's parent closes only when all of its children are closed.
Two representation properties. The AIG forms a temporal DAG over activations and events that can represent
any finite DAG of activation dependencies, with each output event having exactly one parent activation. The six task tools
can represent any finite DAG of task versions, adding at most one auxiliary leaf per version with a single predecessor.
Timestamps consistent with the dependencies are recorded as annotations.
Evidence and task progress. For a task version q, the attached evidence, its
providers, and its content together provide signals of task progress. Interpretations from different agents may agree,
partially agree, or disagree, and different interpreters may weight sources differently.
Representation experiments
Interactions and task history. (a) Interaction graphs and AIG traces for six LangGraph collaboration
patterns. (b) TAG is a general-purpose task representation: individual or composed task actions express Git's
collaborative coding semantics - commits, branches, merges, rebases, reverts, amends, notes, and branch deletion.
(c) AIG represents teams of any size, illustrated by a 50-worker supervisor and a 40-agent hierarchy.
Recovering the Collaboration Policy
A collaboration policy is an interaction topology plus coordination rules. Given a TEA record, aggregate the activations
in a window by agent: agents become nodes, and message counts become edge weights. Activation timing and message paths
then reveal execution order, parallel overlap, repetition, conditional routing, and joins.
Recovering a hidden policy. (a) Task evolution (T), agent activity
(D), and environment transitions (Z) recorded under a hidden
collaboration policy. (b) The interaction graph recovered by aggregating activations by agent, with edges weighted by
interaction frequency. (c) Execution rules inferred as the record grows (P: planner; W: workers; V: verifier).
Asking "who worked on task k2, and why?" traces back to the planner
splitting k1, messaging both workers, and worker1 acting in the environment first.
Read more: the definition, the rules with a worked example, queries, and policies that change over time
Collaboration policy for LLM agents. For a fixed agent team, task, and execution setup, a collaboration
policy Γ is an interaction topology along with a set of coordination rules that together
guide or constrain task execution, shaping the joint distribution of agent activations, tool calls, and event exchanges
over time.
Topology recovery. Take the activations that start and finish within a window
[t, t′] and merge them by agent. The result is a weighted graph
G[t,t′] = (V, E, w) whose edges connect agents that exchanged messages,
weighted by message counts. The full topology is recovered once the records cover every edge the policy can create.
Execution rules. Consecutive activations whose starts fall within τ of
each other and whose intervals overlap are grouped. For completed activations with start
sh and finish fh:
Aggregation gives a ↔ b and a ↔ c, two messages per
direction. With τ = 0.6, hb,1 and
hc,1 are grouped: {a} → {b, c} → {a} → {b} → {a} → {c} → {a}.
Sequential
Apply ≺ along activation paths. ha,1 ≺ hc,1 ≺ ha,3 and ha,3 ≺ hc,2 ≺ ha,4 support the local order a → c → a.
Parallel
hc,1 overlaps hb,1, ha,2, and hb,2: a can receive b's result and send more work to b while c remains active.
Repetition
Follow message paths that return to later activations of the same agent. ha,1 → hb,1 → ha,2 → hb,2 → ha,3 contains a → b → a twice.
Conditional
The graph shows whom an agent sends to; the conditions behind each message are inferred from its input messages and subsequent sending behavior.
Joins
Observe which senders feed an agent's activations and how consistently those sets recur: recurring both-inputs support an all-input rule, either alone an any-input rule.
Queries across layers
Agents' activities induce directed edges (→) and implicit links
(↦) across TEA. Queries traverse them in either direction, within or across layers;
each path between two entities gives an interpretation of their relationship. Tracing task versions reveals who
worked on a task, what they changed and in what order, and the linked evidence and environmental feedback; tracing
further back to activations shows whom they interacted with and what information they received.
Relations and tools in TEA
AIG
Actor
(h, u) ↦ h ↦ ih
Event flow
e → h; h → e′
TAG
Identity
q ↦ k
Open / Close
a → q; a ↦ k
Edit / Update
q → a → q′
Split
q → a → {q′j}
Join
{qj} → a → q′
TEA
Task action
(h, u) ↦ a, u ∈ UT
Environment
(h, u) ↦ u, u ∈ UZ
sendh
(ih, e, j), j ∈ J
attachh
(h, u) ↦ φ
Feedback
(h, u) ↦ Eu ⊆ E
Evidence
φ ↦ h; φ ↦ q; φ ↦ x
Policies that change mid-task. (a) Collaboration policies for three milestones and (b) their
execution trace. (c) Perfect chunking and (d) equal-length chunking, with their recovered policies. Chunks that cross
phases mix interactions from different policies and distort the recovered topology - an opening for detecting
policy changes and selecting aggregation windows.
Routing recovery. (a) A routing policy and (b) its weighted adjacency matrix of edge probabilities;
(c) one execution's AIG and (d) the L∞ distance between the recovered
edge-frequency matrix and (b) as the run proceeds. Crosses mark edges not yet run.
Infuse: From the Record to a Better Policy
TEA is also the interface through which agents communicate, transform tasks, and act in the environment. Infuse takes the
trace of an episode - starting from an empty policy - and proposes the next one.
1
Execute and record
DT ← fexecute(V, Γ)
The team V attempts the task under policy Γ, and TEA records the collaboration as it unfolds.
2
Recover / discover
Γ̂ ← frecover(DT)
Aggregate activations by agent to recover the observed topology and infer candidate execution rules.
3
Analyze and propose
Γ′ ← fanalyze(Γ, Γ̂, DT)
Query the record to compare the recovered policy with the trace, assess coordination against task progress and environmental outcomes, and decide what to keep, add, revise, or leave open.
In the paper's implementation, a separate gpt-6-sol agent consolidates this guidance into global lessons for
the team and private lessons for individual agents, forming Γ′. Infuse itself is
agnostic to how Γ′ is produced.
Read more: how agents run on TEA
A team of agents collaborates asynchronously through TEA. During an activation, an agent acts on its environment view,
its task-graph view, received messages, and private memory; its tool calls support communication, task transformations,
evidence attachment, and environmental actions, all recorded through TEA. Three mechanisms control when agents activate:
After each round of reasoning and tool use, the agent decides whether to continue its current activation or deactivate.
The sleep operation ends the current activation and schedules reactivation after a specified delay.
When sending a message, the agent specifies whether it should interrupt the recipient. Delivery activates an inactive recipient; if the recipient is active and interruption is requested, its current activation ends and a new one begins with the incoming message.
Results
Two domains, five gpt-5.6-terra agents with no predefined roles: code refactoring (modularize a
redundant codebase; LoC reduction while passing fixed tests, six repositories, 100-step budget) and
Silo-Bench (30 distributed reasoning tasks where every agent must answer a query that needs the whole team's
information). A separate gpt-6-sol agent derives the Infuse guidance.
Does the shared task graph help?
Access to TAG and to direct communication is ablated in four conditions.
Code (shop)
Silo-Bench
Setting
Pass ↑
Rall ↑
Rpass ↑
Success ↑
S ↑
P ↑
TAG + comm
4/5
46.9
43.0
0.77
0.81
0.85
TAG only
3/5
35.3
35.2
0.77
0.81
0.83
Comm only
3/5
20.2
19.5
0.60
0.73
0.78
Neither
4/5
21.3
20.8
0.03
0.15
0.20
Code refactoring on shop: five runs per condition; Pass counts runs passing all tests;
Rall and Rpass are mean LoC reductions (%)
over all and passing runs. Silo-Bench: 30 tasks; S is the fraction of exactly correct answers,
P the partial-credit score. Best values in bold.
LoC reduction over agent steps on shop: mean of five runs per condition, shading
±1 SD. The gap between the TAG conditions and the others widens during execution.
TEA's shared task graph outperforms direct communication alone on both benchmarks, while combining the two yields the
strongest overall results.
Does Infuse improve the policy?
TEA with and without one Infuse update, five runs per repository and condition.
0.63 → 0.70
Average pass rate across the six refactoring repositories.
25.5 → 30.4%
Repository-average LoC reduction among passing runs; it rises in all six repositories.
0.77 → 0.83
Silo-Bench success with TAG + comm.
0.81 → 0.86
Silo-Bench exact correctness S with TAG + comm.
Infuse enables TEA to refine its collaboration policy from experience and improve overall task performance.
Read more: full Infuse tables, communication patterns under ablation, and five rounds of Infuse
Code refactoring before (TEA) and after one Infuse update (Inf.)
Easy
Medium
Hard
Average
Codebase
shop
plan
ledger
freight
stay
payroll
TEA
Inf.
TEA
Inf.
TEA
Inf.
TEA
Inf.
TEA
Inf.
TEA
Inf.
TEA
Inf.
Pass rate ↑
0.80
0.80
0.80
0.80
0.60
0.80
0.60
0.80
0.40
0.60
0.60
0.40
0.63
0.70
LoC reduction (%) ↑
43.0±10.2
56.6±8.9
28.3±5.9
31.6±3.7
22.1±3.4
25.1±1.7
20.6±4.0
22.1±3.5
20.2±2.6
24.8±4.4
18.8±2.5
22.2±9.0
25.5±9.2
30.4±13.3
Average is the unweighted repository mean; ± denotes sample SD across passing runs, or across repositories for
Average. Better paired values in bold; ties unbolded. Pass rates rise on ledger, freight, and
stay, remain unchanged on shop and plan, and only fall on payroll.
Silo-Bench before (TEA) and after one Infuse update (Inf.), 30 tasks
TAG + comm
TAG only
Comm only
Neither
TEA
Inf.
TEA
Inf.
TEA
Inf.
TEA
Inf.
Success ↑
0.77
0.83
0.77
0.73
0.60
0.67
0.03
0.17
S ↑
0.81
0.86
0.81
0.75
0.73
0.79
0.15
0.19
P ↑
0.85
0.88
0.83
0.78
0.78
0.82
0.20
0.24
Bold marks the best value in each row. Infuse raises success in three of four settings; TAG-only success slightly falls from 0.77 to 0.73.
Who sent to whom under ablation, summed over five seeds per setting on one scale. Removing TAG
raises traffic rather than lowering it: 244 events over 20 edges without TAG against 94 over 20 with it.
Panels (c) and (d) are empty by construction - with send removed, there is no edge to draw.
Iterating Infuse over five rounds. Top: the team's interaction graph in each round as it inherits the
previous round's lessons. Bottom: score, steps, and LoC reduction relative to the vanilla setting.