Task · Environment · Agents

TEA: Structurally Representing Arbitrary LLM Agents Collaboration

A structural representation that makes emergent LLM-agent collaboration explicit - tasks, environment, and agents on one record - without prescribing agent reasoning, behavior, or communication order. Infuse turns that record into a better collaboration policy.

Record → roll up → infuse. Execution writes the collaboration onto a scroll with three lanes: tasks (T), agent activations and messages (D), and the environment (Z). Rolling the scroll back up merges each agent's activations into a node and tallies messages into weighted edges - the recovered topology and rules. Infuse steeps that record and pours out an updated policy Γ′ for the next episode.

Collaboration Structure Emerges During Execution

A collaboration policy - an executable graph, or code and instructions - rarely controls everything. Agents deviate from workflows, orchestrators decompose work on the fly, and in open-ended settings agents organize collaboration themselves. Flexibility helps performance, but the structure that emerges is hard to track, and therefore hard to improve.

1

Represent

TEA couples an Agent Interaction Graph (asynchronous activations and events) to a Task Activity Graph (immutable task versions and the actions that create them) and the shared environment. Task changes are attributed to activations; evidence attaches to exact task versions.

2

Recover

Merge each agent's activations into a node and count messages per edge to recover the observed topology. Activation timing and message paths then expose sequential, parallel, repeated, conditional, and join rules - including the ones the policy never specified.

3

Infuse

Query the record to check the recovered policy against task progress and environmental outcomes, then propose an updated policy Γ′. Starting from an empty policy, one Infuse update improves task performance in both benchmarks.

43.0% vs 19.5%
LoC reduction in code refactoring with TEA's task graph + communication, against communication alone.
0.77 vs 0.03
Silo-Bench success with the shared task graph, against neither task graph nor communication.
0.63 → 0.70
Average pass rate across six repositories after one Infuse update.
0.77 → 0.83
Silo-Bench success after one Infuse update.
Read more: the paper's overview figure and contributions
TEA and Infuse overview: a collaboration policy is executed, TEA records task evolution, environment outcomes, and agent interactions, a policy is inferred from the record, analyzed, and an updated policy is proposed for the next episode
TEA and Infuse. (1) A collaboration policy Γ may fully, partially, or minimally specify coordination. (2) TEA combines views of Task evolution, Environmental outcomes, and Agent interactions; solid edges are prescribed by the policy, dashed edges emerge from agents. (3) Aggregation and trace analysis recover and discover the topology and candidate execution rules. (4) Infuse queries these records to assess coordination against task progress and environmental outcomes, and (5) proposes an updated policy Γ′ to execute.

Contributions. (1) The TEA representation, which grounds asynchronous agent interactions and agent-created task evolution in the shared environment. (2) A formal definition of collaboration policies, a proof that TEA can represent a large class of agent interactions, and policy recovery and discovery through topology reconstruction, inferred execution rules, and cross-layer queries. (3) Representation experiments showing that TEA records diverse interaction and task structures and correctly recovers their topology and policy properties. (4) Infuse, which learns from open-ended collaboration by recovering a policy from TEA traces, assessing the traces against task progress and environmental outcomes, and updating the policy.

Task-Environment-Agents

DT = (D, T, Z, Φ, U)    interaction graph · task graph · environment · evidence attachments · tool space
T

Task Activity Graph TAG

Task versions and the actions that create them. A version carries a persistent identity, a goal and evaluation rule, and a state. Versions are immutable, so branching and merging work keeps its full history.

D

Agent Interaction Graph AIG

Events and activations. An activation is one reasoning cycle of an agent: its interval, received events, and ordered tool calls with results. Events carry any payload and can wake other agents.

Z

Shared Environment Env

The world agents observe and change. Environment calls return fresh events into the AIG, so task declarations and evidence stay tied to externally verifiable effects.

The three layers of TEA: task rows with versions and actions on top, agent activations exchanging events in the middle, and environment state transitions at the bottom, aligned over time
The three layers, aligned over time. Task rows hold versions (orange circles) and actions (squares). Wide rectangles are activations, inner squares are tool calls, circles are events. The environment row records state transitions. Every tool call is recorded in its issuing activation; task and environment calls additionally update the task graph and the environment, tying the layers together.
Read more: activations, tools, transitions, and what TEA can represent
One AIG activation: input events, a time interval, and an ordered sequence of tool calls each paired with returned events
One activation. Agent i's k-th activation hi,k = ([t:t′], Eh, Uh) records its input events, its duration, and the ordered tool calls ur, each paired with the events Er it returned.

Besides task and environment tools, TEA provides sendh(e, J), which delivers an event to agents J (a new message or a forwarded event), and attachh(φ), which attaches evidence φ = (h, x, q) to an exact task version. Evidence stays with that version and does not carry over to later revisions.

Tool callAIGTAGEnv.Evidence
task tool u ∈ UTD′T′ZΦ
environment tool u ∈ UZD′TZ′Φ
attachh(φ)D′TZΦ′
sendh(e, J)D′TZΦ

TEA transitions: highlighted cells are updated by the call.

Six task tools

OPEN(δ, σ)∅ → q

Create a fresh task identity with an initial version.

EDITδ(q, δ′)q → q′

Revise the goal or evaluation rule while preserving the reported state.

UPDATEσ(q, σ′)q → q′

Change the reported state while preserving the task specification.

SPLIT(q, {(kj, δj, σj)})q → {q1, …, qm}

Decompose one version into versions with distinct task identities.

JOIN({qj}, k′, δ′, σ′){q1, …, qm} → q′

Integrate versions from distinct identities into a single successor version.

CLOSE(k)propagate closure

Closure propagates backward through recorded ancestry; a split's parent closes only when all of its children are closed.

Two representation properties. The AIG forms a temporal DAG over activations and events that can represent any finite DAG of activation dependencies, with each output event having exactly one parent activation. The six task tools can represent any finite DAG of task versions, adding at most one auxiliary leaf per version with a single predecessor. Timestamps consistent with the dependencies are recorded as annotations.
Evidence and task progress. For a task version q, the attached evidence, its providers, and its content together provide signals of task progress. Interpretations from different agents may agree, partially agree, or disagree, and different interpreters may weight sources differently.

Representation experiments

Six collaboration patterns with their AIG traces, Git operations expressed as TAG task tools, and AIG records of a 50-worker supervisor and a 40-agent hierarchy
Interactions and task history. (a) Interaction graphs and AIG traces for six LangGraph collaboration patterns. (b) TAG is a general-purpose task representation: individual or composed task actions express Git's collaborative coding semantics - commits, branches, merges, rebases, reverts, amends, notes, and branch deletion. (c) AIG represents teams of any size, illustrated by a 50-worker supervisor and a 40-agent hierarchy.

Recovering the Collaboration Policy

A collaboration policy is an interaction topology plus coordination rules. Given a TEA record, aggregate the activations in a window by agent: agents become nodes, and message counts become edge weights. Activation timing and message paths then reveal execution order, parallel overlap, repetition, conditional routing, and joins.

A TEA trace generated by a hidden collaboration policy, the interaction graph recovered from it, and the execution rules inferred as the record grows
Recovering a hidden policy. (a) Task evolution (T), agent activity (D), and environment transitions (Z) recorded under a hidden collaboration policy. (b) The interaction graph recovered by aggregating activations by agent, with edges weighted by interaction frequency. (c) Execution rules inferred as the record grows (P: planner; W: workers; V: verifier). Asking "who worked on task k2, and why?" traces back to the planner splitting k1, messaging both workers, and worker1 acting in the environment first.
Read more: the definition, the rules with a worked example, queries, and policies that change over time
Collaboration policy for LLM agents. For a fixed agent team, task, and execution setup, a collaboration policy Γ is an interaction topology along with a set of coordination rules that together guide or constrain task execution, shaping the joint distribution of agent activations, tool calls, and event exchanges over time.

Topology recovery. Take the activations that start and finish within a window [t, t′] and merge them by agent. The result is a weighted graph G[t,t′] = (V, E, w) whose edges connect agents that exchanged messages, weighted by message counts. The full topology is recovered once the records cover every edge the policy can create.

Execution rules. Consecutive activations whose starts fall within τ of each other and whose intervals overlap are grouped. For completed activations with start sh and finish fh:

h ≺ h′  iff  fh ≤ sh′ (execution order) h ∥τ h′  iff  |sh − sh′| ≤ τ (overlap)

Example trace

ActivationIntervalReceives from
ha,1[0, 1]Task input
hb,1[1, 4]ha,1
hc,1[1.5, 7]ha,1
ha,2[4, 5]hb,1
hb,2[5, 6]ha,2
ha,3[7, 8]hb,2, hc,1
hc,2[8, 10]ha,3
ha,4[10, 11]hc,2

Aggregation gives a ↔ b and a ↔ c, two messages per direction. With τ = 0.6, hb,1 and hc,1 are grouped: {a} → {b, c} → {a} → {b} → {a} → {c} → {a}.

Sequential

Apply ≺ along activation paths. ha,1 ≺ hc,1 ≺ ha,3 and ha,3 ≺ hc,2 ≺ ha,4 support the local order a → c → a.

Parallel

hc,1 overlaps hb,1, ha,2, and hb,2: a can receive b's result and send more work to b while c remains active.

Repetition

Follow message paths that return to later activations of the same agent. ha,1 → hb,1 → ha,2 → hb,2 → ha,3 contains a → b → a twice.

Conditional

The graph shows whom an agent sends to; the conditions behind each message are inferred from its input messages and subsequent sending behavior.

Joins

Observe which senders feed an agent's activations and how consistently those sets recur: recurring both-inputs support an all-input rule, either alone an any-input rule.

Queries across layers

Agents' activities induce directed edges (→) and implicit links (↦) across TEA. Queries traverse them in either direction, within or across layers; each path between two entities gives an interpretation of their relationship. Tracing task versions reveals who worked on a task, what they changed and in what order, and the linked evidence and environmental feedback; tracing further back to activations shows whom they interacted with and what information they received.

Relations and tools in TEA

AIGActor(h, u) ↦ h ↦ ih
Event flowe → h;  h → e′
TAGIdentityq ↦ k
Open / Closea → q;  a ↦ k
Edit / Updateq → a → q′
Splitq → a → {q′j}
Join{qj} → a → q′
TEATask action(h, u) ↦ a,  u ∈ UT
Environment(h, u) ↦ u,  u ∈ UZ
sendh(ih, e, j),  j ∈ J
attachh(h, u) ↦ φ
Feedback(h, u) ↦ Eu ⊆ E
Evidenceφ ↦ h;  φ ↦ q;  φ ↦ x
Three collaboration graphs run in sequence, their execution trace, and the policies recovered with perfect chunking versus equal-length chunking
Policies that change mid-task. (a) Collaboration policies for three milestones and (b) their execution trace. (c) Perfect chunking and (d) equal-length chunking, with their recovered policies. Chunks that cross phases mix interactions from different policies and distort the recovered topology - an opening for detecting policy changes and selecting aggregation windows.
A stochastic routing policy, its edge-probability matrix, one execution's AIG, and the distance between recovered edge frequencies and the true matrix as the run proceeds
Routing recovery. (a) A routing policy and (b) its weighted adjacency matrix of edge probabilities; (c) one execution's AIG and (d) the L∞ distance between the recovered edge-frequency matrix and (b) as the run proceeds. Crosses mark edges not yet run.

Infuse: From the Record to a Better Policy

TEA is also the interface through which agents communicate, transform tasks, and act in the environment. Infuse takes the trace of an episode - starting from an empty policy - and proposes the next one.

1

Execute and record

DT ← fexecute(V, Γ)

The team V attempts the task under policy Γ, and TEA records the collaboration as it unfolds.

2

Recover / discover

Γ̂ ← frecover(DT)

Aggregate activations by agent to recover the observed topology and infer candidate execution rules.

3

Analyze and propose

Γ′ ← fanalyze(Γ, Γ̂, DT)

Query the record to compare the recovered policy with the trace, assess coordination against task progress and environmental outcomes, and decide what to keep, add, revise, or leave open.

In the paper's implementation, a separate gpt-6-sol agent consolidates this guidance into global lessons for the team and private lessons for individual agents, forming Γ′. Infuse itself is agnostic to how Γ′ is produced.

Read more: how agents run on TEA

A team of agents collaborates asynchronously through TEA. During an activation, an agent acts on its environment view, its task-graph view, received messages, and private memory; its tool calls support communication, task transformations, evidence attachment, and environmental actions, all recorded through TEA. Three mechanisms control when agents activate:

  1. After each round of reasoning and tool use, the agent decides whether to continue its current activation or deactivate.
  2. The sleep operation ends the current activation and schedules reactivation after a specified delay.
  3. When sending a message, the agent specifies whether it should interrupt the recipient. Delivery activates an inactive recipient; if the recipient is active and interruption is requested, its current activation ends and a new one begins with the incoming message.

Results

Two domains, five gpt-5.6-terra agents with no predefined roles: code refactoring (modularize a redundant codebase; LoC reduction while passing fixed tests, six repositories, 100-step budget) and Silo-Bench (30 distributed reasoning tasks where every agent must answer a query that needs the whole team's information). A separate gpt-6-sol agent derives the Infuse guidance.

Does the shared task graph help?

Access to TAG and to direct communication is ablated in four conditions.

Code (shop)Silo-Bench
SettingPass ↑Rall ↑Rpass ↑Success ↑S ↑P ↑
TAG + comm4/546.943.00.770.810.85
TAG only3/535.335.20.770.810.83
Comm only3/520.219.50.600.730.78
Neither4/521.320.80.030.150.20

Code refactoring on shop: five runs per condition; Pass counts runs passing all tests; Rall and Rpass are mean LoC reductions (%) over all and passing runs. Silo-Bench: 30 tasks; S is the fraction of exactly correct answers, P the partial-credit score. Best values in bold.

Mean LoC reduction over agent steps on the shop repository for the four ablation conditions, with standard-deviation shading
LoC reduction over agent steps on shop: mean of five runs per condition, shading ±1 SD. The gap between the TAG conditions and the others widens during execution.

TEA's shared task graph outperforms direct communication alone on both benchmarks, while combining the two yields the strongest overall results.

Does Infuse improve the policy?

TEA with and without one Infuse update, five runs per repository and condition.

0.63 → 0.70
Average pass rate across the six refactoring repositories.
25.5 → 30.4%
Repository-average LoC reduction among passing runs; it rises in all six repositories.
0.77 → 0.83
Silo-Bench success with TAG + comm.
0.81 → 0.86
Silo-Bench exact correctness S with TAG + comm.

Infuse enables TEA to refine its collaboration policy from experience and improve overall task performance.

Read more: full Infuse tables, communication patterns under ablation, and five rounds of Infuse

Code refactoring before (TEA) and after one Infuse update (Inf.)

EasyMediumHardAverage
Codebaseshopplanledgerfreightstaypayroll
TEAInf.TEAInf.TEAInf.TEAInf.TEAInf.TEAInf.TEAInf.
Pass rate ↑ 0.800.80 0.800.80 0.600.80 0.600.80 0.400.60 0.600.40 0.630.70
LoC reduction (%) ↑ 43.0±10.256.6±8.9 28.3±5.931.6±3.7 22.1±3.425.1±1.7 20.6±4.022.1±3.5 20.2±2.624.8±4.4 18.8±2.522.2±9.0 25.5±9.230.4±13.3

Average is the unweighted repository mean; ± denotes sample SD across passing runs, or across repositories for Average. Better paired values in bold; ties unbolded. Pass rates rise on ledger, freight, and stay, remain unchanged on shop and plan, and only fall on payroll.

Silo-Bench before (TEA) and after one Infuse update (Inf.), 30 tasks

TAG + commTAG onlyComm onlyNeither
TEAInf.TEAInf.TEAInf.TEAInf.
Success ↑0.770.830.770.730.600.670.030.17
S ↑0.810.860.810.750.730.790.150.19
P ↑0.850.880.830.780.780.820.200.24

Bold marks the best value in each row. Infuse raises success in three of four settings; TAG-only success slightly falls from 0.77 to 0.73.

Interaction graphs of the five agents summed over five seeds for the full, no-TAG, no-communication, and neither conditions; the last two are empty
Who sent to whom under ablation, summed over five seeds per setting on one scale. Removing TAG raises traffic rather than lowering it: 244 events over 20 edges without TAG against 94 over 20 with it. Panels (c) and (d) are empty by construction - with send removed, there is no edge to draw.
Five rounds of Infuse updates: the team's interaction graph in each round and score, steps, and LoC reduction relative to the vanilla setting
Iterating Infuse over five rounds. Top: the team's interaction graph in each round as it inherits the previous round's lessons. Bottom: score, steps, and LoC reduction relative to the vanilla setting.

Back to top