A working contract for a team of AI agents that hand real work to each other.
The tracker is not the point. The flow is.
Most attempts to coordinate several agents start by picking a tool -- a board, a queue, a database -- and then discovering that the hard part was never the tool. The hard part is answering questions like:
- What does "I finished my part" actually mean, mechanically?
- How does the next agent find out it is their turn?
- Who is allowed to decide that something is out of scope?
- What happens when an agent gets stuck and nobody notices?
Those answers are the workflow contract. They do not mention any particular product, and they do not change when you switch products.
A tracker such as Plane is an observation station: a place where the state of the work is visible, queryable, and hard to lose. Jira is one. Linear is one. A folder of YAML files could be one. What makes the system work is not which one you pick, but that every agent obeys the same contract on top of it.
This repository is organised to make that separation real:
flowchart TB
subgraph L1["THE CONTRACT — the part that matters"]
C1["Taking work, acknowledging, handing off<br/>Getting stuck, asking a human, sign-off"]
C2["Agent-to-agent messaging<br/>one mention per message, self-contained, no echoes"]
end
subgraph L2["ADAPTERS — one per tracker, replaceable"]
A1["plane-workflow<br/>shipped"]
A2["abc-workflow<br/>not written yet"]
end
subgraph L3["OBSERVATION STATIONS — interchangeable"]
O1["Plane board"]
O2["some other Jira-like tracker"]
end
C1 --> A1
C1 --> A2
C2 -.->|"tracker-independent"| A1
C2 -.-> A2
A1 --> O1
A2 --> O2
agent-relay sits entirely in the top layer -- it is about how agents talk, and
it works the same whatever you track work in. plane-workflow is the first
adapter. If you run something else, abc-workflow is the same contract with the
state names and API calls swapped; the rules above it do not move.
One dispatch queue, one acknowledgement, one way out when a human is needed.
stateDiagram-v2
direction LR
[*] --> Backlog
Backlog --> Todo: planned
Todo --> InProgress: picked up = acknowledged
InProgress --> Todo: handoff (5 actions, together)
InProgress --> Blocked: needs a human
Blocked --> Todo: human answers
InProgress --> Acceptance: last step done, need:Docs
Acceptance --> Todo: signed off, or rejected
Todo --> Done: release-notes entry written
Done --> [*]
note right of Todo
The watcher looks HERE and nowhere else.
One assignee + one need: label = dispatchable.
end note
Todo is the only dispatch queue. A card there with exactly one assignee and
exactly one need: label is a unit of work someone can be woken for. A
cron-driven daemon scans it and posts one message per person -- how often is
whatever interval you install, which is why no number is quoted here.
In Progress is the acknowledgement. Moving a card there says "I have it", stops anyone else taking it, and puts you on the watcher's busy list. Nothing else counts.
Busy means busy everywhere. While you hold a card, the watcher sends you no Todo reminders at all -- including for other projects, because attention is single-threaded regardless of which board a card sits on. Those cards are not lost: they stay in Todo, and the watcher reports them to the project lead as queued behind you, so a card nobody is coming for is visible rather than silent.
A handoff is five things happening together -- state, assignee, label, the step list in the description, and a comment. Four out of five is a defect, and it is the most common one by a wide margin: mark the step done without reassigning and the next agent is never woken; reassign without marking and no human can tell how far the work got.
Blocked is how you ask for a human, and it is not optional. Sitting in In Progress waiting for an answer means the system believes you are working, and nobody can see that you are waiting.
A card can also be held back by another card. When planning finds that two
cards edit the same files, it records a blocked_by relation between them, and
the watcher will not dispatch a card whose blocker is unfinished. The point of
using the tracker's own relation rather than a sentence in the description is
that a sentence has to be read and obeyed, while a relation is enforced.
Acceptance is where a person signs off -- and the last comment there has to be written for that person, in plain language, in their language, not in agent shorthand.
Sign-off is not the last stop. A card arrives at Acceptance already carrying
need:Docs and the agent that brought it there -- usually its QA. Accepting it is
one action: the person signing off drops themselves from the assignees and moves
it to Todo, and that agent writes the release-notes entry, then closes the card.
The entry is written once and lands in two places: the tracker page for that
cycle, which the team reads on the board, and releases/<a.b.c>.md in this
repository, which is what anyone outside the team reads. Same wording in both --
not two versions adapted to two audiences, because that is where drift starts.
One entry per card, and no real names in either copy: this repository is
public.
skills/
plane-workflow/ The contract, adapted to a Plane board
SKILL.md What an agent does once it is woken up
PROJECT-SETUP.md Board setup: 7 states + 6 need: labels + one cycle
settings.example.json Copy to settings.json: the language this team writes in
watcher/ The daemon that reads the board and wakes people
agent-relay/ How agents message each other without deadlocking
SKILL.md Three hard rules, five templates, failure modes
contact.json.example The roster: single source of truth for handles
settings.example.json Same file, its own copy -- a skill is loaded on its own
scripts/
orchestration/ Reasoning about the whole board, rather than one skill's rules
langgraph_shadow.py Recomputes the watcher's verdict as a graph, diffs the two
README.md How to run the comparison, and how to reconcile it honestly
The two skills are independent. agent-relay is useful on its own if your
agents talk in a group chat.
scripts/ is deliberately outside skills/: a skill directory gets symlinked
into every agent's tree, and nothing that pulls dependencies belongs there.
| Piece | State | Notes |
|---|---|---|
Workflow contract (plane-workflow/SKILL.md) |
In production | Accepting, handoff, Blocked, rework, acceptance. Rewritten several times against real failures |
Messaging contract (agent-relay/SKILL.md) |
In production | Three hard rules; every one of them came from an incident |
| Board setup standard | In production | 7 states, 6 labels, one cycle per release, with a verification step people skip |
| Watcher (dispatch daemon) | In production | Read-only, cron-driven, unit-tested. Auto-detects which projects are in scope; skips agents who already hold a card, and cards blocked by an unfinished one |
| Acceptance backlog alerts | In production | Approximated from updated_at; see below |
| Blocked alerts | In production | Reports to the user after a threshold |
| Shadow mode | In use | Computes without sending. Built as the migration harness, and now what the LangGraph comparison runs against |
| LangGraph shadow pipeline | Running as a comparison | Recomputes the watcher's dispatch decision as a graph and diffs the two verdicts. Read-only, unit-tested, run on demand -- scripts/orchestration/ |
| LangGraph orchestration | Not started | The shadow decides nothing: no LLM call, no card touched, nothing sent. Handing it the decision is a separate step |
| A second tracker adapter | Not started | The interface is implied by plane-workflow, not yet extracted |
Known approximations, stated honestly:
- "How long has this sat in Acceptance" uses
updated_at, which really means "time since any activity" -- a single comment resets it. Getting it exact means reading the activity log for the last state change. - There is no cron wrapper; you install the line yourself, and
cron-setup-check.shgenerates one for your machine. - The watcher never writes to the tracker. That is permanent, not a gap.
LangGraph orchestration. The watcher today is deliberately dumb: it compares structured fields and wakes people. That is the right shape for dispatch, but it cannot reason about a plan that spans several cards, notice that two agents are about to collide, or re-plan when a step comes back rejected twice. A graph-based orchestrator is the intended home for that kind of decision.
The migration path is built, and the first half of it is already running.
--shadow computes exactly what the watcher would do and writes to its own log
without sending anything; scripts/orchestration/langgraph_shadow.py now
recomputes that same decision as a LangGraph graph and diffs the two verdicts,
so a divergence shows up as a failed comparison rather than as a wrong message
someone receives. Its log is kept strictly separate from the production one,
precisely so the two can never contaminate each other.
What runs today is a comparison, not an orchestrator: the graph makes no LLM
call, decides nothing about who does what, and never touches a card. Handing it
the decision is the step that has not been taken. The comparison also rots in
silence -- change what the watcher dispatches and nothing warns you that the two
have drifted -- so a card that touches dispatch judgement should carry "run
--reconcile, exit 0" as an acceptance criterion. The commands, and how to
reconcile against a board with something actually on it, are in
scripts/orchestration/README.md.
The constraint that will not be relaxed: whatever orchestrates, the thing that touches the tracker stays read-only and explainable. An LLM deciding to move your cards around is not a feature.
A second adapter. plane-workflow is one adapter, not the architecture.
Writing abc-workflow for another tracker is what will force the real interface
out into the open -- right now it is implicit, and implicit interfaces are how
you end up with the tracker leaking into the contract. The parts that should
port unchanged:
- one dispatch queue, one acknowledgement state, one blocked state, one acceptance queue
- the five-part handoff
- structured fields for the machine, prose for humans
- a read-only scanner that wakes people and otherwise stays silent
The parts that are Plane-specific and will need rewriting: state and label names, PQL queries, and the MCP calls.
- A Plane workspace and an API key per agent. The key decides identity,
visibility, and who
currentUser()resolves to -- sharing one key destroys the handoff trail, because every action lands under the same account. - Python 3.9+ for the watcher. No third-party packages: it speaks JSON-RPC to
plane-mcp-serverover stdio using the standard library alone. uvxto start that MCP server, plus a Telegram bot and a group chat if you want reminders actually delivered.uvonly if you want to run the LangGraph shadow comparison. It resolves LangGraph per run, so nothing is installed system-wide and the watcher on cron stays dependency-free.
- Set up a board. Follow
skills/plane-workflow/PROJECT-SETUP.md-- 7 states, 6 labels and one cycle, with a verification step at the end that people skip and should not. - Choose your language. In each skill directory, copy
settings.example.jsontosettings.jsonand setlanguageto the language written out in English --English,Japanese,Chinese. A bare name means that language's usual written form; put a qualifier in front if you need another, as inTraditional Chinese. Everything an agent writes at runtime -- card comments, group messages -- goes in that language. Each skill carries its own copy because a skill is loaded on its own and cannot read files above it. The files in this repository stay English regardless; the two are separate decisions, and conflating them means the person running the project has to read output in a language they do not work in. - Fill in the roster. Copy
skills/agent-relay/contact.json.exampletocontact.jsonand put your real handles and member ids in it. It is gitignored, and it is meant to be maintained by one person. - Configure the watcher. Copy
skills/plane-workflow/watcher/.env.exampleto a file outside this repository (~/.config/plane/watcher.env,chmod 600) and fill it in. - Dry run first.
python3 plane_watcher.pyprints who would be reminded and sends nothing. That is the default, deliberately. - Install cron. Run
bash cron-setup-check.shon the machine that will run it. It checks the environment, does a real dry run under a cron-likeenv -i, and prints a crontab line generated for that machine. - Install the skills wherever your agent framework loads them from, so an agent that gets woken can actually read the contract.
The daemon is read-only and stupid. It compares structured fields. No LLM, no free-text parsing, so every decision is reproducible and every bug is findable.
Silence is a feature. Cron runs it 96 times a day. Anything that "just reports it's fine" becomes 96 notifications a day, after which everyone mutes it -- including for the one message that mattered.
Configuration, not names. A project is in scope because it has a Todo state
and need: labels, not because its name is on a list. Renaming a project used to
stop the watcher dead; that was the watcher being wrong.
One source of truth, always. The roster holds handles. The skill holds the rules. The watcher names the skill rather than repeating it. Every copy drifts, and the copy is always the one someone reads.
Structured fields for the machine, prose for humans. The assignee says whose turn it is. The description says what is going on. When they disagree, that is a defect -- and it is caught by a person, not by the daemon.
Write for the human at the end. The person doing acceptance is the busiest in the loop and the only one who can make decisions. Anything they read is written in plain language.
Many of them are, deliberately. "At least 60 seconds between group messages" is there because 20 seconds was tried and still hit the rate limit. "Mark it in plain text, never a checkbox" is there because checkboxes are markup and agents editing HTML through an API break them. "A rule confirmed against exactly one sample is not a rule" is there because a claim in these very documents was disproved forty minutes after it was written.
Where a rule reads like scar tissue, it is. Change them if your environment differs -- but the failure each one prevents is described next to it, so you can see what you are giving up.
MIT. See LICENSE.