Skip to content

Latest commit

 

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

agent-workflow

A working contract for a team of AI agents that hand real work to each other.

The tracker is not the point. The flow is.


The idea

Most attempts to coordinate several agents start by picking a tool -- a board, a queue, a database -- and then discovering that the hard part was never the tool. The hard part is answering questions like:

  • What does "I finished my part" actually mean, mechanically?
  • How does the next agent find out it is their turn?
  • Who is allowed to decide that something is out of scope?
  • What happens when an agent gets stuck and nobody notices?

Those answers are the workflow contract. They do not mention any particular product, and they do not change when you switch products.

A tracker such as Plane is an observation station: a place where the state of the work is visible, queryable, and hard to lose. Jira is one. Linear is one. A folder of YAML files could be one. What makes the system work is not which one you pick, but that every agent obeys the same contract on top of it.

This repository is organised to make that separation real:

flowchart TB
    subgraph L1["THE CONTRACT — the part that matters"]
        C1["Taking work, acknowledging, handing off<br/>Getting stuck, asking a human, sign-off"]
        C2["Agent-to-agent messaging<br/>one mention per message, self-contained, no echoes"]
    end

    subgraph L2["ADAPTERS — one per tracker, replaceable"]
        A1["plane-workflow<br/>shipped"]
        A2["abc-workflow<br/>not written yet"]
    end

    subgraph L3["OBSERVATION STATIONS — interchangeable"]
        O1["Plane board"]
        O2["some other Jira-like tracker"]
    end

    C1 --> A1
    C1 --> A2
    C2 -.->|"tracker-independent"| A1
    C2 -.-> A2
    A1 --> O1
    A2 --> O2
Loading

agent-relay sits entirely in the top layer -- it is about how agents talk, and it works the same whatever you track work in. plane-workflow is the first adapter. If you run something else, abc-workflow is the same contract with the state names and API calls swapped; the rules above it do not move.


The flow

One dispatch queue, one acknowledgement, one way out when a human is needed.

stateDiagram-v2
    direction LR
    [*] --> Backlog
    Backlog --> Todo: planned

    Todo --> InProgress: picked up = acknowledged
    InProgress --> Todo: handoff (5 actions, together)
    InProgress --> Blocked: needs a human
    Blocked --> Todo: human answers
    InProgress --> Acceptance: last step done, need:Docs
    Acceptance --> Todo: signed off, or rejected
    Todo --> Done: release-notes entry written
    Done --> [*]

    note right of Todo
        The watcher looks HERE and nowhere else.
        One assignee + one need: label = dispatchable.
    end note
Loading

Todo is the only dispatch queue. A card there with exactly one assignee and exactly one need: label is a unit of work someone can be woken for. A cron-driven daemon scans it and posts one message per person -- how often is whatever interval you install, which is why no number is quoted here.

In Progress is the acknowledgement. Moving a card there says "I have it", stops anyone else taking it, and puts you on the watcher's busy list. Nothing else counts.

Busy means busy everywhere. While you hold a card, the watcher sends you no Todo reminders at all -- including for other projects, because attention is single-threaded regardless of which board a card sits on. Those cards are not lost: they stay in Todo, and the watcher reports them to the project lead as queued behind you, so a card nobody is coming for is visible rather than silent.

A handoff is five things happening together -- state, assignee, label, the step list in the description, and a comment. Four out of five is a defect, and it is the most common one by a wide margin: mark the step done without reassigning and the next agent is never woken; reassign without marking and no human can tell how far the work got.

Blocked is how you ask for a human, and it is not optional. Sitting in In Progress waiting for an answer means the system believes you are working, and nobody can see that you are waiting.

A card can also be held back by another card. When planning finds that two cards edit the same files, it records a blocked_by relation between them, and the watcher will not dispatch a card whose blocker is unfinished. The point of using the tracker's own relation rather than a sentence in the description is that a sentence has to be read and obeyed, while a relation is enforced.

Acceptance is where a person signs off -- and the last comment there has to be written for that person, in plain language, in their language, not in agent shorthand.

Sign-off is not the last stop. A card arrives at Acceptance already carrying need:Docs and the agent that brought it there -- usually its QA. Accepting it is one action: the person signing off drops themselves from the assignees and moves it to Todo, and that agent writes the release-notes entry, then closes the card.

The entry is written once and lands in two places: the tracker page for that cycle, which the team reads on the board, and releases/<a.b.c>.md in this repository, which is what anyone outside the team reads. Same wording in both -- not two versions adapted to two audiences, because that is where drift starts. One entry per card, and no real names in either copy: this repository is public.


What is in here

skills/
  plane-workflow/       The contract, adapted to a Plane board
    SKILL.md              What an agent does once it is woken up
    PROJECT-SETUP.md      Board setup: 7 states + 6 need: labels + one cycle
    settings.example.json Copy to settings.json: the language this team writes in
    watcher/              The daemon that reads the board and wakes people
  agent-relay/          How agents message each other without deadlocking
    SKILL.md              Three hard rules, five templates, failure modes
    contact.json.example  The roster: single source of truth for handles
    settings.example.json Same file, its own copy -- a skill is loaded on its own
scripts/
  orchestration/        Reasoning about the whole board, rather than one skill's rules
    langgraph_shadow.py   Recomputes the watcher's verdict as a graph, diffs the two
    README.md             How to run the comparison, and how to reconcile it honestly

The two skills are independent. agent-relay is useful on its own if your agents talk in a group chat.

scripts/ is deliberately outside skills/: a skill directory gets symlinked into every agent's tree, and nothing that pulls dependencies belongs there.


Status

Piece State Notes
Workflow contract (plane-workflow/SKILL.md) In production Accepting, handoff, Blocked, rework, acceptance. Rewritten several times against real failures
Messaging contract (agent-relay/SKILL.md) In production Three hard rules; every one of them came from an incident
Board setup standard In production 7 states, 6 labels, one cycle per release, with a verification step people skip
Watcher (dispatch daemon) In production Read-only, cron-driven, unit-tested. Auto-detects which projects are in scope; skips agents who already hold a card, and cards blocked by an unfinished one
Acceptance backlog alerts In production Approximated from updated_at; see below
Blocked alerts In production Reports to the user after a threshold
Shadow mode In use Computes without sending. Built as the migration harness, and now what the LangGraph comparison runs against
LangGraph shadow pipeline Running as a comparison Recomputes the watcher's dispatch decision as a graph and diffs the two verdicts. Read-only, unit-tested, run on demand -- scripts/orchestration/
LangGraph orchestration Not started The shadow decides nothing: no LLM call, no card touched, nothing sent. Handing it the decision is a separate step
A second tracker adapter Not started The interface is implied by plane-workflow, not yet extracted

Known approximations, stated honestly:

  • "How long has this sat in Acceptance" uses updated_at, which really means "time since any activity" -- a single comment resets it. Getting it exact means reading the activity log for the last state change.
  • There is no cron wrapper; you install the line yourself, and cron-setup-check.sh generates one for your machine.
  • The watcher never writes to the tracker. That is permanent, not a gap.

Roadmap

LangGraph orchestration. The watcher today is deliberately dumb: it compares structured fields and wakes people. That is the right shape for dispatch, but it cannot reason about a plan that spans several cards, notice that two agents are about to collide, or re-plan when a step comes back rejected twice. A graph-based orchestrator is the intended home for that kind of decision.

The migration path is built, and the first half of it is already running. --shadow computes exactly what the watcher would do and writes to its own log without sending anything; scripts/orchestration/langgraph_shadow.py now recomputes that same decision as a LangGraph graph and diffs the two verdicts, so a divergence shows up as a failed comparison rather than as a wrong message someone receives. Its log is kept strictly separate from the production one, precisely so the two can never contaminate each other.

What runs today is a comparison, not an orchestrator: the graph makes no LLM call, decides nothing about who does what, and never touches a card. Handing it the decision is the step that has not been taken. The comparison also rots in silence -- change what the watcher dispatches and nothing warns you that the two have drifted -- so a card that touches dispatch judgement should carry "run --reconcile, exit 0" as an acceptance criterion. The commands, and how to reconcile against a board with something actually on it, are in scripts/orchestration/README.md.

The constraint that will not be relaxed: whatever orchestrates, the thing that touches the tracker stays read-only and explainable. An LLM deciding to move your cards around is not a feature.

A second adapter. plane-workflow is one adapter, not the architecture. Writing abc-workflow for another tracker is what will force the real interface out into the open -- right now it is implicit, and implicit interfaces are how you end up with the tracker leaking into the contract. The parts that should port unchanged:

  • one dispatch queue, one acknowledgement state, one blocked state, one acceptance queue
  • the five-part handoff
  • structured fields for the machine, prose for humans
  • a read-only scanner that wakes people and otherwise stays silent

The parts that are Plane-specific and will need rewriting: state and label names, PQL queries, and the MCP calls.


Requirements

  • A Plane workspace and an API key per agent. The key decides identity, visibility, and who currentUser() resolves to -- sharing one key destroys the handoff trail, because every action lands under the same account.
  • Python 3.9+ for the watcher. No third-party packages: it speaks JSON-RPC to plane-mcp-server over stdio using the standard library alone.
  • uvx to start that MCP server, plus a Telegram bot and a group chat if you want reminders actually delivered.
  • uv only if you want to run the LangGraph shadow comparison. It resolves LangGraph per run, so nothing is installed system-wide and the watcher on cron stays dependency-free.

Getting started

  1. Set up a board. Follow skills/plane-workflow/PROJECT-SETUP.md -- 7 states, 6 labels and one cycle, with a verification step at the end that people skip and should not.
  2. Choose your language. In each skill directory, copy settings.example.json to settings.json and set language to the language written out in English -- English, Japanese, Chinese. A bare name means that language's usual written form; put a qualifier in front if you need another, as in Traditional Chinese. Everything an agent writes at runtime -- card comments, group messages -- goes in that language. Each skill carries its own copy because a skill is loaded on its own and cannot read files above it. The files in this repository stay English regardless; the two are separate decisions, and conflating them means the person running the project has to read output in a language they do not work in.
  3. Fill in the roster. Copy skills/agent-relay/contact.json.example to contact.json and put your real handles and member ids in it. It is gitignored, and it is meant to be maintained by one person.
  4. Configure the watcher. Copy skills/plane-workflow/watcher/.env.example to a file outside this repository (~/.config/plane/watcher.env, chmod 600) and fill it in.
  5. Dry run first. python3 plane_watcher.py prints who would be reminded and sends nothing. That is the default, deliberately.
  6. Install cron. Run bash cron-setup-check.sh on the machine that will run it. It checks the environment, does a real dry run under a cron-like env -i, and prints a crontab line generated for that machine.
  7. Install the skills wherever your agent framework loads them from, so an agent that gets woken can actually read the contract.

Design principles

The daemon is read-only and stupid. It compares structured fields. No LLM, no free-text parsing, so every decision is reproducible and every bug is findable.

Silence is a feature. Cron runs it 96 times a day. Anything that "just reports it's fine" becomes 96 notifications a day, after which everyone mutes it -- including for the one message that mattered.

Configuration, not names. A project is in scope because it has a Todo state and need: labels, not because its name is on a list. Renaming a project used to stop the watcher dead; that was the watcher being wrong.

One source of truth, always. The roster holds handles. The skill holds the rules. The watcher names the skill rather than repeating it. Every copy drifts, and the copy is always the one someone reads.

Structured fields for the machine, prose for humans. The assignee says whose turn it is. The description says what is going on. When they disagree, that is a defect -- and it is caught by a person, not by the daemon.

Write for the human at the end. The person doing acceptance is the busiest in the loop and the only one who can make decisions. Anything they read is written in plain language.


A note on the rules that look oddly specific

Many of them are, deliberately. "At least 60 seconds between group messages" is there because 20 seconds was tried and still hit the rate limit. "Mark it in plain text, never a checkbox" is there because checkboxes are markup and agents editing HTML through an API break them. "A rule confirmed against exactly one sample is not a rule" is there because a claim in these very documents was disproved forty minutes after it was written.

Where a rule reads like scar tissue, it is. Change them if your environment differs -- but the failure each one prevents is described next to it, so you can see what you are giving up.

License

MIT. See LICENSE.

About

A workflow contract for teams of AI agents — dispatch, handoff, blocked and human sign-off, with pluggable tracker adapters (Plane shipped)

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages