Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
7ee64cf
feat(ship-check): surface-diff script for comparing MCP tool lists
aliasunder Oct 3, 2026
ccdf911
feat(ship-check): surface-diff prints a tool's definition as readable…
aliasunder Oct 3, 2026
275c103
feat(ship-check): tool-definition-reviewer agent and tool-definition-…
aliasunder Oct 3, 2026
84e2d88
docs: list the tool-definition reviewer and the bundled-script conven…
aliasunder Oct 3, 2026
8f84eb0
refactor(ship-check): surface-diff runs on Bun only
aliasunder Oct 3, 2026
131b208
fix(ship-check): the tool-definition reviewer stops at eight tools an…
aliasunder Oct 3, 2026
6ca1364
fix(ship-check): the reviewer's scope line cannot be mistaken for the…
aliasunder Oct 3, 2026
cd348fa
fix(ship-check): review findings on the surface-diff script and workf…
aliasunder Oct 3, 2026
cee7553
fix(ship-check): duplication threshold applies to the trimmed overlap…
aliasunder Oct 3, 2026
b1f6e97
docs: root README names all eight skills, the Bun prerequisite, and h…
aliasunder Oct 3, 2026
013a6fb
fix(ship-check): the tool-definition reviewer fails when its script c…
aliasunder Oct 3, 2026
4404242
style: clarify comments in surface-diff, workflow path shapes, and RE…
aliasunder Oct 3, 2026
4b93729
docs(ship-check): the reviewer's usage example defines its prompt lin…
aliasunder Oct 3, 2026
f6a2418
test(ship-check): cover every InputError path, CLI mode, and schema b…
aliasunder Oct 3, 2026
cec0249
fix(ship-check): the reviewer lists each traced failure as listed or …
aliasunder Oct 3, 2026
fdcb8ef
docs(ship-check): two skill sentences a test reader could take two ways
aliasunder Oct 3, 2026
df18884
fix(ship-check): surface-diff reports a removed duplicate line and tr…
aliasunder Oct 3, 2026
3f7bd43
fix(ship-check): surface-diff rejects a tool name that could carry sh…
aliasunder Oct 3, 2026
023e243
docs: the two READMEs and the marketplace description use plain punct…
aliasunder Oct 3, 2026
6354e56
fix(ship-check): the reviewer marks a schema-shadowed throw unreachab…
aliasunder Oct 3, 2026
efcc51a
fix(ship-check): --show quotes an unknown tool name the same way the …
aliasunder Oct 3, 2026
dc1b650
style: replace remaining em dash in plan-check marketplace description
aliasunder Oct 3, 2026
1f1e39f
ci: one release run at a time in each release workflow
aliasunder Oct 3, 2026
a1fbac0
docs: the remaining public copy uses plain punctuation in place of em…
aliasunder Oct 3, 2026
963f400
fix(ci): check out current main in manual release, not the dispatch-t…
aliasunder Oct 3, 2026
5b8f2a3
fix(ci): a queued manual release checks out the branch tip, and the t…
aliasunder Oct 3, 2026
efd214f
docs: restore the original punctuation in the READMEs, manifests, and…
aliasunder Oct 3, 2026
315985c
docs: restore the last three README lines the punctuation rewrite cha…
aliasunder Oct 3, 2026
3a74463
docs: the first adapting step in the root README is two plain sentenc…
aliasunder Oct 3, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
{
"name": "ship-check",
"source": "./plugins/ship-check",
"description": "Dedicated review agents for the ship-check pipeline. Five agent types (pr-reviewer, code-quality-reviewer, test-auditor, bug-checker, fresh-eyes) plus the orchestrator skill.",
"description": "Dedicated review agents for the ship-check pipeline. Six agent types (pr-reviewer, code-quality-reviewer, test-auditor, bug-checker, fresh-eyes, tool-definition-reviewer) plus eight skills, including the pipeline orchestrator.",
"version": "1.1.1",
"keywords": [
"code-review",
Expand Down
16 changes: 11 additions & 5 deletions .github/workflows/auto_release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,10 @@ jobs:
if: "!endsWith(github.actor, '[bot]')"
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
# This job only reads the checked-out files, so it keeps no token in the git config.
persist-credentials: false

- name: Extract version from tag
id: tag
Expand Down Expand Up @@ -69,13 +72,13 @@ jobs:
steps:
- name: Generate app token
id: app-token
uses: actions/create-github-app-token@v3
uses: actions/create-github-app-token@bcd2ba49218906704ab6c1aa796996da409d3eb1 # v3.2.0
with:
client-id: ${{ secrets.RELEASE_APP_CLIENT_ID }}
private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }}
permission-contents: write

- uses: actions/checkout@v7
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0
token: ${{ steps.app-token.outputs.token }}
Expand All @@ -93,19 +96,21 @@ jobs:

- name: Build artifacts
run: |
# Each plugin.json sits at ./plugins/<name>/.claude-plugin/plugin.json; two dirnames reach the plugin root.
while IFS= read -r PLUGIN_JSON; do
PLUGIN_DIR=$(dirname "$(dirname "$PLUGIN_JSON")")
PLUGIN_NAME=$(basename "$PLUGIN_DIR")

echo "Building $PLUGIN_NAME.zip..."
(cd "$PLUGIN_DIR" && zip -r "$GITHUB_WORKSPACE/${PLUGIN_NAME}.zip" . \
-x "evals/*" "dist/*" ".gitignore")
-x "evals/*" "dist/*" ".gitignore" "*/__tests__/*")

for SKILL_DIR in "$PLUGIN_DIR"/skills/*/; do
[ -d "$SKILL_DIR" ] || continue
SKILL_NAME=$(basename "$SKILL_DIR")
echo "Building $SKILL_NAME.skill..."
(cd "$PLUGIN_DIR/skills" && zip -r "$GITHUB_WORKSPACE/${SKILL_NAME}.skill" "$SKILL_NAME/")
(cd "$PLUGIN_DIR/skills" && zip -r "$GITHUB_WORKSPACE/${SKILL_NAME}.skill" "$SKILL_NAME/" \
-x "*/__tests__/*")
done
done < <(find . -path '*/.claude-plugin/plugin.json' -not -path './.claude-plugin/*')

Expand All @@ -123,6 +128,7 @@ jobs:
VERSION: ${{ steps.tag.outputs.version }}
run: bash .github/scripts/update-changelog.sh "$VERSION" /tmp/release-notes.md

# The tag push runs on a detached HEAD, so the changelog commit is moved onto main.
- name: Commit changelog update
run: |
git config user.name "github-actions[bot]"
Expand Down
19 changes: 15 additions & 4 deletions .github/workflows/manual_release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -16,20 +16,29 @@ on:
permissions:
contents: write

# Two dispatches started together would compute the same next version and race on the push,
# so a second run waits for the first.
concurrency:
group: ${{ github.workflow }}
cancel-in-progress: false

jobs:
release:
runs-on: ubuntu-latest
steps:
Comment thread
aliasunder marked this conversation as resolved.
- name: Generate app token
id: app-token
uses: actions/create-github-app-token@v3
uses: actions/create-github-app-token@bcd2ba49218906704ab6c1aa796996da409d3eb1 # v3.2.0
with:
client-id: ${{ secrets.RELEASE_APP_CLIENT_ID }}
private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }}
permission-contents: write

- uses: actions/checkout@v7
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
# Check out the branch tip as it is when this run starts. The default is the commit at
# dispatch time, and a run that waited in the queue would find that commit one release behind.
ref: ${{ github.ref }}
fetch-depth: 0
token: ${{ steps.app-token.outputs.token }}

Expand Down Expand Up @@ -93,19 +102,21 @@ jobs:

- name: Build artifacts
run: |
# Each plugin.json sits at ./plugins/<name>/.claude-plugin/plugin.json; two dirnames reach the plugin root.
while IFS= read -r PLUGIN_JSON; do
PLUGIN_DIR=$(dirname "$(dirname "$PLUGIN_JSON")")
PLUGIN_NAME=$(basename "$PLUGIN_DIR")

echo "Building $PLUGIN_NAME.zip..."
(cd "$PLUGIN_DIR" && zip -r "$GITHUB_WORKSPACE/${PLUGIN_NAME}.zip" . \
-x "evals/*" "dist/*" ".gitignore")
-x "evals/*" "dist/*" ".gitignore" "*/__tests__/*")

for SKILL_DIR in "$PLUGIN_DIR"/skills/*/; do
[ -d "$SKILL_DIR" ] || continue
SKILL_NAME=$(basename "$SKILL_DIR")
echo "Building $SKILL_NAME.skill..."
(cd "$PLUGIN_DIR/skills" && zip -r "$GITHUB_WORKSPACE/${SKILL_NAME}.skill" "$SKILL_NAME/")
(cd "$PLUGIN_DIR/skills" && zip -r "$GITHUB_WORKSPACE/${SKILL_NAME}.skill" "$SKILL_NAME/" \
-x "*/__tests__/*")
done
done < <(find . -path '*/.claude-plugin/plugin.json' -not -path './.claude-plugin/*')

Expand Down
25 changes: 25 additions & 0 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
name: Test

on:
push:
branches: [main]
pull_request:

permissions:
contents: read

jobs:
test:
runs-on: ubuntu-latest
# The suite runs in under a second; five minutes covers a slow runner start.
timeout-minutes: 5
name: script tests
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false

- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2.2.0

# The scripts have no dependencies, so there is nothing to install.
- run: bun test plugins
Comment thread
aliasunder marked this conversation as resolved.
9 changes: 9 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ that bundle agents, skills, commands, and hooks as distributable packages.
workflows/
auto_release.yml # v* tag push → validate versions, build artifacts, GitHub release
manual_release.yml # workflow_dispatch → bump version, tag, build, release
test.yml # push to main and PRs → run the bundled scripts' tests
umm_review.yml # PR review via umm-actually (configurable via repo variables)
scripts/ # Shared release-note and changelog helpers
plugins/
Expand All @@ -28,6 +29,7 @@ plugins/
test-auditor.md
bug-checker.md
fresh-eyes.md # Phase 2 — stranger read, report only
tool-definition-reviewer.md # On demand — MCP tool definitions, report only
skills/ # Skills (SKILL.md in subdirectories)
ship-check/ # Pipeline orchestrator
pr-review/ # Phase 1 — correctness, security, conditional checks
Expand All @@ -36,6 +38,8 @@ plugins/
test-audit/ # Phase 4 — test quality + coverage gaps
bug-check/ # Phase 5 — systematic bug hunt
pr-monitor/ # Phase 6 — CI, bot comments, merge readiness
tool-definition-review/ # On demand — MCP tool-definition review (report only)
scripts/ # surface-diff.ts and its __tests__/
README.md
plan-check/ # Pre-implementation plan review plugin
.claude-plugin/
Expand Down Expand Up @@ -70,6 +74,11 @@ SECURITY.md # Vulnerability reporting policy
- **Plugin manifests** use semver versioning
- Agent `tools:` fields are allowlists — omit to give all tools, list explicitly to restrict
- Agent `skills:` preloads skill content from any installed plugin or `~/.claude/skills/`
- **Bundled scripts** live in a skill's `scripts/` directory as TypeScript (`.ts`)
with no npm dependencies, run with Bun. Their tests (`*.test.ts` files) live in
`scripts/__tests__/`; run them with `bun test plugins` (`plugins` is the directory
Bun searches). The release archives leave `__tests__` out, through the `-x`
patterns on the `zip` commands in both release workflows.

## Skill authoring

Expand Down
18 changes: 10 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,14 +16,14 @@ Personal plugin marketplace for Claude Code and Claude Cowork — review agents,

| Plugin | Description |
|--------|-------------|
| [ship-check](plugins/ship-check/) | Post-implementation review pipeline: five dedicated review agents (pr-reviewer, code-quality-reviewer, test-auditor, bug-checker, fresh-eyes) plus seven skills covering PR review, code quality, test audit, bug hunting, stranger reads, and PR monitoring |
| [ship-check](plugins/ship-check/) | Post-implementation review pipeline: six dedicated review agents (pr-reviewer, code-quality-reviewer, test-auditor, bug-checker, fresh-eyes, tool-definition-reviewer) plus eight skills: the pipeline orchestrator and one each for PR review, code quality, test audit, bug hunting, stranger reads, MCP tool-definition review, and PR monitoring |
| [plan-check](plugins/plan-check/) | Pre-implementation plan review: a fresh-eyes agent (plan-reviewer) plus the plan-review skill — premise audit, alternatives comparison, guard/control arithmetic, concurrent-writer analysis, mechanism-cost proportionality, and verification-plan safety before any code exists |

## Structure

- **`.claude-plugin/marketplace.json`** — marketplace manifest listing all plugins
- **`plugins/`** — the plugins themselves (agents, skills, manifests)
- **`.github/workflows/`** — release automation and PR review (`umm_review.yml`)
- **`.github/workflows/`** — release automation, script tests (`test.yml`), and PR review (`umm_review.yml`)

## Installation

Expand All @@ -34,23 +34,25 @@ claude plugin marketplace add aliasunder/agent-plugins
claude plugin install ship-check@agent-plugins
```

Or browse via `/plugin > Discover`.
Or run `/plugin` inside Claude Code and open the Discover tab.

For local development, register the repo directory instead:

```
claude plugin marketplace add ~/Code/agent-plugins
```

The ship-check tool-definition reviewer runs a bundled script, which needs [Bun](https://bun.sh) installed.

## Adapting for your own use

If you want to use these plugins as a starting point:

1. Replace the vault-cortex loading steps ([vault-cortex](https://github.com/aliasunder/vault-cortex)) in the agents and skills — `vault_read_note` calls on `Reference/code-standards-*.md` and `vault_memory_recall`/`vault_get_memory` preference retrieval — with your own standards docs and memory/preference source (or remove them)
2. Remove `mcp__claude_ai_Vault_Cortex__*` entries from the agents' `tools:` allowlists if you dropped vault-cortex
3. Install [fable-mode](https://github.com/mrtooher/fable-mode) as a skill, or remove it from the agents' `skills:` lists
4. Install the [sequential-thinking](https://github.com/modelcontextprotocol/servers/tree/main/src/sequentialthinking) MCP server — the agents use it for triage reasoning (or drop it from their `tools:` allowlists)
5. The pipeline structure, review dimensions, and procedural triggers in the skills are workflow-agnostic and should transfer as-is
1. Replace the [vault-cortex](https://github.com/aliasunder/vault-cortex) loading steps in the agents and skills with your own standards docs and memory or preference source, or remove them. Those steps are the `vault_read_note` calls on `Reference/code-standards-*.md` and the `vault_memory_recall`/`vault_get_memory` preference retrieval.
2. If you dropped vault-cortex, remove its entries from the agents' `tools:` allowlists. Claude Code names an MCP tool `mcp__<server>__<tool>`, so the vault-cortex entries start with `mcp__claude_ai_Vault_Cortex__` or `mcp__vault-cortex__` (the same server, connected two ways).
3. Install [fable-mode](https://github.com/mrtooher/fable-mode) as a skill, or remove it from the agents' `skills:` lists.
4. Install the [sequential-thinking](https://github.com/modelcontextprotocol/servers/tree/main/src/sequentialthinking) MCP server. The skills tell the agents to call the server's `sequentialthinking` tool before they decide what to do with a finding. To go without the server, drop that tool from the agents' `tools:` allowlists and the skills' `allowed-tools:` lists, and remove the skill steps that call it.
5. The pipeline structure, review dimensions, and procedural triggers in the skills are workflow-agnostic and should transfer as-is.

## License

Expand Down
2 changes: 1 addition & 1 deletion plugins/ship-check/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"name": "ship-check",
"version": "1.1.1",
"description": "Dedicated review agents for the ship-check pipeline. Each agent approaches the codebase without prior context and returns structured findings; the four phase agents load project conventions and user preferences independently, and fresh-eyes deliberately loads none."
"description": "Dedicated review agents for the ship-check pipeline. Each agent approaches the codebase without prior context and returns structured findings; the four phase agents that load conventions do so independently, fresh-eyes deliberately loads none, and tool-definition-reviewer reviews MCP tool definitions on demand."
}
44 changes: 36 additions & 8 deletions plugins/ship-check/README.md
Original file line number Diff line number Diff line change
@@ -1,33 +1,43 @@
# ship-check

Dedicated review agents for the ship-check pipeline. Each agent approaches the
codebase without prior context and returns structured findings. The five phase
codebase without prior context and returns structured findings. The four phase
agents that load conventions do so independently; `fresh-eyes` (Phase 2)
deliberately loads none.
deliberately loads none. A sixth agent, `tool-definition-reviewer`, is dispatched
on demand and is not a pipeline phase.

## Agents

| Agent | Phase | Color | Role |
|-------|-------|-------|------|
| `pr-reviewer` | 1 | cyan | Correctness, security, conditional checks (Tool Description Quality Score (TDQS), feature surface, stale paths) |
| `pr-reviewer` | 1 | cyan | Correctness, security, conditional checks (Tool Definition Quality Score (TDQS), feature surface, stale paths) |
| `fresh-eyes` | 2 | purple | Stranger read: every place a newcomer pauses, per function. Report only — no conventions, no edits, no history. Pauses feed into Phase 3. |
| `code-quality-reviewer` | 3 | green | Naming, structure, comments, simplicity, module conventions. Resolves fresh-eyes pauses. |
| `test-auditor` | 4 | yellow | Test quality audit + coverage gap analysis (writes missing tests) |
| `bug-checker` | 5 | red | 7-dimension systematic bug hunt (description-vs-code, SQL, type safety, etc.) |
| `tool-definition-reviewer` | on demand | orange | MCP tool definitions read as the client receives them: TDQS rubric marks, text changed in tools nobody meant to touch, dropped facts, description text that repeats the schema, and failures the description never lists. Report only. |

Phase 6 (pr-monitor) runs inline in the orchestrator — it needs user interaction
and continuous monitoring, which agents can't do. `fresh-eyes` can also be dispatched
standalone to see what a newcomer experiences without the pipeline.

`tool-definition-reviewer` is not dispatched by the pipeline. Dispatch it yourself
when a change touches an MCP server's tool descriptions or input schemas.

## External Dependencies

Each agent preloads skills via `skills:` frontmatter. The `pr-review`,
`code-quality`, `test-audit`, `bug-check`, and `fresh-eyes` skills are bundled in
this plugin; [fable-mode](https://github.com/mrtooher/fable-mode) is external and
must be installed separately (e.g. in `~/.claude/skills/`). `fresh-eyes` preloads
only its own skill and uses no MCP tools.
`code-quality`, `test-audit`, `bug-check`, `fresh-eyes`, and
`tool-definition-review` skills are bundled in this plugin;
[fable-mode](https://github.com/mrtooher/fable-mode) is external and must be
installed separately (e.g. in `~/.claude/skills/`). `fresh-eyes` and
`tool-definition-reviewer` each preload only their own skill and use no MCP tools.

The `tool-definition-review` skill bundles one script, `scripts/surface-diff.ts`.
It has no dependencies and runs with [Bun](https://bun.sh). Without Bun the agent
reports the review as `failed`.

The convention-loading phase agents (all except `fresh-eyes`) also use MCP tools
The four convention-loading phase agents also use MCP tools
loaded at runtime via `ToolSearch`:

- `vault_get_memory` ([vault-cortex](https://github.com/aliasunder/vault-cortex) MCP) — user preferences
Expand All @@ -51,3 +61,21 @@ it needs the file list in its prompt:
```
Agent({ subagent_type: "ship-check:fresh-eyes", prompt: "Read src/a.ts and src/b.ts at <sha> as a stranger..." })
```

`tool-definition-reviewer` needs the tool list as a file (called a "surface" in the
prompt): the JSON a client gets from the MCP `tools/list` method, or a snapshot the
project commits to its repository. Give it the file from before the change as well,
when there is one:

```
Agent({ subagent_type: "ship-check:tool-definition-reviewer", prompt: "Current surface: /tmp/tools-now.json\nBase surface: /tmp/tools-before.json\nIntended tools: search_notes, read_note\nRepository root: /path/to/server" })
```

Only `Current surface` is required. Each other line unlocks checks:

- `Base surface` is the same file from before the change. Without it the agent
reviews every tool and skips the checks that compare the two files.
- `Intended tools` names the tools the change means to alter. The agent reports a
changed tool outside this list as an unintended change.
- `Repository root` is where the server's source lives. The agent traces each
tool's handler there to find failures the description does not list.
Loading
Loading