Releases: escoffier-labs/brigade
Release list
v0.25.0
v0.24.0
Added
brigade version --componentsreports managed native component installation state, with Windows native acceptance for supported paths.- GraphTrail v0.4.0 and MiseLedger v0.6.0 source now live under
engines/. Standalone repositories and release pipelines remain unchanged until Phase 4. brigade.code-reference.v1defines exact code-reference evidence lookups before lexical fallback.brigade code sync|context|impactandbrigade evidence crawl|search|doctornow execute the imported engines.brigade searchaliases remain for two minor releases or 90 days, whichever is longer.
Fixed
brigade runterminalizes interrupted and stale runs, refuses read-only execution on incapable seats, and rejects empty tasks.
Full details: CHANGELOG.md
v0.23.3
v0.23.2
v0.23.1
Added
- Versioned component manifest v1 with pinned GraphTrail platform artifacts and provenance checks. (#353 / PR #367)
brigade setupinstalls those pinned native components transactionally, with dry-run, offline cache, rollback, smoke checks, and state recording. (#355 / PR #368)
Changed
- Clarified cross-platform installation documentation. (#351)
v0.23.0
Fixed
- ACPX late permission prompts after a completed
end_turnanswer no longer discard the usable final result. Brigade now preserves the terminal text, stdout/stderr, exit code, session/request metadata, and records a typedtransport_warningwhile marking the workerokwith warning instead of failing the run. Post-permission assistant chunks and permission errors without pre-final text no longer qualify as usable-with-warning. Pre-final permission failures remain failed. Stderr fallback infers late-32072only when the stream has no structured JSON-RPC error; any structured error takes precedence over stderr markers. (#337) proc.run()decodes child stdout and stderr as explicit UTF-8 on every platform, includingTimeoutExpiredpartial output, normalizes missing streams to empty text, preserves recoverable bytes with deterministic UTF-8 replacement, exposes explicit decode status onproc.Result, and surfaces typed harness decode failures throughrun_agentand ACP without counting a model verdict. Provider, output-validation, and chronology failures stay distinct. (#336)- Run control selects a supported local transport per platform: Unix-domain sockets remain the default on Linux and macOS when the socket path fits platform limits; otherwise Brigade falls back to authenticated loopback TCP with an ephemeral port. Run metadata records a typed
control_transportdescriptor for clients, legacycontrol_socketremains for Unix runs, and malformed descriptors fail closed without leaking private paths or owner tokens into public summaries. (#341) - Agent adapter detection and dispatch now share one resolved executable identity on Windows, so doctor can mark a seat runnable without
run_agentlater failing withcommand not foundafter discarding an npm shim. Only native.exefiles and extensionless paths with a verified PE signature are runnable;.cmd,.bat, and unverified npm shims fail before inference withunsupported-command-shimand path-safe remediation to put the package native executable directory earlier on PATH. Launched provider workspace-trust refusals classify asprovider-preflightinstead of command resolution, and remaining process-creation failures no longer leak absolute paths. (#340)
Added
-
brigade repos adoptionseparates configured harness wiring from observed work-loop use across the local repository fleet. It reportsunwired,partial,advisory-only,enforced-idle,active,bypassed, andstalerows, correlates Claude session writes with brief, verify, outcome, GraphTrail, MiseLedger, and handoff evidence, exposes fleet denominators and stable JSON monitor keys, and provides a read-onlyrepairplan. (#317) -
brigade operator checkupaccepts repeatable--surfaceselectors, lists stable surface names with--list-surfaces, and provides anevidence-looppreset for work receipt integrity and outcome capture, GraphTrail health and receipt deltas, and MiseLedger work-receipt import state. Scoped JSON separatesselected_readyfrom unevaluatedoverall_readyand reports selected, skipped, and per-surface elapsed data. The default six-doctor checkup is unchanged. (#298) -
The versioned
harness-contract.v1schema provides safe availability and opt-in version probes with bounded, redacted output. Codex Desktop and Cursor GUI remain external-only, while Antigravity availability resolves only theagyandantigravitycommand candidates without executing vendor commands. (#342) -
Project-scoped Claude Code work-loop hooks can brief once per session and repository, route direct verification commands through Brigade, and check completion receipts after write work.
brigade work hooks install|update|status|uninstallmanages the package without replacing unrelated Claude settings or hooks. -
brigade receipts export miseledger --jsonprints a typedbrigade.miseledger_export_result.v1summary (statusempty,nothing-new,exported, orfailed, plus export/skip/error counts) to stdout while writing JSONL to a named--outfile. A zero-receipt non-JSON export exits 0 with no output.--fleet --jsonruns the same export across enabled[[repo]]entries from.brigade/repos.toml, aggregates rows into one batch, keeps per-repo cursors, imports valid partial output at most once, returnsbrigade.miseledger_fleet_export_result.v1with per-repo counts and privacy-safe errors, and treats aggregate import failure as top-levelfailedwithout incrementingfailed_count. (#254) -
Route as a third outcome-ledger cohort axis (Phase 1: record and surface, no scoring change).
brigade outcome capturestamps a coarse route manifest (followed,path,size, sortedsignals,coverage) and aroute_fingerprinton each record, resolved from the owningbrigade run's route. A bare verify-capture with no owning run is the honest unrouted cohort.brigade outcome explainsplits records into routed-vs-unrouted cohorts with a help-rate, so the ledger can finally answer whether runs that followed a composed route verify greener. Ratchet unchanged; seedocs/design/route-outcome-cohort.md. -
Recency-weighted outcome ranking (context-aware outcomes follow-up):
outcome rank --recency(default half-life 45 days,--recency-half-life DAYSto override) weights the score it sorts by over recency-weighted counts, so a signal's weight halves every half-life and credit earned under a drifted-away environment fades without rewriting the append-only log. It applies to the pooled score by default and the capability-shrunk score with--by-capability; the two flags compose. Off by default, so rank output stays byte-identical. The Wilson bound and shrinkage prior now accept fractional counts; the promotion ratchet never uses recency. See docs/design/context-blind-spot.md.
v0.22.0
Fixed
- Security health and release readiness now honor the latest accepted-risk closeout by exact finding fingerprint, while new or changed findings become active again. Harness-wiring health also follows the configured template-inclusion policy instead of blocking a release on excluded managed-template URLs.
- Route calibration from a task corpus: a docs rewrite (
rewrite the QUICKSTART) no longer takes the full code route (a barerewrite/redesignno longer beats a docs hint), UI wordspanel/dashboard/sidebar/widget/view/tooltipnow pull the UI review lenses, and a migration earns tests the way an auth change does. - Route derivation no longer misroutes conventional-commit code work as docs:
fix(install): ... referencing docsandfeat(skills): ship ... CHANGELOG.mdroute as code, while prose likefix typo in READMEandfix(docs): ...stay docs. A code fix losing its review lenses to a stray "docs" keyword was the one misroute direction that cost coverage.
Added
brigade run --route-signal +auth-surface/~ship-requestedand the same flag onbrigade routeforce-add or suppress a derived route signal when the heuristic is wrong on one task, a middle ground short of--no-route. Forced signals still pull their dependents (a forced+auth-surfaceearns tests and a security review).brigade route --jsonnow carriestriggered_by(the live signal that pulled each stage in) andoverrides, alongside the existing signals/approvals/route/waves/held.- Property tests fuzz the router over 2,000 random catalogs, asserting the invariants a route must always hold (held stages never route, inputs always satisfiable, waves a strict topological order, cycles raise). Plan attempts now record
unknown_covers: acoverstag naming a stage not in the route is surfaced instead of silently ignored. --route-signalrefuses to add or suppress a path signal (code/docs/system): the path is a derive-time decision, and stripping it collapsed the route to a pathless remnant. Found by dogfooding the newlatent-premisesreview skill against the router's own diff.brigade run --wait[=SECONDS]can wait for an active per-target run lock with a finite timeout; lock conflicts remain fail-fast unless the flag is present.- Explicit adapter execution modes keep read-only and writable CLI flags aligned, reject unsupported direct Cursor Composer plan runs, and point those seats to the reviewed ACP transport. Authenticated ACP checks passed
composer-2.5andgrok-4.5; Brigade does not rewrite the direct-onlygrok-4.5-xhighalias because acpx exposes no separate reasoning flag. - The GraphTrail personalized ranking benchmark corpus adds checked synthetic MRR and latency thresholds; the managed snapshot pins that reviewed revision.
- Internal typed run transport and receipt modules separate adapter execution from artifact writing without changing CLI or artifact contracts.
- Capability-aware outcome scoring (Phase 2 of context-aware outcomes):
outcome rankandoutcome explainresolve the current runtime capability and score each artifact's content-current records earned under it, with a thin cohort pulled toward the pooled rate by deterministic shrinkage(helped + kappa*pooled_rate)/(total + kappa)(one documentedkappa=4.0).outcome rank --by-capabilitysorts by "what worked under my current context" (shrunk estimate, then on-capability sample size, then pooled). Records with no capability fingerprint are grandfathered into the current-capability cohort, so a pre-context ledger's default rank/explain output stays byte-identical until signals under a different capability accumulate. The default sort and the promotion ratchet are unchanged; the ratchet still scores the pooled current-fingerprint cohort only. Recency decay and a regression-attribution guard are tracked as follow-ups in docs/design/context-blind-spot.md. - Runtime-context manifest on outcome records (Phase 1 of context-aware outcomes):
outcome captureandoutcome recordnow stamp each new record with a coarsecontextmanifest (Brigade version, interpreter major.minor, platform, plus best-effort harness and model tagged with their source) and acapability_fingerprintover the low-cardinality vector{harness, model_family, python, platform}.outcome explainsurfaces a per-capability breakdown. This is capture-and-surface only: scores, rank order, and the promotion ratchet are unchanged, and records without a manifest are grandfathered. A content hash cannot see the harness a signal was earned under; this is the data foundation for cohort-aware scoring. Design and three-model rationale in docs/design/context-blind-spot.md. - Transitive card-link fingerprints: a card's
content_fingerprintnow folds in the transitive closure of the cards it[[links]](each reachable card's content hash), so editing a linked card invalidates every card that reaches it, the same logic_tracking idea applied to memory cards. The walk is cycle-safe, deterministic, and tolerant of dead links (a[[missing]]link contributes nothing until that card exists). A card with no resolvable links hashes to exactlysha256(card content), byte-identical to the prior scheme, so existing single-card records are never invalidated. - Skill bundle fingerprints (the ledger's
logic_tracking): a skill'scontent_fingerprintnow covers its whole bundle (every file's path plus content hash, skipping.DS_Storeand theskill.jsonsidecar), so editing a bundled helper invalidates the skill's signals the same way editingSKILL.mddoes. A skill whose only content file isSKILL.mdhashes to exactlysha256(SKILL.md), byte-identical to the prior scheme, so existing single-file records are never invalidated. Cards stay single-file; transitive[[wiki-link]]tracking remains future work. - Fingerprint-aware promotion ratchet:
outcome reconcileandoutcome forknow score the current-fingerprint cohort when deciding, so an edited skill must re-earninstall_min_helpedagainst the text that now ships instead of coasting on signals for text that no longer exists. The decision rules (thresholds, cooldown, forward-only ratchet) are unchanged; only the score fed in narrows, and grandfathering keeps a never-edited artifact's decision, receipt, and output byte-identical to before. Decisions that drop proven-stale evidence carry audit fields (content_fingerprint,lifetime_*,stale_records,legacy_records). - Skill content fingerprints in the outcome ledger (CocoIndex's memo-key idea applied to the ratchet):
outcome captureandoutcome recordstamp each new record with the SHA-256 of the artifact's content (content_fingerprint, the installed copy first and the registry master as fallback), andoutcome rank/outcome explaindefault to a score that drops proven-stale records (fingerprinted against a different revision) while keeping lifetime counts visible, so an edited skill earns its score back instead of inheriting signals for text that no longer exists. Pre-fingerprint records are grandfathered into the current score, so a never-edited skill keeps its score andrankoutput stays byte-identical until a skill is actually edited. The ledger remains append-only and the digest chain unchanged; reconcile/promotion rules are untouched. - Deterministic route brief for
brigade run:route_catalog.derive_signalsmaps the task to signals (auth surface, UI, migration, bug, docs, system, ship request), a pure-function router (brigade/router.py, algorithm adapted from alp-river, MIT) composes the required stages into dependency-ordered parallel waves with while/until holds, and the plan prompt carries the result. Plans tag assignments withcovers; a plan missing required stages gets one corrective retry, and gaps are recorded inplan-attempts.jsonrather than failing the run. The route lands inrun.jsontelemetry. Opt out withbrigade run --no-route; release a held ship stage with--approve-ship. brigade route <task> [--template --changed-path --approve-ship --json]shows the composed route for a task: signals, stages with the signal that pulled each in, parallel waves, and held stages. At run time the route also reads git-changed paths (segment-matched) and an optional--route-templatehint.- Registry skills
latent-premises(unguarded-assumption review across input/contract/environment/ordering/cardinality) andretry-safety(side-effect re-run safety, migration deploy-window lens), vendored from skillet. brigade stations verify <path>checks an explicitly selectedstation.jsonwithout running install commands. On POSIX, read-only surfaces and narrowly checked help/version probes run without a shell in an isolated temporary home, with process-group cleanup, finite positive timeouts, strict JSON parsing, and a 64 KiB combined output ceiling. Windows fails closed before probe creation. JSON results omit raw child output, control characters are removed from details, unreadable manifests return structured exit 2, and managed-catalog drift is advisory unless--check-managedis set.
v0.21.1
Added
- Embedded memory-doctor verbs under
brigade memory:status,lint(dead wiki-links),compact(flatten/tighten MEMORY.md), andinit-git. Also available aspython -m brigade.memory_doctor. Handoff promotion stays onbrigade ingest. - Search station CLI:
brigade search status|doctorand review-onlysync planfor GraphTrail plus optional code-search. Shared station health schema (station_health) powers status/doctor/plan across stations. - Tokens station CLI:
brigade tokens status|doctorand review-onlywire planfor Token Glace (current name; TokenJuice is the old name) plus optionalusage-trackermanaged tool under the tokens station. brigade stations discoverfinds localstation.jsoncatalogs (brigade.station.v1) and printsbrigade add <path>next commands.- Plating managed tool under the guard station for optional demo render / leak scan / drift verify helpers.
- Evidence station CLI:
brigade evidence status|doctor, review-onlycrawl planandexport plan, andbrigade add evidencenext-step banner. MiseLedger stays a process-boundary Go binary; crawl/import are operator-run. - Pantry station first-class path:
brigade pantry doctor, status/expiry/setup plans emit explicitnextcommands and product docs links, andbrigade add pantryprints the multi-machine setup sequence. Agent Pantry stays a process-boundary Go sidecar. - MCP adapter for Hermes: user-scoped
~/.hermes/config.yamlundermcp_servers(YAML subset, zero-dep surgical merge). Requires--user-scope. - MCP adapters for Grok CLI: project
.grok/config.toml(grok) and user~/.grok/config.toml(grok-user, requires--user-scope). Same Codex-like[mcp_servers.<name>]TOML shape. (#183)
Fixed
- Release preflight now bumps and synchronizes every version stamp before the full verification and cold-start gates, and the checklist confirms the published PyPI version after tagging.
- MCP Codex/Grok empty-
argsfingerprint conflicts:to_providerno longer emitsargs: [], matching TOML render which omits empty arrays, so force-sync stays idempotent. (#181) - MCP import of url-only servers no longer lands as invalid
stdio+url; coerce tohttp/sse, including OpenClaw sources with a bogustransport: stdioor a command that is actually a URL. (#182) - Productized GraphTrail ↔ Brigade ↔ MiseLedger dogfood path:
brigade operator checkupreports optional loop health (graph/ledger/ last and meanbrief_hit_ratefrom run receipts) without blocking readiness;brigade add graphtrailinstalls the code-graph tool under the search station; QUICKSTART documents install → checkup → export → rank. - Outcome rank/reconcile surfaces mean
context_eval.brief_hit_rateper skill as a quality signal (secondary sort among equal Wilson scores; install/rollback thresholds still use verified exit codes only). brigade-workskill teaches the full loop: verify with capture → outcome from run → MiseLedger export → evidence brief next time.brigade addaccepts a managed tool name (e.g.graphtrail) as well as a station name, so optional loop sidecars install without pulling every tool on the station.
Changed
- Memory station no longer installs the external memory-doctor package. Maintenance is built into brigade-cli; bootstrap-doctor remains an optional sidecar for OpenClaw prefix trim.
v0.21.0
Added
- Verified binary pins for runbooks: a runbook can pin each step's binary by sha256,
runbook planshows pin status without executing,runbook runrefuses on mismatch unless--allow-pin-mismatchis passed and records the check (and any override) in the receipt, andbrigade runbook pin <runbook.json>writes pins from the current binaries. Unpinned runbooks behave exactly as before. (#156) - Tamper-evident receipts: verify, runbook, and outcome artifacts carry a sha256
digestsblock binding the receipt payload and every referenced log file;memory/outcome/records.jsonlis hash-chained withprev_digest/digestso an edited or deleted middle record breaks visibly;brigade receipts verifywalks everything and reports OK, MISMATCH, MISSING, or legacy, wired intobrigade doctorandoutcome doctor. Pre-existing artifacts report as legacy, not failures. (#157) - Optional receipt signing:
brigade receipts keygenwrites a local 0600 key, receipts then carry an HMAC-SHA256signatureandkey_idover the receipt digest, andreceipts verifyreports SIGNED-OK, flips SIGNATURE-MISMATCH nonzero, or marks foreign-key signatures unverifiable. Without a key, receipts are byte-identical to before. (#164) - Code-graph delta receipts: verify runs and
brigade runsnapshot the target's GraphTrail database (WAL-safe), re-sync and diff after, and attach a compactcode_graph_deltasummary to the receipt with the full diff as a digest-coveredgraph-delta.jsonsidecar. Snapshots are deleted after the diff with their sha256s kept as attestations. Fail-open everywhere; read-only, dry-run, and--no-code-graphruns skip capture. (#158) - MiseLedger receipt export:
brigade receipts export miseledgeremits onemiseledger.adapter.v1JSONL line per verify-run and run receipt, withraw.hashreusing the receipt digest so re-imports dedupe on content identity.--new-onlyadds a cursor so reruns export only new receipts, and--importpipes the export straight throughmiseledger import adapter, fail-open when the binary is absent and skipping the import subprocess entirely when nothing new was exported. (#159, #162, #165) - Git provenance on receipts: verify receipts and run payloads capture
head,branch, anddirty_filesinside the digest-covered content, and the MiseLedger export emits the GitHub commit URL as a link. (#162) - Delta-aware outcome ledger:
outcome capturecopies the receipt's compact code-graph delta onto the hash-chained record,outcome capture --run-receipt <run-id|latest>captures from run receipts (where code actually changes) with status mapped to signal, andoutcome rank/reconcilereportgraph: N changing / M no-opper subject. Promotion decision rules are unchanged. (#163, #169) - Context evals on run receipts: when a run had a code-graph brief and its delta captured cleanly,
context_evalrecords whether the pre-run context named the files the run actually touched (hits,missed,brief_hit_rate). Absent when there is no brief, no clean delta, or nothing changed. (#167) - Evidence briefs close the receipts-to-context loop:
brigade runcan attach a capped, fail-open brief of recent verified evidence for the target repo pulled from MiseLedger, always framed as untrusted evidence rather than instructions, gated by--no-evidence, and recorded inrun.json.brigade work import context --from-miseledger '<query>'runs the same fetch on demand. (#171) - Workflow sequence scanner: mines run artifacts for recurring command sequences and proposes runbooks (with empty pin stubs) from what operators actually repeat. (#155)
- Detached run control:
brigade run --detach, live watch, and steer/interrupt verbs for app-server runs. (#154) - Vendored content-guard under
brigade.guard, addedpython -m brigade.guardandbrigade guard ..., and packaged the guard policies so scrub works without a separate checkout. (#166)
Changed
brigade scrubnow defaults to the embeddedbrigade.guardscanner and policy set. SettingCONTENT_GUARD_DIRstill preserves the old external checkout behavior, includingpython -m content_guardwith that checkout'ssrconPYTHONPATH. (#166)
Fixed
- Run receipts are written atomically, closing a race where a concurrent reader could catch a partially written
run.json. (#161) - Friction imports use shared import identity keys, so re-imports no longer duplicate items. (#153)
- The technical guide no longer overstates runbook approval semantics. (#152)
v0.19.0
Added
- Managed station tools now declare machine-readable surfaces: live doctor JSON, bounded markdown briefs, summary JSON, and verify commands where each tool supports them.
brigade stations list --jsonincludes those surfaces so profiles can show what a station can feed into automation before it is installed. brigade add <path>can discover a localstation.jsonmanifest, report its install command and machine surfaces, and refuse to run manifest install commands unless--installis passed. Built-in station installation by station name still works as before.brigade runnow has a run-level brief budget. Code graph and upstream drift impact briefs are ranked by task type, clipped when needed, and recorded inrun.jsonwith attached brief names, sizes, and truncation flags.brigade runcan attach a fail-open upstream drift impact brief when a repo has a GraphTrail database and pending upstream-drift state. The brief combines drift report excerpts withgraphtrail impactoutput so workers see likely blast radius before editing.brigade memory care scannow validatesevidence:frontmatter for receipt-like paths and MiseLedger evidence refs, reportingmissing-evidence-refissues and evidence reference counts.brigade pantry expiry-alertreports near-expiry Agent Pantry sessions and plans anagent-notifymessage by default. It sends only when--sendis passed.brigade outcome rebuild-status,fork, anddiffadd a drift oracle for outcome receipts. Operators can rebuildstatus.jsonfrom records and compare alternate reconciliation configs without mutating live state.
Changed
- The three largest modules are now packages:
repos_cmd(6,417 lines -> 8 files),tools_cmd(6,079 -> 18), andphases_cmd(5,621 -> 7), each behind a facade that preserves the full external surface and monkeypatch semantics; no source file exceeds 2,000 lines. The mypy override list dropped from 45 entries to 21, and type-ratcheting surfaced and fixed four latent defects along the way (a friction-show crash on malformed JSON, a notifications tuple-shape crash, a fleet-health wrong-module read, and a research runner returning an object where callers expected text). - Agent Pantry health now prefers
agentpantry doctor --json, uses the old status JSON as a compatibility fallback, and exposes the newinventory --markdownbrief surface for near-expiry session summaries.