Test Commander is a generic, product-domain-agnostic testing tool. It ships with universal English and software-engineering defaults only — no e-commerce, healthcare, finance, research, or other product-domain vocabulary is baked into the shipped rubric, tags, methodology, fixtures, or examples. This is a deliberate design choice; see Decision D19 in the phased plan.
This guide shows how a consuming project extends Test Commander for its own domain — what to add, where to add it, and what to leave alone.
Test Commander does not know in advance whether you are testing a banking app, a hospital information system, a research data platform, an online retailer, or an internal tool. The shipped rubric, tag taxonomy, and examples make no assumptions. Product-specific knowledge enters at runtime through hooks you control.
A direct consequence: out of the box, Test Commander will not flag PCI-, HIPAA-, or commerce-specific defects in your requirements. It catches universal quality problems (clarity, testability, dependencies, ambiguity, generic-security anti-patterns like plain text). To get domain-aware checks, extend the configuration as described below.
Four explicit hooks. None are required — Test Commander works without any of them on the universal defaults. Use as few or as many as your project needs.
| Hook | What you supply | When it ships |
|---|---|---|
| 1 | <workspace>/config.yaml extensions to rubric keyword sets |
Phase 2 (first surface), extended by later phases |
| 2 | Your project's documents under .test-commander/documents/uploaded/ |
Phase 2 |
| 3 | Phase 3 project knowledge ingestion (/tc:learn-from-*) |
Phase 3 |
| 4 | Project-defined values inside shipped tag namespaces (@area:, @risk:, @persona:) |
Phase 5 |
This guide is updated by every phase that adds an extensible surface. If a phase introduced a new hook or schema key, you will find a worked example here.
The workspace's config.yaml is the primary configuration surface. Test Commander reads it on every helper invocation and unions project-supplied lists with the shipped universal core. Extensions never replace defaults — only add to them.
Phase 2 ships three extensible rubric dimensions under tc-requirements: (read by /tc:review-requirements for data-rules and risk, and by /tc:review-acceptance-criteria for roles-permissions):
tc-requirements:
data-rules:
sensitive-keywords:
- PAN
- primary account number
- PHI
- SSN
risk:
compliance-keywords:
- PAN
- primary account number
- PHI
- social security
roles-permissions:
permission-verbs:
- issue
- refund
- dispense
- prescribe
role-qualifiers:
- customer
- store-manager
- investigator
- reviewerMissing keys = no extension. The helper falls back to the universal core only. The shipped seeded fixture in tests/fixtures/seeded-flawed-requirements/ does not rely on any extension; every seeded defect triggers via the universal core alone.
You are testing an online retail platform. Your requirements regularly mention credit card, PAN, refund, customer, store manager. Extend the configuration:
tc-requirements:
data-rules:
sensitive-keywords: [PAN, primary account number, credit card, credit-card]
risk:
compliance-keywords: [PAN, primary account number, credit card, social security]
roles-permissions:
permission-verbs: [issue, refund, charge, dispute]
role-qualifiers: [customer, store-manager, fulfillment-agent]Effect: /tc:review-requirements now flags requirements that mention refund without naming the role authorized to issue it, and flags requirements that mention credit card storage without a constraint keyword (encrypted, tokenized).
tc-requirements:
data-rules:
sensitive-keywords: [PHI, medical record number, MRN, prescription history, diagnosis code]
risk:
compliance-keywords: [PHI, HIPAA, patient identifier, MRN]
roles-permissions:
permission-verbs: [dispense, prescribe, refer, transfer]
role-qualifiers: [physician, nurse, pharmacist, technician, patient]Effect: requirements mentioning prescribe without a clinician role qualifier are flagged; storage references to PHI without an encryption constraint are flagged for HIPAA risk.
tc-requirements:
data-rules:
sensitive-keywords: [participant identifier, IRB protocol, consent record]
risk:
compliance-keywords: [IRB, deidentification, re-identification risk]
roles-permissions:
permission-verbs: [enroll, terminate, deidentify, archive]
role-qualifiers: [investigator, study-coordinator, IRB-reviewer, participant]Effect: requirements that move participant data without naming the IRB-authorized role are flagged; storage references that lack deidentification constraints are flagged for compliance.
- Extensions are additive. The helper unions defaults with your list. You cannot remove a universal-core keyword via configuration.
- Keyword matching is case-insensitive. Single-token keywords match at word boundaries; multi-token phrases (e.g.
credit card) match literally. - Add the same vocabulary in multiple sections if the same term carries meaning under multiple dimensions (e.g.
PANis bothsensitive-keywordsfor data-rules andcompliance-keywordsfor risk — that is correct and intentional). - Document your additions in your project's
.test-commander/methodology.mdso the team uses a shared taxonomy.
Phase 3 ships four extensible sub-blocks under tc-knowledge: (one per learn helper that has tunable detection; /tc:learn-from-specs has no specs: schema in v1 because the OpenAPI / Postman structural keys are themselves a universal vocabulary):
tc-knowledge:
documents:
# Additive case-sensitive entity names. Surface anywhere they appear in
# prose, even outside an entity-tokened heading.
entity-keywords: [Patient, Provider, Claim]
# Additive case-insensitive heading tokens. A heading containing any of
# these substrings is treated as a journey heading.
journey-headings: [story, flow]
code:
# Where to walk. Default: documents/uploaded/code. Resolves relative to
# <workspace>. Use ../src to point at the consuming project's actual
# source if it lives outside the workspace.
source-root: src
# v1 parses python only. Setting [] short-circuits the AST walk while
# still flagging unsupported-language files.
enabled-languages: [python]
# Substring matches against the relative path; default includes
# __pycache__, .git, .venv, node_modules.
ignored-paths: [migrations, fixtures, .venv]
# Reserved for the v2 spec-vs-decorator cross-check; v1 ignores.
endpoint-decorator-patterns: ["@app.{method}", "@router.{method}"]
api:
# recorded (default) or live. Live mode is refused under pytest via the
# PYTEST_CURRENT_TEST env var.
mode: recorded
# Default: documents/uploaded/recorded-api/responses.json. Resolves
# relative to <workspace>.
recorded-path: documents/uploaded/recorded-api/responses.json
# Live-mode-only (v2). v1 ignores both keys when mode is recorded.
base-url: http://localhost:8000
auth-header: "Authorization: Bearer ${TC_API_TOKEN}"
tests:
# Default: documents/uploaded/tests. Resolves relative to <workspace>.
source-root: tests
# Substring matches against the relative path.
ignored-paths: [fixtures, __pycache__]Missing keys = no extension. The helpers fall back to documented defaults. The shipped seeded fixture in tests/fixtures/seeded-sample-project/ does not rely on any extension; every seeded finding triggers via the universal cores plus the default paths.
The consuming project's source lives at the repository root under src/; the OpenAPI spec lives at openapi.yaml; tests live under tests/.
tc-knowledge:
documents:
entity-keywords: [Patient, Provider, Claim, EncounterRecord]
journey-headings: [story]
code:
source-root: ../src
enabled-languages: [python]
ignored-paths: [migrations, alembic, .venv, __pycache__]
endpoint-decorator-patterns: ["@app.{method}", "@router.{method}"]
api:
mode: recorded
recorded-path: ../tests/recorded-api/responses.json
tests:
source-root: ../tests
ignored-paths: [conftest.py, fixtures]Effect: /tc:learn-from-code walks the consuming project's src/ (not the workspace's local code directory); /tc:learn-from-tests reads the consuming project's tests/; /tc:learn-from-api plays back a recording the project's own test harness produced. Patient, Provider, Claim, EncounterRecord surface as entities anywhere they appear in prose.
The consuming project is a Node/Express service. v1 of /tc:learn-from-code parses Python only — but enabled-languages: [] documents the project's intent and still flags every .ts/.js file as language-unsupported-in-v1 so the gap signal is visible while a future phase adds the TS/JS parser.
tc-knowledge:
documents:
entity-keywords: [Order, LineItem, Shipment, Refund]
code:
source-root: ../src
# v1 cannot parse JS/TS. The empty list documents intent; the helper
# still emits language-unsupported-in-v1 gaps for every detected file,
# which keeps the surface visible until the parser ships.
enabled-languages: []
ignored-paths: [node_modules, dist, build]
api:
mode: recorded
recorded-path: ../e2e/recorded.json
tests:
# Jest convention: __tests__/*.test.js or src/**/*.test.js. v1 detects
# *.spec.ts and *_test.py but does NOT recognize .test.js yet; the
# ignored-paths key is unused here, but documents the future intent.
source-root: ../e2e
ignored-paths: [node_modules, fixtures]Effect: /tc:learn-from-code flags every .js/.ts file under ../src/ as a language-unsupported-in-v1 gap; the consuming project knows what is uncovered in v1. The same code helper still runs cleanly for any Python files that happen to be in the tree (utility scripts, automation, etc.).
The consuming project exposes its API through Postman collections rather than OpenAPI. /tc:learn-from-specs auto-detects Postman v2.1 by file extension; no specs config key is needed.
tc-knowledge:
documents:
entity-keywords: [Subscription, Plan, Invoice]
api:
mode: recorded
# Postman collection runs typically write a separate recording per
# request. The consuming project concatenates them into a single
# responses.json at this path.
recorded-path: ../qa-artifacts/postman-recordings.json
tests:
source-root: ../qa
ignored-paths: [fixtures]The consuming project drops their-api.postman_collection.json under documents/uploaded/ and runs /tc:learn-from-specs. The helper extracts endpoints (from item.request), schemas (from request.body.raw JSON top-level keys), and auth schemes (from distinct request.auth.type values across requests). The Postman URL is normalized — {{base_url}} prefixes are stripped via _strip_postman_variables so the captured path is just the API path. /tc:learn-from-api's cross-check against the spec model works the same way it does for OpenAPI consumers.
- Universal cores.
/tc:learn-from-docsdetects entities, terms, journeys, business rules, and assumptions through universal English heading tokens (entit/model/noun/glossaryfor entities;glossary/terminologyfor terms;journey/flow/walkthrough/scenariofor journeys) plus RFC-2119 modals (must/shall/should/may) and assumption markers (assume/expected/presumed/likely)./tc:learn-from-specsreads OpenAPI 3 / Postman v2.1 structural keys directly./tc:learn-from-codeuses stdlibastfor Python; non-Python extensions (.ts,.tsx,.js,.jsx,.go,.java,.rb) are flagged./tc:learn-from-apireads recorded{method, path, status, headers, body}entries and classifies status family./tc:learn-from-testsparsestest_*.py/*_test.pyviaastand counts*.spec.tsfiles via regex without parsing TypeScript. - Schema keys (per the YAML block above):
tc-knowledge.documents.{entity-keywords, journey-headings},tc-knowledge.code.{source-root, enabled-languages, ignored-paths, endpoint-decorator-patterns},tc-knowledge.api.{mode, recorded-path, base-url, auth-header},tc-knowledge.tests.{source-root, ignored-paths}. - Tests that would fail if helpers ignored extensions.
tests/test_learn_from_docs.py::test_entity_keywords_extension_appliedand::test_journey_headings_extension_appliedassert thattc-knowledge.documents.{entity-keywords, journey-headings}extend the universal cores additively.tests/test_learn_from_code.py::test_source_root_extension_appliedand::test_ignored_paths_extension_excludes_matchesassert the same for code.tests/test_learn_from_api.py::test_recorded_path_extension_appliedand::test_live_mode_refused_under_pytestassert recorded-path + live-mode refusal.tests/test_learn_from_tests.py::test_source_root_extension_appliedasserts the tests source-root extension. Each test points the helper at a non-default path and asserts the helper found the content there; if the helpers ignored extensions, every one of these tests would fail with "no source found" or equivalent.
Phase 4 ships three extensible sub-blocks under tc-explore: (read by /tc:create-charter for charter risk and area heuristics; by /tc:explore for exploration mode and recorded-session path; by the internal exploration-review sub-mode for review rubric extensions). The v1 /tc:test-ideas helper has no test-ideas: schema — the universal-English stopword list and five-character-stem matching are non-tunable in v1; project-specific keyword tuning is a deferred v2 surface.
tc-explore:
charters:
# Additive case-insensitive risk-vocabulary tokens. Lines in
# <workspace>/risk-register/risk-register.md matching any universal-core
# OR extended keyword surface as risk-areas in the generated charter.
risk-keywords: [PCI, PHI, GDPR]
# Additive case-insensitive area tokens used by /tc:explore (Step 4.3)
# for journey detection. Universal core covers sign-in, sign-out,
# dashboard, search, upload, profile, settings, session, workspace.
area-keywords: [checkout, refund, claim, prescription]
exploration:
# recorded (default) or live. Live mode is refused under pytest via the
# PYTEST_CURRENT_TEST env var. Live mode v1 is documented but not yet
# implemented; recorded playback is sufficient for every Phase-4 contract.
mode: recorded
# Default: documents/uploaded/recorded-sessions. Resolves relative to
# <workspace>. Use a different path when the project's CI captures
# recordings under a sibling artifact directory.
recorded-path: documents/uploaded/recorded-sessions
# Live-mode-only (v2). v1 ignores both keys when mode is recorded.
mcp-endpoint: http://localhost:9999
target-url: http://localhost:8000
review:
# Reserved for v2 review-rubric extensions. v1 ships only the two
# universal-core checks: missing-evidence (asymmetric +/-3 second window)
# and charter-coverage-shortfall (every AC marked unobserved).
rubric-extensions: []Missing keys = no extension. The helpers fall back to the documented defaults. The shipped seeded fixture in tests/fixtures/seeded-exploration-session/ does not rely on any extension; every seeded anomaly, every coverage verdict, and every candidate scenario triggers via the universal cores plus the default paths.
The consuming project is a Python web app with a Playwright + pytest test suite. Recordings are captured by the project's CI to a sibling artifact directory; the project supplies its own PCI-risk vocabulary for the charter generator to surface.
tc-explore:
charters:
risk-keywords: [PCI, primary account number, fraud]
area-keywords: [checkout, refund, payment-method]
exploration:
mode: recorded
recorded-path: ../ci-artifacts/recorded-sessionsEffect: /tc:create-charter surfaces any risk-register line mentioning PCI / primary account number / fraud as a charter risk-area. /tc:explore reads recordings from ../ci-artifacts/recorded-sessions/<CH-ID>.json rather than the workspace-default path. /tc:test-ideas runs against the resulting session summaries with no per-project tuning — the universal-stem matching catches payment / payments, auth / authenticate / authentication, and so on.
The consuming project is a React Native iOS app. The team uses a custom Appium-driving MCP instead of Playwright. v1 of /tc:explore reads any recorded-session JSON whose shape matches the universal core ({timestamp, event_type, page_url, ...} per event); the team's MCP records to that shape.
tc-explore:
charters:
risk-keywords: [keychain, biometric, background-state]
area-keywords: [onboarding, push-notification, deep-link]
exploration:
mode: recorded
recorded-path: ../e2e/appium-recordings
# mcp-endpoint and target-url are live-mode-only (v2); recorded mode
# ignores both even when present.
mcp-endpoint: http://localhost:4723
target-url: ios://com.example.app
review:
rubric-extensions: []Effect: the team's existing Appium recordings drive /tc:explore. The event_type: page_load semantics map to "screen" rather than "URL", but the helper does not care about the value semantics — it only counts events by type. Anomalies, screenshots, and coverage verdicts work the same. The mcp-endpoint and target-url keys are accepted as documentation of future v2 wiring; v1 ignores them under mode: recorded.
The consuming project is a public REST API with no UI. Exploration is driven from the OpenAPI spec rather than a browser-recorded session. The team writes a thin script that exercises every spec-derived endpoint and records each {method, path, status, headers, body} to the universal-core JSON shape.
tc-explore:
charters:
risk-keywords: [rate-limit, payload-size, idempotency-key]
area-keywords: [authentication, pagination, webhook]
exploration:
mode: recorded
recorded-path: ../qa/spec-derived-recordingsEffect: charters scope exploration around spec-derived endpoint families. /tc:explore reads the spec-driven recordings exactly as it would Playwright MCP recordings — the universal-core event types (page_load, click, fill, screenshot, console_message, network_request, anomaly) cover the API-only case naturally because API exercise records network_request and anomaly rows; the helper does not require the other event types to be populated. Coverage verdicts use the standard URL-and-keyword-match heuristic — for API-only projects URL paths dominate the heuristic, which is the desired behavior. /tc:test-ideas enriches Phase-2 REQ seeds via the same stem-matching as web-app projects.
- Universal cores.
/tc:create-charterships universal risk-vocabulary (security,auth,performance,data-integrity,accessibility,compliance,session,permission,token,leak) and universal area-vocabulary (sign-in,sign-out,dashboard,search,upload,profile,settings,session,workspace); both extend additively./tc:exploreships six universal anomaly categories (slow-response,console-error,broken-link,missing-evidence,auth-mismatch,unexpected-state) with severities (low,medium,high,critical) and ten universal trigger words for coverage downgrade (expiration,expired,expire,leak,leakage,concurrent,race,timeout,timed-out,rollback). The internal exploration-review sub-mode ships two universal checks (missing-evidencewith the asymmetric ±3-second rule;charter-coverage-shortfallfor every AC markedunobserved)./tc:session-summaryships three universal candidate-scenario types (happy,edge,negative)./tc:test-ideasships a universal-English stopword list and a five-character stem-match algorithm for charter-coverage → REQ-ID cross-reference. - Schema keys (per the YAML block above):
tc-explore.charters.{risk-keywords, area-keywords},tc-explore.exploration.{mode, recorded-path, mcp-endpoint, target-url},tc-explore.review.{rubric-extensions}. Notc-explore.test-ideas:schema in v1 — the stopword list and stem length are deliberately non-tunable in v1; the deferred v2 surface would exposetc-explore.test-ideas.{stopwords-extend, stem-length}. - Tests that would fail if helpers ignored extensions.
tests/test_create_charter.py::test_risk_keywords_extension_appliedasserts thattc-explore.charters.risk-keywordssurfaces a project-specific PCI risk-register line in the generated charter;::test_area_keywords_extension_acceptedasserts the area extension is accepted alongside--mission.tests/test_explore.py::test_recorded_path_extension_appliedassertstc-explore.exploration.recorded-pathpoints the helper at a custom location;::test_live_mode_refused_under_pytestasserts the live-mode refusal under the test harness. Each test points the helper at non-default keywords or a non-default path and asserts the helper found or surfaced the result; if the helpers ignored extensions, every one of these tests would fail with "keyword not surfaced" or "recording not found" or equivalent.
Phase 5 ships the tc-bdd skill (/tc:generate-bdd, /tc:review-bdd) and
tc-traceability (/tc:traceability-map). The first configurable surface,
shipped with /tc:generate-bdd in Step 5.2, lets a project union additional
universal class tags onto every generated scenario:
tc-bdd:
tags:
# Class tags unioned onto every generated scenario, in addition to the
# type-derived class (@smoke for happy/positive, @regression for
# edge/negative). A leading @ is optional; it is added if missing.
extra-classes: ["@automated-candidate"]Worked example — a project that wants every generated scenario marked as an
automation candidate. With the block above, a scenario the helper would
otherwise tag @area:checkout @req:REQ-014 @cs:CS-014-002 @regression becomes
@area:checkout @req:REQ-014 @cs:CS-014-002 @regression @automated-candidate,
so a downstream /tc:automation-plan (Phase 6) can select the whole set by tag.
The second surface, shipped with /tc:review-bdd in Step 5.3, lets a project
add vague-word and UI-word tokens to the review rubric:
tc-bdd:
review:
rubric-extensions:
vague-words: ["tbd", "wip"] # flagged as ambiguous-step
ui-words: ["tap", "swipe"] # flagged as ui-coupled-stepWorked example — a mobile project. A mobile team whose specs say "tap" and
"swipe" adds those to ui-words so the ui-coupled-step check catches
mobile-gesture leakage into behavior specs; adding tbd/wip to vague-words
flags placeholder steps left in during drafting. The extensions union with the
universal cores additively.
The universal class core (@smoke, @regression, @manual, @exploratory,
@automated-candidate) and the machine-readable linkage tags (@req:, @cs:,
@anomaly:) are not project-tunable — they are the traceability contract. The
project namespace values (@area:, @risk:, @persona:) are picked per
project; see "Hook 4: project-defined tag namespaces" below.
The two surfaces above tune differently depending on how a project's BDD is shaped. Three materially-different shapes:
-
Web app with browser-driven scenarios. Scenarios describe user-visible behavior across pages. The default
ui-wordsrubric (click,button,url,selector, …) is exactly right — no extension needed; it keeps selectors and URLs out of the behavior specs. The project tunestc-bdd.tags.extra-classes: ["@automated-candidate"]so generated scenarios flow into the Phase-6 automation plan, and picks@area:values per page area (@area:sign-in,@area:dashboard).tc-bdd: tags: extra-classes: ["@automated-candidate"]
-
API-only project with spec-derived scenarios. Scenarios describe request/response behavior, so "URL" is legitimately part of the language — the project does not add URL terms to
ui-words(that would mis-flag endpoint references). Instead it adds API-draft placeholders tovague-wordsso unfinished steps are caught, and tags risk by data class.tc-bdd: review: rubric-extensions: vague-words: ["tbd", "returns stuff", "some payload"]
-
Mobile app with screen-flow scenarios. Gesture verbs leak into specs, so the project extends
ui-wordswith mobile gestures and marks exploratory device-specific scenarios@manualduring refinement.tc-bdd: review: rubric-extensions: ui-words: ["tap", "swipe", "pinch", "long-press"]
Each shape exercises a different surface as its primary tuning point: the web
shape tunes tags.extra-classes, the API shape tunes vague-words, the mobile
shape tunes ui-words.
- Universal cores.
/tc:generate-bddships a type→class map (happy/positive→@smoke;edge/negative→@regression) and the@req:/@cs:/@anomaly:linkage-tag convention./tc:review-bddships the six-category rubric (ambiguous-step,missing-tag,untraceable,ui-coupled-step,missing-examples,conjunction-overload) with universal vague-word and UI-word sets./tc:traceability-mapships the shared requirements-map renderer and the scenario-level test-map withpendingdownstream links. - Schema keys.
tc-bdd.tags.extra-classes(list) andtc-bdd.review.rubric-extensions.{vague-words, ui-words}(lists). The linkage tags and the type→class map are not tunable — they are the traceability contract./tc:traceability-mapships noconfig.yamlsurface (the maps are mechanical). - Tests that would fail if helpers ignored extensions.
tests/test_generate_bdd.py::test_config_extra_classes_unionasserts a configured@automated-candidateappears on every generated scenario;tests/test_review_bdd.pydrives the rubric against the seededflawed.featureand would fail if a category were dropped. If the helpers ignored theconfig.yamlextensions, the union assertion would fail with "tag not present".
Phase 6 ships the automation skills. The first configurable surface, shipped
with /tc:automation-plan in Step 6.3, lets a project re-weight the seven-factor
automation-suitability rubric:
tc-automate:
suitability:
weights:
# Only these seven factor names are recognized; an omitted or unknown
# name keeps its default weight. Defaults: traceable 3, regression-value 2,
# risk-flagged 2, deterministic 2, right-sized 1, data-ready 1,
# persona-scoped 1.
risk-flagged: 4Worked example — a project that automates by risk first. The default rubric
weights traceable highest (3). A team whose priority is "automate the riskiest
behavior before anything else" raises risk-flagged to 4, so a @risk:high
scenario that scores consider under the defaults is promoted to automate.
Because the factor names are fixed and unknown names are ignored, a typo silently
keeps the default rather than corrupting the score — check the generated
automation-plan/<area>.md table to confirm the new weighting took effect.
The seven factor names, the rank thresholds (>= 8 automate, >= 5
consider), and the two hard overrides (@automated-candidate always automate,
@manual always manual) are not tunable — they are the rubric contract. Only the
per-factor weights are project-tunable.
tc-automate.suitability.weights is the only config.yaml surface Phase 6
adds. The other two Phase 6 inputs are not config.yaml keys:
PLAYWRIGHT_BASE_URL(environment variable) — the target the generatedplaywright.config.tspoints at. It is a runtime value, not a workspace setting, so it lives in the environment, never inconfig.yaml.- The
test-data/seed shape — universal and fixed; a project fills in real field values by hand-editing the generatedseed/<area>.json(preserved if marker-less) or adding a Python factory undertest-data/factories/, not viaconfig.yaml.
The suitability weights tune differently depending on what a project optimizes for. Three materially-different shapes:
-
Web app prioritizing smoke coverage. A team that wants its critical-path smoke scenarios automated first raises
regression-valueso@smoke-tagged scenarios outrank everything else.tc-automate: suitability: weights: regression-value: 4
-
Risk-driven / compliance project. A team that automates by risk class first raises
risk-flagged, so a@risk:highscenario is promoted ahead of lower-risk but otherwise-strong candidates.tc-automate: suitability: weights: risk-flagged: 4
-
Large suite favoring stable, data-driven specs. A team drowning in flaky candidates de-emphasizes traceability (everything is traceable anyway) and rewards stability, so
right-sizedanddata-readyscenarios rank higher.tc-automate: suitability: weights: traceable: 1 right-sized: 3 data-ready: 3
Each shape changes which scenarios cross the >= 8 automate threshold; the
threshold and factor names themselves stay fixed. Confirm the effect in the
generated automation-plan/<area>.md table.
- Universal cores.
/tc:build-frameworkships a universal Playwright config (target viaPLAYWRIGHT_BASE_URL) and four.tsobject templates./tc:automation-planships the seven-factor rubric with fixed names, thresholds, and hard overrides./tc:automateships the provenance + fixture conventions;/tc:review-automationships the seven-category review rubric;/tc:generate-test-dataships the universal seed shape. - Schema keys.
tc-automate.suitability.weights(mapping of the seven factor names to integer weights) — the onlyconfig.yamlsurface. Unknown names are ignored (a typo keeps the default, never corrupts the score). - Not tunable. The rubric factor names, thresholds, and overrides; the six review categories; the framework layout; the seed shape. The target URL is an env var, not a config key.
- Tests that would fail if the helper ignored the weights.
tests/test_automation_plan.py::test_config_weights_change_rankingwrites a borderline non-candidate scenario and asserts boostingtraceableflips it fromconsidertoautomate; if the helper ignoredtc-automate.suitability.weights, the rank would not change and the assertion would fail.
Phase 7 ships execution, evidence, and reporting. Its one config.yaml surface
is the quality-gate thresholds — the bar a release must clear. The run modes,
the evidence policy, the triage rubric, the report section catalog, and the gate
criteria are universal; only the gate thresholds are project-tuned, under
tc-quality-report.gate.thresholds:
| Key | Default | Meaning | Breach |
|---|---|---|---|
min-pass-rate |
1.0 |
minimum passed / total of the latest run |
FAIL |
max-failed |
0 |
maximum failing tests | FAIL |
max-flaky |
0 |
maximum flaky (pass-on-retry) tests | WARN |
max-open-questions |
0 |
maximum unresolved open questions | WARN |
Any unset key keeps its default; unknown keys and unparseable values are ignored
(the gate falls back to the default for that criterion). The verdict is the worst
per-criterion verdict (PASS < WARN < FAIL); the CLI exits 1 on FAIL so CI can
branch on it. Tune by project maturity and risk tolerance — vary the bar, not
the universal criteria.
A safety- or compliance-critical project ships nothing with a known failure and treats flakiness as a defect:
tc-quality-report:
gate:
thresholds:
min-pass-rate: 1.0 # 100% of the suite must pass
max-failed: 0 # no failing tests
max-flaky: 0 # no flaky tests (a flake is a defect)
max-open-questions: 0 # every open question resolved before releaseA fast-moving product tolerates a small transient-failure margin and a few known-flaky tests, but still blocks on hard failures:
tc-quality-report:
gate:
thresholds:
min-pass-rate: 0.95 # allow a 5% transient margin
max-failed: 0 # but never ship a hard failure
max-flaky: 5 # tolerate a few known flakies (WARN above)
max-open-questions: 25 # WARN above 25 unresolved questionsAn internal dashboard wants the gate as a signal, not a blocker, so the soft criteria warn generously while the hard floor stays meaningful:
tc-quality-report:
gate:
thresholds:
min-pass-rate: 0.80 # a low but non-trivial floor
max-failed: 2 # a couple of known failures are acceptable
max-flaky: 50 # effectively advisory
max-open-questions: 100 # effectively advisory- Universal cores.
/tc:runships the run modes and the result schema;/tc:analyze-resultsships the four-category triage rubric; thetc-evidenceindexer ships the commit-versus-ignore evidence policy;/tc:reportships the fifteen-section catalog and the[fact]/[interpretation]/[review]separation;/tc:quality-gateships the four gate criteria. - Schema keys.
tc-quality-report.gate.thresholds(min-pass-rate,max-failed,max-flaky,max-open-questions) — the onlyconfig.yamlsurface. Unknown keys are ignored (a typo keeps the default). - Not tunable. The run modes, the result schema, the triage categories, the
evidence policy, the report section catalog, and the gate criteria. The
Playwright target is the
PLAYWRIGHT_BASE_URLenv var, not a config key. - Tests that would fail if the helper ignored the thresholds.
tests/test_quality_gate.py::test_loose_thresholds_passand::test_flaky_only_warnswrite atc-quality-report.gate.thresholdsblock and assert the verdict flips (FAIL → PASS, and FAIL → WARN); if the helper ignored the config, the verdict would stay FAIL and both assertions would fail.
Phase 8 ships the governed learning loop (tc-learning: /tc:learn, the three
/tc:learn-from-*, /tc:review-lessons, /tc:promote-lessons). It adds no new
config.yaml surface — the tc-lesson/v1 taxonomy, the review rubric, and the
promotion gate are the universal governance contract (Decision D19). The one
place a project tunes the loop is where the lessons come from: the artifacts
under documents/uploaded/feedback.md and the resolved requirements/open-questions.md
entries the capture commands read. Promotion is governed by a fixed --apply
human gate, not a config key, and the loop writes only under learning/ — never
the shipped methodology or any third-party skill (Open Question Q6), so there is
nothing to misconfigure. Project-specific guidance enters through
learning/promoted-guidance.md (what you promote), not through a schema.
Phase 9 ships visual documentation (tc-visualize: /tc:visualize, the eight
/tc:diagram-* commands, /tc:generate-infographic, /tc:render-visuals). It
adds no new config.yaml surface — the diagrams are generated mechanically
from committed artifacts, with no theme, layout, or keyword knobs in v1 (Decision
D19 keeps the shipped output universal). A project tunes the visuals only by
changing the underlying artifacts the generators read (its requirements,
risk register, traceability maps, system model, automation plan, and quality
report); there is nothing to misconfigure. Rendering uses the Mermaid CLI
provisioned by make install; a project that wants a custom Mermaid theme can
post-process the visuals/mermaid/*.md sources or pass its own mmdc config,
but that is outside Test Commander's schema.
Phase 10 ships the read-only web console (tc-web: /tc:web-init,
/tc:web-start, /tc:web-sync, /tc:web-index-artifacts, /tc:web-export). It
adds no new <workspace>/config.yaml surface — the console renders the same
committed artifacts every other phase produces, and its only configuration
(.test-commander/.web/console.json: the API base and web port, written by
/tc:web-init) is the console's own derived config, not a project-domain
extension point. A project tunes what the console shows only by changing the
underlying artifacts it reads. The stack ports and the served workspace are set
through TC_WORKSPACE and the compose file, not a schema.
Phase 10.5 ships the genuinely-extensible governance surface: which roles may do
what, and which actions require human approval. Both live under
<workspace>/policy/ and are read by the runtime; the shipped defaults are the
universal baseline (Decision D19), overridable per deployment.
policy/permissions.yaml — role → allowed permission levels (default deny):
# Grant your "QA Lead" role through execute-tests, but not destructive/admin.
Viewer:
- read-only
QA Lead:
- read-only
- safe-write
- code-write
- execute-tests
Admin:
- read-only
- safe-write
- code-write
- execute-tests
- external-network
- destructive
- adminA role absent from the file, or a level absent from a role's list, is denied — there is no implicit grant. The five default roles (Viewer → Tester → Automation Engineer → Maintainer → Admin) apply when the file is absent.
policy/approvals.yaml — which levels require an approval card:
# A stricter deployment: require approval even for safe-write.
require_approval:
- safe-write
- code-write
- execute-tests
- external-network
- destructive
- adminWhen the file is absent, the five privileged levels (code-write and above)
require approval and safe-write does not. The seven permission levels and the
action classifier themselves are the fixed governance contract — a project tunes
who and what-needs-approval, not the levels. See
governance.md and
../security-and-permissions.md.
Phase 11 ships the Runtime API and the
MCP server — two new front-ends to the same governance
pipeline. They add no new <workspace>/config.yaml, tag, or policy surface:
both are governed by the existing policy/permissions.yaml and
policy/approvals.yaml above, so customizing those two files customizes the API
and the MCP server at the same time. The only Phase-11 configuration is
deployment-level, not a project-domain extension point: TC_WORKSPACE (the
served project root) and the operator's choice of governance adapter (the
deterministic mock by default; the real Claude adapter is an explicit server-side
choice). See integrating.md.
Phase 12 ships a genuinely-extensible surface: the sandbox config that
/tc:sandbox-init writes and the safety guards read. A consuming project edits
it to point the sandbox at their application and lock down targeting:
# .test-commander/sandbox/config.yaml
schema: tc-sandbox/v1
provider: docker-compose # or a configured container host
environment_label: "Acme QA Sandbox (ephemeral)"
target:
base_url: "https://qa.acme.example"
allowed_domains: # only these hosts may be targeted
- acme.example
- "*.acme.example"
block_private_ranges: true # refuse 10/8, 172.16/12, 192.168/16, 127/8, 169.254/16
approvals_required: # levels that need an approval before the sandbox acts
- external-network
- destructiveA target whose host is not on allowed_domains is refused, and a private/
loopback/link-local address is refused while block_private_ranges is true — so
a sandbox can never be aimed at an internal network. The seven permission levels
and the approval semantics are the fixed governance contract (the sandbox runs
the same Phase-10.5 pipeline); a project tunes the target and the allow-list,
not the levels. See sandbox.md.
Phase 13 ships the autonomy-mode surface: how much the continuous quality agent may do on its own. A consuming project sets the mode and the PR label:
# .test-commander/continuous/config.yaml
schema: tc-continuous/v1
autonomy_mode: 0 # 0 advisor, 1 assisted, 2 approved-execution,
# 3 PR-automation, 4 governed-autonomy
pr_label: "acme/auto-tests"The mode is a ceiling on what auto-approves in the Phase-10.5 pipeline
(cumulative: advisor → safe-write → execute-tests → code-write →
external-network); destructive and admin never auto-approve at any mode, and
only modes 3–4 may open pull requests. A team starts at mode 0 (advice only) and
raises it deliberately. The autonomy progression and the approval semantics are
the fixed governance contract; a project tunes the mode and the PR label, not
what each level means. See continuous-quality.md and
../autonomy-levels.md.
In addition, continuous mode reads two workspace artifacts your earlier phases
produce — the impact map (product-knowledge/impact-map.yaml: file-path pattern
→ impacted features) and the coverage map (traceability/coverage.yaml: feature
→ automated?). These are your project's knowledge; the shipped fixtures are
only illustrative.
The Phase 2 helpers read every Markdown file in .test-commander/documents/uploaded/ that matches their convention — REQ-\d+ markers for requirements, US-\d+ for stories, AC-\d+ for acceptance criteria. Drop your real product requirements there as Markdown files. No tool configuration is needed; the helpers find and parse them.
This is the single most important customization: the requirements Test Commander reviews are your requirements. The shipped fixture exists only so Test Commander can be tested against itself.
Phase 3 ships the tc-knowledge skill with five commands plus a shared synthesizer: /tc:learn-from-docs, /tc:learn-from-specs, /tc:learn-from-code, /tc:learn-from-api, /tc:learn-from-tests. They scan your narrative documents, OpenAPI/Postman specs, Python source, recorded API responses, and existing tests to build a structured knowledge model under .test-commander/product-knowledge/ (10 artifacts: 5 per-source models, 4 cross-cutting indexes, plus system-model.md regenerated by every run). Downstream Phase 4, 5, 6, and 7 commands read this knowledge to make their findings product-aware without you hand-curating keyword lists. End-to-end walkthrough: building-project-knowledge.md.
The schema for tuning the five helpers is the tc-knowledge: block documented above under "Phase 3 schema (tc-knowledge)". Worked extension examples cover Python/FastAPI, Node/Express (where v1 cannot parse JS yet), and Postman-only projects.
This is the preferred long-term path for domain awareness. config.yaml extensions to tc-requirements are a useful Phase 2 bootstrap; Phase 3 ingestion is how Test Commander learns your product without manual taxonomy work.
Test Commander (Phase 5+) ships three project-defined namespaces:
| Namespace | Purpose | Example values your project picks |
|---|---|---|
@area:<feature> |
Feature area your project tests | @area:sign-in, @area:reports, @area:billing |
@risk:<class> |
Risk classification | severity: @risk:high/medium/low; category: @risk:data-loss/availability/integrity; domain: @risk:compliance/fraud/safety |
@persona:<role> |
User persona | @persona:admin, @persona:customer, @persona:investigator |
Test Commander ships the namespaces; you pick the values. Document your values in your project's .test-commander/methodology.md. Tag-driven gates (e.g. "block release if any @risk:high test is failing") are configured per-project.
- Do not fork Test Commander. The universal core is a contract — extending via
config.yamlis the supported path. Forking divorces you from upstream improvements. - Do not put product-specific vocabulary in the shipped fixtures (
tests/fixtures/). Those test Test Commander itself; your domain belongs in your project'sdocuments/uploaded/andconfig.yaml. - Do not expect overrides to remove defaults. Extensions add to defaults, never replace them. If a universal keyword (e.g.
plain textfor risk) is not relevant to your product, your finding still includes it — you simply ignore the suggestion in your team's review. - Do not encode product knowledge inline in helpers or methodology docs. Phase 3's knowledge ingestion is the canonical place for that information;
config.yamlis the canonical place for vocabulary extensions.
- Workspace reference — where each artifact lives, including
config.yaml. - Phased plan, Decision D19 — the genericness principle.
- Workflow walkthrough — first end-to-end run.
- Command reference — every command, by phase.
- Getting started — install and verify.