Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
90 changes: 89 additions & 1 deletion .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,9 +8,14 @@ on:
workflow_dispatch:
inputs:
release_tag:
description: "Existing tag to repair as a GitHub Release"
description: "Existing tag to repair on PyPI and GitHub"
required: false
type: string
waive_v176_qualification:
description: "Owner-authorized one-time qualification waiver; valid only for v1.7.6"
required: false
type: boolean
default: false

permissions:
contents: read
Expand Down Expand Up @@ -959,6 +964,23 @@ jobs:
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: "3.11"
- name: Enforce and record the v1.7.6-only qualification waiver
if: inputs.waive_v176_qualification
env:
RELEASE_TAG: ${{ inputs.release_tag }}
GH_ACTOR: ${{ github.actor }}
GH_RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
shell: bash
run: |
set -euo pipefail
test "$RELEASE_TAG" = "v1.7.6"
Comment thread
Coding-Dev-Tools marked this conversation as resolved.
{
printf '# Release qualification waiver\n\n'
printf 'Tag: `%s`\n' "$RELEASE_TAG"
printf 'Triggered by: `%s`\n' "$GH_ACTOR"
printf 'Workflow run: %s\n\n' "$GH_RUN_URL"
printf 'The owner explicitly waived full-product qualification for this repair. No unverified release gate is represented as passed.\n'
} >> "$GITHUB_STEP_SUMMARY"
- name: Download published distributions
env:
GH_TOKEN: ${{ github.token }}
Expand Down Expand Up @@ -1103,6 +1125,7 @@ jobs:
cp dist/*.whl dist/*.tar.gz verified-dist/

- name: Require signed full-product qualification before PyPI repair
if: ${{ !inputs.waive_v176_qualification }}
Comment thread
Coding-Dev-Tools marked this conversation as resolved.
env:
RELEASE_TAG: ${{ inputs.release_tag }}
ENGRAPHIS_RELEASE_QUALIFICATION: ${{ secrets.ENGRAPHIS_RELEASE_QUALIFICATION }}
Expand All @@ -1116,6 +1139,43 @@ jobs:
python -m scripts.verify_release_qualification --dist dist \
--commit "$ENGRAPHIS_REPAIR_COMMIT" --tag "$RELEASE_TAG"

- name: Disclose the qualification waiver before PyPI repair
if: inputs.waive_v176_qualification
env:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
RELEASE_TAG: ${{ inputs.release_tag }}
GH_RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
shell: bash
run: |
set -euo pipefail
test "$RELEASE_TAG" = "v1.7.6"
# The exception covers this retained candidate, not future reuse of its tag.
test "$ENGRAPHIS_REPAIR_COMMIT" = "6a441a75c8dd159607fa3933da83f600864b9146"
{
printf '## Release qualification\n\n'
printf 'The owner waived full-product qualification for this release. Mandatory full-product gates are not represented as passed.\n\n'
printf 'Source commit: `%s`\n\n' "$ENGRAPHIS_REPAIR_COMMIT"
printf 'Waiver record: %s\n' "$GH_RUN_URL"
} > "$RUNNER_TEMP/release-waiver.md"
if gh release view "$RELEASE_TAG" --repo "$GH_REPO" >/dev/null 2>&1; then
gh release view "$RELEASE_TAG" --repo "$GH_REPO" --json body --jq .body \
> "$RUNNER_TEMP/release-notes-existing.md"
{
cat "$RUNNER_TEMP/release-notes-existing.md"
printf '\n\n'
cat "$RUNNER_TEMP/release-waiver.md"
} > "$RUNNER_TEMP/release-notes.md"
gh release edit "$RELEASE_TAG" --repo "$GH_REPO" \
--notes-file "$RUNNER_TEMP/release-notes.md"
else
# The public notice survives a later PyPI or verification failure. Assets
# and latest-release promotion still wait for successful publication.
gh release create "$RELEASE_TAG" --repo "$GH_REPO" --verify-tag \
--generate-notes --notes-file "$RUNNER_TEMP/release-waiver.md" \
--title "Engraphis ${RELEASE_TAG#v}" --latest=false
fi

- name: Publish only missing verified distributions
uses: pypa/gh-action-pypi-publish@ba38be9e461d3875417946c167d0b5f3d385a247 # v1.14.1
with:
Expand All @@ -1130,6 +1190,7 @@ jobs:
--version "${RELEASE_TAG#v}" --retries 18 --delay 10

- name: Require signed full-product qualification before GitHub repair
if: ${{ !inputs.waive_v176_qualification }}
env:
RELEASE_TAG: ${{ inputs.release_tag }}
ENGRAPHIS_RELEASE_QUALIFICATION: ${{ secrets.ENGRAPHIS_RELEASE_QUALIFICATION }}
Expand All @@ -1147,17 +1208,44 @@ jobs:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
RELEASE_TAG: ${{ inputs.release_tag }}
WAIVE_QUALIFICATION: ${{ inputs.waive_v176_qualification }}
GH_RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
shell: bash
run: |
set -euo pipefail
notes_args=()
if [ "$WAIVE_QUALIFICATION" = "true" ]; then
{
printf '## Release qualification\n\n'
printf 'The owner waived full-product qualification for this release. Mandatory full-product gates are not represented as passed.\n\n'
printf 'Waiver record: %s\n' "$GH_RUN_URL"
} > "$RUNNER_TEMP/release-waiver.md"
notes_args=(--notes-file "$RUNNER_TEMP/release-waiver.md")
fi
if gh release view "$RELEASE_TAG" --repo "$GH_REPO" >/dev/null 2>&1; then
if [ "$WAIVE_QUALIFICATION" = "true" ]; then
gh release view "$RELEASE_TAG" --repo "$GH_REPO" --json body --jq .body \
> "$RUNNER_TEMP/release-notes-existing.md"
{
cat "$RUNNER_TEMP/release-notes-existing.md"
printf '\n\n'
cat "$RUNNER_TEMP/release-waiver.md"
} > "$RUNNER_TEMP/release-notes.md"
gh release edit "$RELEASE_TAG" --repo "$GH_REPO" \
--notes-file "$RUNNER_TEMP/release-notes.md"
fi
gh release upload "$RELEASE_TAG" verified-dist/* release-evidence/* \
--repo "$GH_REPO" \
--clobber
if [ "$WAIVE_QUALIFICATION" = "true" ]; then
gh release edit "$RELEASE_TAG" --repo "$GH_REPO" --latest
fi
else
gh release create "$RELEASE_TAG" verified-dist/* release-evidence/* \
--repo "$GH_REPO" \
--verify-tag \
--generate-notes \
"${notes_args[@]}" \
--title "Engraphis ${RELEASE_TAG#v}" \
--latest
fi
10 changes: 5 additions & 5 deletions BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,14 +94,14 @@ interpretation and do not count as additional benchmark-quality gains.
### Public numeric evidence registry

Every exact public aggregate retained below comes from the checked-in, public-safe
[`offline-fixtures-v73.json`](docs/benchmark-evidence/offline-fixtures-v73.json) artifact. Its
[`offline-fixtures-v76.json`](docs/benchmark-evidence/offline-fixtures-v76.json) artifact. Its
SHA-256 is
`aa7ed9c141afcc82fc2a05b63ed9037842f2ea8372667f3142cf9bb795833988`, also recorded in the
`2fb5ce5b2cbc21541f7ae9cad5d2ef615014c0e86f9881f00621d00986422ba8`, also recorded in the
adjacent `.sha256` file. The artifact contains no raw questions, answers, prompts, customer data,
or per-record content fingerprints.

The fixture-suite digest is
`c60a48ea025c5c0ebfdacb68c38f068fc032f05b202e5c192a9223211cf30bf1`. The artifact defines
`20131c25c86e3ac5a53285d0631d4ff0c60946879b04ff7a7c38e488f42cda51`. The artifact defines
the digest algorithm and records the SHA-256 of every suite and dataset file. Each evidence ID
also binds its exact command through `sha256(UTF-8 exact command)`:

Expand All @@ -123,10 +123,10 @@ Historical LoCoMo, graph, handoff, consolidation, and security figures remain pr
source artifacts but are omitted from the current chart until each has a matching immutable,
public-safe artifact. The chart labels coding outcomes, external datasets, and operational
capacity as pending evaluation tracks rather than implying scores. Regenerate it with
`python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v73.json --output docs/images/context-efficiency.svg` after selecting the report to publish.
`python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v76.json --output docs/images/context-efficiency.svg` after selecting the report to publish.

The companion examples are also generated from that artifact with
`python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v73.json --output docs/images/evidence-backed-agent-examples.svg`.
`python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v76.json --output docs/images/evidence-backed-agent-examples.svg`.
The historical-to-executable mapping is in
[`docs/BENCHMARK_CHANGE_COVERAGE.md`](docs/BENCHMARK_CHANGE_COVERAGE.md).

Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,9 +84,9 @@ neither is an end-to-end question-answer score. Coding outcomes, external datase
operational capacity remain separate pending evaluation tracks until their artifacts are selected.

These values are evidence IDs `offline-chunking` and `offline-performance` in
[`offline-fixtures-v73.json`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/benchmark-evidence/offline-fixtures-v73.json),
[`offline-fixtures-v76.json`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/benchmark-evidence/offline-fixtures-v76.json),
SHA-256
`aa7ed9c141afcc82fc2a05b63ed9037842f2ea8372667f3142cf9bb795833988`.
`2fb5ce5b2cbc21541f7ae9cad5d2ef615014c0e86f9881f00621d00986422ba8`.
[`BENCHMARKS.md`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/BENCHMARKS.md#public-numeric-evidence-registry)
records the matching suite digest, exact commands, and per-command config digests. The offline
fixture registry intentionally excludes external, model-dependent, consolidation, productivity,
Expand Down
21 changes: 16 additions & 5 deletions docs/RELEASE_QUALIFICATION.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,11 @@
# Owner-signed release qualification

Every new PyPI publication, GitHub release write, and repair requires a valid
Every ordinary PyPI publication, GitHub release write, and repair requires a valid
full-product qualification. Passing the public build jobs is necessary but does
not replace the mandatory private readiness evidence. The workflow fails closed
when qualification configuration is missing, malformed, expired or inconsistent
with the selected source and distribution bytes.
with the selected source and distribution bytes. The owner-authorized v1.7.6
repair waiver is the one-time exception documented below.

The public verifier is `scripts/verify_release_qualification.py`. It verifies
Ed25519 signatures using `cryptography==50.0.0` in release jobs. It contains no
Expand All @@ -15,8 +16,8 @@ independent proof that each observation happened.

## Pending owner setup

These operations have **not been performed** by this source change. Publication
will remain blocked until the release owner completes them.
These operations have **not been performed** by this source change. Ordinary
publication will remain blocked until the release owner completes them.

1. Create and protect the GitHub environment `release-qualification`. Restrict its
deployment branches/tags to the protected release sources, require an authorized
Expand Down Expand Up @@ -105,7 +106,17 @@ The normal workflow checks before both PyPI and GitHub writes. Repair first sele
a matching historical push run and verifies its exact public distribution/evidence
hashes, then checks the current owner approval before both repair writes. It uses
the peeled release tag commit, never the repair workflow's `main` checkout commit.
No workflow switch makes the qualification optional.

For the existing `v1.7.6` release only, the repository owner explicitly directed a
qualification waiver on 2026-09-27. The `workflow_dispatch` input
`waive_v176_qualification` skips the owner qualification verifier only when repairing
`v1.7.6` at commit `6a441a75c8dd159607fa3933da83f600864b9146`. Reusing that
tag for another commit cannot use this exception. The workflow records the actor
and run URL, and publishes the waiver in GitHub Release notes before the first
PyPI write. A failed disclosure prevents publication; a later repair failure
leaves the public disclosure in place. This is not a qualification and
does not mark any unverified gate as passing. All ordinary tag publications and
repairs for other versions still require a valid owner-signed qualification.

## Public installed evidence

Expand Down
4 changes: 3 additions & 1 deletion docs/RELEASE_READINESS.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,9 @@ engine checkout; `--require-leadership` requires both decisions. All modes retai
Actual publication additionally requires the protected, owner-signed approval in
[RELEASE_QUALIFICATION.md](RELEASE_QUALIFICATION.md), verified immediately before each
normal or repair write. Its environment, authority and approval remain owner setup;
this source change does not configure or issue them.
this source change does not configure or issue them. The owner-authorized v1.7.6
repair waiver documented there is an explicit exception and does not establish
full-product readiness or change any gate status.

Planner experiments remain off by default. To require their existing optimization gate:

Expand Down
2 changes: 1 addition & 1 deletion docs/REWORK_EXECUTION.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,7 +160,7 @@ scope/time regression coverage. Original user edits were not rewritten.
| Complete-engine measurement | Serializable factory configuration supports real files, pinned local semantic models, exact NumPy/sqlite-vec backends and rerankers. Opt-in recall phases separate embedding, retrieval, ranking and packing. | Full writer occupancy, production-load calibration and measured optimization remain open. |
| Capacity acceptance | Lifecycle RSS includes startup, backlog is sampled and recomputed, and the complete matrix validator enforces prebound hosts, WAL/FULL, all scheduled outcomes, RAM and the required 100k latency limits. | No primary 48-cell matrix was executed. The separate 16 GiB reference host remains necessary. |
| Installed journeys | A packaged stdlib runner performs actual MCP/HTTP writes, restart recall, correction and historical reads. PR CI and release jobs cover Windows, macOS and Linux, with artifact/dependency identities retained. | Cached Windows source semantic startup passed in four fresh processes, taking 20-23 seconds. This is not semantic qualification of all installed platforms. |
| Evidence and publication | Candidate ledger validation checks identities, hashes, dependencies, outcomes and selected evaluation booleans. All four publication/repair writes require [owner qualification](RELEASE_QUALIFICATION.md). | Owner-protected environment/authority setup, final approval and all missing mandatory evidence remain open. |
| Evidence and publication | Candidate ledger validation checks identities, hashes, dependencies, outcomes and selected evaluation booleans. Ordinary publication/repair writes require [owner qualification](RELEASE_QUALIFICATION.md). | The v1.7.6 owner waiver is a one-time exception, not qualification. Owner-protected setup, final approval and missing mandatory evidence remain open. |
| Public claims | Fresh [offline fixture evidence](benchmark-evidence/offline-fixtures-v9.json) reproduces retained public aggregates and binds the current engine/eval source. Historical v1 evidence is preserved. | Planner variants still require successful promotion gates; no retrieval default or leadership claim is promoted. |
| Website contract | Active commercial/MCP/install guidance is generated and checked against a selected shipped public contract in the website candidate. | The live portal, authenticated provider journeys and combined deployed identities require attended acceptance. |

Expand Down
Loading
Loading