Verification-gated skill routing and self-improvement harness for Hermes-style agent skills
-
Updated
Sep 26, 2026 - Python
Verification-gated skill routing and self-improvement harness for Hermes-style agent skills
Release gate that catches "green build, healthy container, broken page" deploys: CSS audit, computed-style assertions, screenshot diff, and Docker candidate promotion with verified rollback.
The deterministic merge gate for AI-generated agent capability changes — a local-first, static Tool-Use Readiness review for MCP, OpenAPI, and SDK tool surfaces. Open-source CLI + GitHub Action.
Black-box-first skill suite for manual test design, risk-based prioritization, minimal test case synthesis, and release gate decisions. / ブラックボックス前提で、手動テストの観点出し・リスク優先度付け・最小ケース化・リリースゲート判定まで行う skills
Deterministic, evidence-backed QA runtime for Codex, Claude Code, and Cursor. Versioned artifact contracts, Playwright browser driver, redacted evidence, reproducible release gates. Never calls an LLM.
Release gate for public GitHub repositories. Give it a repo URL and it returns a ship-or-block verdict with file-and-line evidence, the fixes that clear the gate, and the deployment plan and launch copy you ship with. HTTP + MCP. Static analysis only; no code execution.
AI-powered release gate for SAP Transport Requests. CLI collects evidence from ADT; SKILL guides LLMs to review code, config, interfaces, and release risk — online or offline.
Risk-based release gate for CRA (Article 14), NIS2 and DORA compliance with SBOM, KEV correlation, and auditable evidence.
Connect Automotive Cybersecurity Risk to Reality — automotive cybersecurity digital-thread & evidence-graph platform: TARA, requirements traceability, ISO/SAE 21434, UN R155/R156, SBOM, change impact analysis and cybersecurity release gates.
Claude Code QA agent with memory, plus a read-only MCP server (verdict-qa-mcp on PyPI): baseline → delta runs (NEW/REGRESSED), a fact harness so the model judges while the system measures, signed run history, flaky quarantine with expiry, and an exit-code release gate for CI. Ships six scored eval fixtures — misses published.
Human-in-the-loop AI release gate: Gemini 3 on Vertex AI reasons over Arize Phoenix telemetry to approve, block, or guardrail a version promotion, with idempotent gated writes. Google Cloud Rapid Agent Hackathon.
Validate, pressure-test, and release-gate SKILL.md packages for OpenAI and portable agent runtimes.
Stateful AI agent rehearsal and outcome verification — inject failures, verify final state, block regressions.
QA Architecture Blueprint — Python/Playwright/Pytest, Docker-first CI Gate Chain, Release-Readiness Governance, Agentic QA Workflow Foundation
ML release-gate checks for drift, performance regression, latency, and JSON/Markdown deployment reports.
Evidence-first product quality checks for web apps
Evidence-backed release audits for AI agent harnesses, covering loop correctness, tool safety, context integrity, recovery, and production readiness. 为 AI Agent 运行底座做有证据的上线审查,验证循环、工具、上下文、故障恢复与生产安全。
Autonomous release-quality agent for AI apps: matched regression testing, counterfactual root-cause replay, CI gate, JUnit/SARIF.
Side-effect governance, approvals, crash-safe idempotency, reconciliation, and audit for AI agents.
AI 写嵌入式固件的闭环验证工作台——代码可由 AI 生成,但"算不算对"由硬件与门禁说了算:需求 FSD 对账 → 四态判定(PASS/XFAIL/XPASS/FAIL)→ G0-G3 发布门禁 → R1-R8 事后审计 → 知识库校准沉淀;真机(STM32F103/OpenOCD/RTT)与无板仿真(QEMU/semihosting)双后端
To associate your repository with the release-gate topic, visit your repo's landing page and select "manage topics."