test: parallel assertion judges with execution receipts and negative controls (J2a) - #65
Merged
Merged
Conversation
Add a tracked-corpus assertion matrix with isolated sources, repeated runs, timeouts and regression reports. Port the seeded differential fuzzer with wrong-value reduction and record the current main findings. Parallelize HIR judge input sweeps with private CPUs/images and ordered diagnostics. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wrwo36SzxKRZYTgpDo7hzo
Enumerate assertions with frontend JSON listings and require execution receipts for isolated checks and shared sandbox units. Add expected-value controls, separate baseline/candidate inventories, flaky classification and strict tool errors. Diagnose malformed C directives and accept trailing comments. Distinguish MIR2/oracle discrepancies, Z80/MIR2 discrepancies and assembly failures in differential fuzzing, and correct the literal-width findings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wrwo36SzxKRZYTgpDo7hzo
Record the three pre-existing Z80 C assertion failures and exempt only exact candidate diagnostics. Fail the gate when an entry passes, changes or loses its assertion, and retain negative control and inventory safeguards. Add tests for matching, stale entries, repeated outcomes and CLI behavior. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wrwo36SzxKRZYTgpDo7hzo
Apply assertion filters and negative controls only when explicitly requested. Check all Z80 tuple returns and preserve assertion source lines after #embed. Run separate Z80 and MIR2 matrix passes across all assertion frontends, and prevent known failures from exempting baseline-passing regressions. Add default-mode source comparisons and regression tests for each fix. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wrwo36SzxKRZYTgpDo7hzo
Make execution receipts opt-in and compare complete emitted files and raw compiler diagnostics in the default sweep. Add tracked tuple assertions that catch a checker ignoring later results, restore optional backend behavior, and support repeated controls and controls-only CI passes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wrwo36SzxKRZYTgpDo7hzo
Generate independent negative controls for every tuple element and verify that the matrix rejects both first-only and first-skipping checkers. Restore Z80 assertion execution in LLVM mode and empty-return acceptance for WASM sandbox assertions, with regressions and CI control guidance. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wrwo36SzxKRZYTgpDo7hzo
PR tests: 🟢 passCommit:
The external CZECH and full Zork I stories are not part of this PR gate. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
J2a: fast, honest assertion judges
The old assertion matrix had false passes: 9 asserts never executed and were still counted as passing. This PR makes the regression gate show what actually ran, and adds negative controls so the gate can show it detects breakage.
What it adds
scripts/assert_matrix.py,scripts/fuzz_diff.py,scripts/compile_sweep.py), covering all 9 frontends: .c, .nanz, .pas, .plm, .abap, .frl, .lanz, .lizp, .m.--list-asserts, including the tuple arity). With--assert-receiptit prints an execution receipt:ASSERTS: executed=N passed=P failed=F. The receipt is opt-in, so default compiler output is unchanged.examples/nanz/tuple_assert_gate.nanzis tracked.scripts/test_assert_mutants.py: builds checker mutants and requires the matrix to reject them:Numbers (vs origin/main 05402e1 with the listing instrumentation applied; -j16, controls on)
--no-controls: 263 s)..pasasserts dropped from the listingBehaviour kept compatible
--asserts wasm|llvmstill runs MIR2 asserts, andllvmmode still runs Z80 asserts. WASM sandbox asserts still skip a missing export and accept void results. P7's CLI behaviour is unchanged.Recommended CI use (follow-up C2)
--runs 2.--no-controlscannot detect a weakened checker.Gate
-race ./pkg/hir, hir/mir2/c89/pascal/plm/cmd,-shortpipeline.scripts/test_judges.py: 25/25.🤖 Generated with Claude Code
https://claude.ai/code/session_01Wrwo36SzxKRZYTgpDo7hzo