Repository navigation
feat(metrics,cli): the trend says when periods are not comparable by size (#102); report wording; codex default - #105
Merged
Conversation
Contributor
|
evidtrail ✅ 1 commit — agent 1 — every commit in this change set carries provenance. Details — scope, provenance, limitsScope: Evidence: 100% coverage — declared 1 · inferred 0 · none 0.
Limits
|
…; report wording; codex default 53, two empty months between, rendered as a +60.9 pt quality change. The arithmetic was right; the reading it invited was not. Maturity answered "same time?"; nothing answered "enough, and alike enough?". - metrics: `latestComparison.comparability` (additive) — eligible files and authored commits per side, status ok | weak (sizes differ beyond 3×) | insufficient (a side under 10 eligible files), reasons spelled out. The delta is still computed: the pair is labelled, never swapped for a friendlier one. - cli: weak → delta shown and labelled; insufficient → withheld with the reason, rows still shown. "How code holds up" → "how often code is touched again"; "below is better than average" → "below is fewer than its share predicts". Data Quality names the AI commits with a tool signal but no autonomy level (`attribution.aiModeUnknown`, additive): 100% coverage is not 100% known autonomy. A `--since` run says its population differs from full history. - core: `codex` joins the default tool names — "generated by Codex" was returning no evidence. The hook does not auto-detect Codex: it documents no environment variable for spawned commands, and a guess that could be wrong is worse than an honest unknown (README). On this repository the 2026-04 → 2026-07 comparison is now labelled weak: 5 vs 18 commits (3.6×). AI-Mode: agent
ceccode
force-pushed
the
fix/trend-comparability
branch
from
October 1, 2026 06:56
deca227 to
7d0905e
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #102. The third of three PRs from the 18 September external review; independent of #103 and #104 (all from
main).Comparability by size
Maturity answers "have both periods had the same time?". Nothing answered "is there enough in each, and are they alike enough, for a delta to mean anything?". Found on this repository: 5 commits over 18 files against 18 over 53, two empty months between, rendered as +60.9 pt. The arithmetic was right; the reading it invited was not.
latestComparison.comparability(additive): eligible files and authored commits per side; statusok·weak(sizes differ beyond 3×) ·insufficient(a side has fewer than 10 eligible files); reasons spelled out.What the label finds on real repositories (same HEAD as the other two PRs)
Three of five headline comparisons were on pairs that do not hold, and one of them was built on a single file. No number changed; what changed is that the report now says so.
Words, and one number the report was missing
attribution.aiModeUnknown(additive) and a Data Quality line: AI commits whose evidence names a tool but no autonomy level. Coverage counts them as evidence of involvement; the autonomy sections count them asunknown. On this repository that is 36 of 99 AI commits under 100% coverage — exactly the figure the review pointed at. On aspire it is 1,614 of 2,466; on react-router 22 of 22. Full coverage was reading as full knowledge.--sincerun now says next to the table that its population differs from a full-history run (a file's clock starts at its first touch inside the window: 68% vs 90% on the same repository). Once fix(metrics): observation stops at --until — truncated files are too recent, truncated months immature #103 lands the label shows the resolved instant rather than the flag text.Codex
codexjoinsDEFAULT_TOOLS: "generated by Codex" returned no evidence while the same sentence with Claude was inferred AI. Theopenaico-author domain was already recognised. On the six dogfood repositories this changes no count — none mention Codex in a message — which is the honest thing to report. The hook does not auto-detect Codex: its documentation lists no environment variable set for spawned commands, and a detection that could be wrong is worse than an honestunknown. The README says to declare withEVIDTRAIL_MODE=agent.Tests
trend.test.ts: weak pair (delta still computed, both reasons named); insufficient pair withheld; alike pair ok.index.test.ts:aiModeUnknowncounts tool-only AI commits and agrees withmodes.unknown.ai-tags.test.ts: Codex inferred, level unknown.pnpm build/typecheck/lintgreen; tests 307 → 312.