Skip to content

feat(metrics,cli): the trend says when periods are not comparable by size (#102); report wording; codex default - #105

Merged
ceccode merged 1 commit into
mainfrom
fix/trend-comparability
Oct 1, 2026
Merged

ceccode merged 1 commit into
mainfrom
fix/trend-comparability

Conversation

@ceccode

@ceccode ceccode commented Sep 19, 2026

Copy link
Copy Markdown
Owner

Closes #102. The third of three PRs from the 18 September external review; independent of #103 and #104 (all from main).

Comparability by size

Maturity answers "have both periods had the same time?". Nothing answered "is there enough in each, and are they alike enough, for a delta to mean anything?". Found on this repository: 5 commits over 18 files against 18 over 53, two empty months between, rendered as +60.9 pt. The arithmetic was right; the reading it invited was not.

  • latestComparison.comparability (additive): eligible files and authored commits per side; status ok · weak (sizes differ beyond 3×) · insufficient (a side has fewer than 10 eligible files); reasons spelled out.
  • Report: a weak delta is shown and labelled (a change of pace is being measured along with any change in the code); an insufficient one is withheld with the reason, rows still shown. The pair is never swapped for a friendlier one — choosing the comparison to make it look valid would be worse than a labelled weak one.

What the label finds on real repositories (same HEAD as the other two PRs)

Repo Headline before Now
commander.js 2026-04 → 2026-05: 0.0% → 5.9% withheld — 2026-04 has 1 eligible file. The old headline was a one-file comparison.
react-router 27.4% → 48.5% weak — 649 vs 165 eligible files (3.9×)
aspire 48.7% → 33.0% weak — 319 vs 104 commits (3.1×)
evidtrail 27.8% → 88.7% weak — 5 vs 18 commits (3.6×)
tailwindcss 63.4% → 39.3% ok

Three of five headline comparisons were on pairs that do not hold, and one of them was built on a single file. No number changed; what changed is that the report now says so.

Words, and one number the report was missing

  • "how code holds up per autonomy level" → "how often code is touched again, per autonomy level"; "below is better than average" → "below is fewer than its share predicts".
  • attribution.aiModeUnknown (additive) and a Data Quality line: AI commits whose evidence names a tool but no autonomy level. Coverage counts them as evidence of involvement; the autonomy sections count them as unknown. On this repository that is 36 of 99 AI commits under 100% coverage — exactly the figure the review pointed at. On aspire it is 1,614 of 2,466; on react-router 22 of 22. Full coverage was reading as full knowledge.
  • A --since run now says next to the table that its population differs from a full-history run (a file's clock starts at its first touch inside the window: 68% vs 90% on the same repository). Once fix(metrics): observation stops at --until — truncated files are too recent, truncated months immature #103 lands the label shows the resolved instant rather than the flag text.

Codex

codex joins DEFAULT_TOOLS: "generated by Codex" returned no evidence while the same sentence with Claude was inferred AI. The openai co-author domain was already recognised. On the six dogfood repositories this changes no count — none mention Codex in a message — which is the honest thing to report. The hook does not auto-detect Codex: its documentation lists no environment variable set for spawned commands, and a detection that could be wrong is worse than an honest unknown. The README says to declare with EVIDTRAIL_MODE=agent.

Tests

  • trend.test.ts: weak pair (delta still computed, both reasons named); insufficient pair withheld; alike pair ok.
  • index.test.ts: aiModeUnknown counts tool-only AI commits and agrees with modes.unknown.
  • ai-tags.test.ts: Codex inferred, level unknown.

pnpm build / typecheck / lint green; tests 307 → 312.

@github-actions

github-actions Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

evidtrail ✅ 1 commit — agent 1 — every commit in this change set carries provenance.

Details — scope, provenance, limits

Scope: base..HEAD @ 14bd61cbabfd — 10 files touched. A change-set view, not repository history or deployed state; time-based signals (rapid retouch, trend) are omitted because a fresh PR has not had a comparable observation window.

Evidence: 100% coverage — declared 1 · inferred 0 · none 0.

Autonomy level Commits
agent 1

Limits

  • Evidence coverage is 100.0%: a missing signal remains unknown; it is not evidence of human authorship and a defaultMode prior does not turn it into observed provenance.
  • Commit scope is pr at 14bd61c: only commits in base..HEAD are included. This is a change-set view, not repository history, merge status, or deployed production state.
  • Rapid-retouch rates and trends are omitted from the PR report because a fresh change set has not had a comparable observation window.
  • AI tagging uses declarations and conservative heuristics; tool use that leaves no commit evidence cannot be recovered from git alone.

…; report wording; codex default

53, two empty months between, rendered as a +60.9 pt quality change. The
arithmetic was right; the reading it invited was not. Maturity answered
"same time?"; nothing answered "enough, and alike enough?".

- metrics: `latestComparison.comparability` (additive) — eligible files
  and authored commits per side, status ok | weak (sizes differ beyond
  3×) | insufficient (a side under 10 eligible files), reasons spelled
  out. The delta is still computed: the pair is labelled, never swapped
  for a friendlier one.
- cli: weak → delta shown and labelled; insufficient → withheld with the
  reason, rows still shown. "How code holds up" → "how often code is
  touched again"; "below is better than average" → "below is fewer than
  its share predicts". Data Quality names the AI commits with a tool
  signal but no autonomy level (`attribution.aiModeUnknown`, additive):
  100% coverage is not 100% known autonomy. A `--since` run says its
  population differs from full history.
- core: `codex` joins the default tool names — "generated by Codex" was
  returning no evidence. The hook does not auto-detect Codex: it
  documents no environment variable for spawned commands, and a guess
  that could be wrong is worse than an honest unknown (README).

On this repository the 2026-04 → 2026-07 comparison is now labelled weak:
5 vs 18 commits (3.6×).

AI-Mode: agent
@ceccode
ceccode force-pushed the fix/trend-comparability branch from deca227 to 7d0905e Compare October 1, 2026 06:56
@ceccode
ceccode merged commit 46bd263 into main Oct 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Trend compares periods that are not comparable by size

1 participant