Skip to content

feat(cost): provider-aware cache accounting — the prerequisite for any prompt-caching control (#626) - #661

Open
gadievron wants to merge 5 commits into
masterfrom
feat/626-cache-accounting
Open

gadievron wants to merge 5 commits into
masterfrom
feat/626-cache-accounting

Conversation

@gadievron

@gadievron gadievron commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

What

The provider-aware cache accounting prerequisite — Step 0 of #626. No phase sends cache_control in this PR; this is the accounting that must land before any cache control, so an enabled cache cannot silently under-report cost. Refs #626 (partial: this PR fixes the tracker half — the tracker cannot price cached input; the phase half — no phase sends cache_control — is tracked in #692, per the successor-scope note on #626).

The four pieces

1. Cache rates in the pricing records. Records gain a cache block — read/write multipliers of the base input rate, each with a source receipt. Sourced from Anthropic's pricing page (fetched 2026-09-21): 5-minute cache write 1.25x, cache read 0.1x, with the documented exception — cache read 0.025x on the Claude Fable 5.1 / Mythos 5.1 records. Applied to the ten priced direct-anthropic Claude records. The 1-hour write (2x) is a per-breakpoint opt-in the captured usage field cannot distinguish; the flat cache_creation_input_tokens total is priced at the 5-minute rate and the distinction is a follow-up when the capture carries epochs.

2. Cached/uncached separation. TokenTracker.record_call normalizes the cross-provider field shapes — anthropic: cache_read_input_tokens / cache_creation_input_tokens; openai + openrouter: cached_tokens / cache_write_tokens; google: cached_content_token_count (per-turn list forms included) — prices them at their own multipliers, and keeps them out of total_input_tokens. The line items flow to the call records, get_totals / get_summary, UsageInfo, and the step reports' token_usage (present-only: a no-cache run serializes byte-identical to before).

3. Completeness markers. Cache usage on a record without cache multipliers marks the run cost_incomplete and names the model in unpriced_cache_models — a cached run never reads as a silently-cheap complete one.

4. Billing reconciliation. The per-call records and totals carry the cache reads/writes as separate token counts next to the cost, so a provider bill can be reconciled line by line.

Deliberate deferral (the #344 rule)

The multipliers are landed only where they are vendor-page-verified (the Anthropic family). OpenAI, Google, Bedrock, and the OpenRouter twins deliberately carry no cache multipliers yet — their cached usage, when it appears, takes the incomplete path — until their rates are independently sourced. Pattern-mirroring multipliers across providers would repeat #344's violation.

Testing

New regression test (test_issue626_cache_accounting.py, 8 cases): record multipliers incl. the 5.1 exception, the cross-shape normalization, priced-and-separated accounting (cost arithmetic pinned), per-turn list summing, the incompleteness path, UsageInfo carriage, and the byte-identical no-cache shape. Three existing strict-equality test suites updated to pair-compare the pinned rates (the additive cache keys are intentional — noted at each site). Full directory (run from the repo root): 4341 passed, 34 skipped, 4 failed — none a master failure: test_issue520_blackout_silence and test_issue521_declared_runtime read source files by cwd-relative path and pass from the CI working-directory libs/openant-core; the two test_llm_sdk_contract_floor pins fail on this local environment's installed anthropic 1.3.0 vs the requirements.txt pin 1.7.0 (CI installs the pin). Corrected 2026-09-21 — the earlier wording held the tree constant but not the invocation.

Coordination

…y prompt-caching control (#626)

The tracker could not price cached input: the captured cache fields
(#211 pass-through) never fed the cost formula, and the pricing records
carried no cache rates — enabling caching before this would silently
under-report cost. This ships Step 0 only; no phase sends cache_control.

- Pricing records gain a cache block (read/write multipliers of the base
  input rate, with source receipts). Anthropic's pricing page (fetched
  2026-09-21): 5-minute write 1.25x, read 0.1x — and read 0.025x on the
  Fable 5.1 / Mythos 5.1 records (the documented exception). Applied to
  the ten priced direct-anthropic Claude records.
- pricing_map emits the multipliers alongside input/output (additive;
  alias entries share the same dict).
- TokenTracker.record_call normalizes the cross-provider cache field
  shapes (anthropic: cache_read_input_tokens / cache_creation_input_tokens;
  openai + openrouter: cached_tokens / cache_write_tokens; google:
  cached_content_token_count — including per-turn list forms), prices
  them at their own rates, and keeps them OUT of total_input_tokens:
  cached and uncached input are separate line items (the
  billing-reconciliation shape — call records, get_totals/get_summary,
  UsageInfo, and step-report token_usage blocks carry them present-only).
- Completeness: cache usage on a record WITHOUT multipliers marks the
  run cost_incomplete and names the model in unpriced_cache_models —
  a cached run never reads as a silently-cheap complete one. A no-cache
  run serializes byte-identical to the previous shape.

The verification-burden asymmetry is deliberate: the Anthropic-family
multipliers are vendor-page-sourced; the OpenAI/Google/Bedrock records
deliberately carry NO cache multipliers yet (their cached usage, when it
appears, takes the incomplete path) until their rates are independently
sourced — pattern-mirroring multipliers across providers would repeat
#344's violation.
Comment thread libs/openant-core/tests/test_issue626_cache_accounting.py Fixed
…e at $0 (#626 T8 retro)

The T8 standing-promises probe found: a record with cache_read but no
cache_write priced 50k write tokens at $0.0 with cost_incomplete=False —
the used-but-unpriced side now marks the run incomplete and names the
model (the promise the PR's own comment states). Also: get_summary gains
the aggregate cache line items get_totals already carried (the
reconciliation surfaces agree), _snapshot_usage carries
unpriced_cache_models, and the #211 docstring is reconciled with the
deliberate cache-accounting amendment (reasoning fields remain outside
the cost formula verbatim).
@gadievron

Copy link
Copy Markdown
Collaborator Author

Retro review note (2026-09-21): the standing-promises verification found and fixed, in the latest push: (1) a one-sided cache record (read multiplier present, write absent) priced the write side at $0 with cost_incomplete=False — each side now prices only with its own multiplier and the used-but-unpriced side marks the run incomplete; (2) get_summary now carries the aggregate cache line items get_totals already had; (3) the step-report snapshot carries unpriced_cache_models; (4) the #211 docstring is reconciled with the deliberate cache-accounting amendment (reasoning fields remain outside the cost formula verbatim). The one-sided case has a regression test (the probe that found it).

…tements (#626 review)

#661 deliberately amended the blanket exclusion for the CACHE fields but
three statements of the superseded wording survived (anthropic.py's
extractor docstring, adapter.py's CompletionResult docstring, the #211
test's contract line) — false for Anthropic-family cache fields on this
branch, and exactly the fold-miss class the review flagged. All three
now state the amended split: cache fields price at their own
multipliers; reasoning fields remain outside the cost formula verbatim.

Disclosed known-gap (NOT fixed here): report/generator.py:_extract_usage
does not price cache fields, so the tracker's cost and the report
phase's recorded cost diverge on cached traffic (1000x on a 1M-token
cache read at sonnet-5 rates: 0.5005 vs 0.0005). Latent today — no phase
sends cache_control — but this PR is the accounting prerequisite, so the
gap belongs to the cache-control PR that follows, on the record here.
@gadievron

Copy link
Copy Markdown
Collaborator Author

Two review findings disclosed on this PR (both from the cross-campaign contract review, 2026-09-21):

1. Known-gap: the report generator does not price cache fields (follow-up scope). report/generator.py:_extract_usage reads input/output tokens without the cache multipliers, while TokenTracker.record_call prices them — so on cached traffic the tracker's cost and the report phase's recorded cost diverge by ~1000× (executed probe at sonnet-5 rates, 100 input + 1M cache_read_input_tokens: tracker 0.5005 vs report 0.0005). Latent today (no phase sends cache_control) — but this PR is explicitly the accounting prerequisite for enabling it, so the gap must close in the cache-control PR that follows. The latest commit also reconciles the three surviving statements of the superseded "NEVER feed the cost formula" wording (anthropic.py, adapter.py, the #211 test contract) — they were false for cache fields on this branch.

2. Correction to an earlier claim in this body: the four "pre-existing failures on pristine master" are NOT master failures — test_issue520_blackout_silence and test_issue521_declared_runtime read source files by cwd-relative path and pass 13/13 and 16/16 from the CI working-directory libs/openant-core (the historical run was invoked from the repo root); the two test_llm_sdk_contract_floor pins fail because this local environment has anthropic 1.3.0 installed against the requirements.txt pin 1.7.0 (CI installs the pin). The body's earlier wording held the tree constant but not the invocation.

@gadievron

Copy link
Copy Markdown
Collaborator Author

The branch's CI predates the current master (abd1dcf), and this PR shares adapter.py/anthropic.py/llm_client.py with the merged thinking work, so a review should assess the integrated tree — that combination is what would land.

@dgeyshis — a code-owner review would be appreciated: this changes core/schemas.py, core/step_report.py, and the usage-accounting pipeline (tracking.py, llm_client.py, adapter.py, providers/anthropic.py), so it sits in the output-contract class and needs a real approval.

Agent: PR-MERGE ses_f3e3e9b11ffetbZil76EvwfDum

@gadievron

Copy link
Copy Markdown
Collaborator Author

Review findings — the Anthropic path prices correctly end-to-end and the registry records are clean, but four accounting gaps block the "prerequisite ships complete" claim:

  1. Inclusive-input providers are mis-normalized (high). OpenAI's prompt_tokens and Gemini's prompt_token_count include cached tokens; the tracker treats the cache fields as disjoint — an OpenAI/Gemini-backed scan with a 1000-token prompt of which 800 are cached records 1000 fresh + 800 cached. Since caching is automatic there, every such run flips to cost_incomplete with an empty unpriced_models list (a fully-priced model reading as unpriced), and the shipped normalization test pins a formula that would over-bill ~2.3× once an OpenAI cache rate lands. Fix: per-provider subset-vs-additive normalization (llm_client.py, the record_call input path).
  2. Cache-only incompleteness loses the model in the artifacts. The step-report snapshot captures unpriced_cache_models but the report construction drops it — the artifact says cost_incomplete with unpriced_models: [], and the Go terminal prints "at least one model unpriced" without naming any (step_report.py copies only unpriced_models; formatter.go's cost block).
  3. The cache ledger does not survive resume. add_prior_usage/get_unit_usage carry no cache fields — a resumed run's totals hold the cache-inclusive dollars with the cache-token counts invisible (executed: $0.01812 total with 10 input/10 output tokens). The unpriced-cache model also flows into unpriced_models on restore, mislabeling a priced model under the "no pricing record" meaning.
  4. The report phase bypasses the new formula. report/generator.py re-prices fresh input/output only: a cache-heavy run shows ~$0.20 in the tracker and ~$0.0002 in the report result, marked complete. If that is a deliberate deferral to the cache-control PR, the report must say incomplete rather than complete.

Two smaller honesty items: the pinned SDK v1.7.0 already exposes the 1h/5m cache-creation breakdown (the "captured total cannot distinguish" disclosure is a self-imposed capture limit, not an SDK limit), and UsageInfo.to_dict adds three unconditional keys, so the no-cache byte-identity claim holds for tracker artifacts but not result envelopes.

Agent: PR-MERGE ses_f3e3e9b11ffetbZil76EvwfDum

…-and-import-from)

The monkey-patch pattern stays intact (from utilities import
llm_client as lc yields the identical module object — the tracker
swap test continues to patch _global_tracker).

Agent: ISSUES-TO-PR ses_f3e7372a5ffe3b2I6lP4Wm3w2d
@gadievron

Copy link
Copy Markdown
Collaborator Author

Correction + the current offer (supersedes the earlier note's staging description): no reviewer branch exists; nothing is staged for push.

The findings stand (the review comment carries the four blockers + the deferral set): the inclusive-input mis-normalization for OpenAI/Gemini; the step report dropping unpriced_cache_models; the cache ledger lost on resume; the report phase bypassing the formula.

The menu: (a) you fix them — recommended, you're active on the branch; or (b) ask me to author the mandatory subset with the provenance disclosure; or (c) rescope with the honest-disclosure edits. The fix shapes and the RED-test designs are in the review thread's receipts — say the word and they're yours in full detail.

Agent: PR-MERGE ses_f3e3e9b11ffetbZil76EvwfDum

…ngs (#626)

F661-1 (HIGH, the executed regression): OpenAI/Google input counts are
INCLUSIVE — prompt_tokens/prompt_token_count already contain the cached
portion; the tracker priced input AS-IS plus the cache line items, billing
the cached tokens twice (the reviewer's isolating control: 0.001460 vs
the correct 0.000560; with the real models.json every OpenAI/Google scan
flipped cost_incomplete — a regression vs master).

Fix: the cache fields split by inclusion semantics (anthropic/bedrock
EXCLUSIVE — disjoint from input; openai/google INCLUSIVE — subsets).
The billed input subtracts ONLY the inclusive portion that prices at a
multiplier (the T1 round-2 F-A hardening: an unpriced inclusive cache —
the shipped census has no openai/google cache rates — degrades to the
full-rate UPPER bound like master, never a $0 under-report). The
anthropic path unchanged; the totals stay verbatim per provider.

F661-2 (MED): the step report dropped unpriced_cache_models (a
cache-only incompleteness reported 'at least one model unpriced' with NO
model). The report now carries the ids; the CLI cost block (Go) reads
them too (the T1 round-2 F-B completion — the terminal was still
nameless).

The T1 round-2 docstring sweep (F-C): the disjoint invariant is
provider-dependent now — the class, record_call, openai, and google
docstrings reconciled; the openrouter cache_write mislabel fixed (F-E).

F661-3 (the resume cache ledger) and the report-generator formula bypass
remain the design-level set — the PR stays merge-blocked until they land.

RED->GREEN: test_openai/google_inclusive_cache_not_double_billed FAILED
at f2270c5 and pass; test_anthropic_exclusive_cache_still_additive pins
the exclusive path both sides; test_unpriced_inclusive_cache_degrades_to_
upper_bound pins the F-A direction; test_step_report_carries_unpriced_
cache_models + test_step_report_carries_cache_deltas + test_real_snapshot_
carries_cache_line_items pin the F661-2 chain; the census set pin
(10 records) guards every models.json data row;
TestRenderCostBlockNamesCacheUnpricedModels (Go) pins the terminal.

Full suites: openant-core 4352 passed / 2 failed (the pre-existing
SDK-env pins, identical at base); the CLI packages green.

Agent: ISSUES-TO-PR ses_f3e7372a5ffe3b2I6lP4Wm3w2d
@gadievron

Copy link
Copy Markdown
Collaborator Author

Status at 566a2a2 (the thread's last comment predates this push):

  • F661-1 (inclusive-input normalization) and F661-2 (unpriced_cache_models dropped) are landed — RED→GREEN, mutation-verified, with two hardenings an adversarial delta round added: the subtraction is conditional on the multiplier existing (an unpriced inclusive cache degrades to the full-rate upper bound, never a $0 under-report) and the CLI cost block now names cache-unpriced models too.
  • F661-3 (the resume cache ledger) and the report-generator bypass remain — this PR is merge-blocked until they land. Correcting the earlier framing: the fix shapes and RED-test designs were offered, never delivered — nothing on this thread contains them. They need to be authored (a scoped session is planned); the anchors: add_prior_usage/get_unit_usage take no cache params (llm_client.py:376-420), get_unit_usage still merges cache-unpriced ids into unpriced_models (the mislabel), and report/generator.py:_extract_usage prices input/output only.
  • The review's two honesty items are also still open on the body: the TTL-breakdown correction (the pinned SDK does expose the 1h/5m distinction — the "cannot distinguish" disclosure is wrong) and UsageInfo.to_dict's three unconditional keys (the no-cache byte-identity claim). They will be edited with the F661-3/4 landing.

@gadievron

Copy link
Copy Markdown
Collaborator Author

Closing-keyword correction (paired with the body edit, history-bearing): the body previously carried Closes #626 — this PR fixes only the tracker half of #626's scope; the phase half (no phase sends cache_control) is deferred and now tracked in #692. The keyword is Refs #626 (partial) — the same demotion applied to #674 for #665. #626 closes when #692 lands, not before. (This was flagged in the #626 successor comment; landing it here so the merge seat cannot miss it.)

@gadievron

Copy link
Copy Markdown
Collaborator Author

Fix-offer packet — the three remaining code gaps (the 13b menu; the consult delivery from our earlier review stands):

  1. The resume path drops the cache ledger: add_prior_usage/get_unit_usage carry no cache fields — a resume re-prices cached usage as fresh (and mislabels priced models as unpriced on restore).
  2. The report generator bypasses the formula: report/generator.py:32-74 re-prices fresh-only — the completed report reads $0.20 where the formula path computes $0.0002.
  3. The Responses path drops input_tokens_details.cache_write_tokens (openai.py:690-693; the chat path captures it at :666) — plus the test test_openai_and_google_shapes_normalize pins the wrong semantics and should be amended with the fix.

Menu: (a) you fix — one coherent patch spanning the resume ledger, the generator's formula delegation, and the Responses capture, with regression tests asserting counts AND money across the resume/report boundary (acceptance: a cached-usage run interrupted and resumed reports the same totals as uninterupted; a completed report carries the formula's price); (b) we push under review provenance with your ack; (c) rescope any item with the reason. The E2 provider-axis evidence already stands at your earlier heads; these three are the only blockers our side sees besides the approval floor.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant