feat(cost): provider-aware cache accounting — the prerequisite for any prompt-caching control (#626) - #661
feat(cost): provider-aware cache accounting — the prerequisite for any prompt-caching control (#626)#661gadievron wants to merge 5 commits into
Conversation
…y prompt-caching control (#626) The tracker could not price cached input: the captured cache fields (#211 pass-through) never fed the cost formula, and the pricing records carried no cache rates — enabling caching before this would silently under-report cost. This ships Step 0 only; no phase sends cache_control. - Pricing records gain a cache block (read/write multipliers of the base input rate, with source receipts). Anthropic's pricing page (fetched 2026-09-21): 5-minute write 1.25x, read 0.1x — and read 0.025x on the Fable 5.1 / Mythos 5.1 records (the documented exception). Applied to the ten priced direct-anthropic Claude records. - pricing_map emits the multipliers alongside input/output (additive; alias entries share the same dict). - TokenTracker.record_call normalizes the cross-provider cache field shapes (anthropic: cache_read_input_tokens / cache_creation_input_tokens; openai + openrouter: cached_tokens / cache_write_tokens; google: cached_content_token_count — including per-turn list forms), prices them at their own rates, and keeps them OUT of total_input_tokens: cached and uncached input are separate line items (the billing-reconciliation shape — call records, get_totals/get_summary, UsageInfo, and step-report token_usage blocks carry them present-only). - Completeness: cache usage on a record WITHOUT multipliers marks the run cost_incomplete and names the model in unpriced_cache_models — a cached run never reads as a silently-cheap complete one. A no-cache run serializes byte-identical to the previous shape. The verification-burden asymmetry is deliberate: the Anthropic-family multipliers are vendor-page-sourced; the OpenAI/Google/Bedrock records deliberately carry NO cache multipliers yet (their cached usage, when it appears, takes the incomplete path) until their rates are independently sourced — pattern-mirroring multipliers across providers would repeat #344's violation.
…e at $0 (#626 T8 retro) The T8 standing-promises probe found: a record with cache_read but no cache_write priced 50k write tokens at $0.0 with cost_incomplete=False — the used-but-unpriced side now marks the run incomplete and names the model (the promise the PR's own comment states). Also: get_summary gains the aggregate cache line items get_totals already carried (the reconciliation surfaces agree), _snapshot_usage carries unpriced_cache_models, and the #211 docstring is reconciled with the deliberate cache-accounting amendment (reasoning fields remain outside the cost formula verbatim).
|
Retro review note (2026-09-21): the standing-promises verification found and fixed, in the latest push: (1) a one-sided cache record (read multiplier present, write absent) priced the write side at $0 with |
…tements (#626 review) #661 deliberately amended the blanket exclusion for the CACHE fields but three statements of the superseded wording survived (anthropic.py's extractor docstring, adapter.py's CompletionResult docstring, the #211 test's contract line) — false for Anthropic-family cache fields on this branch, and exactly the fold-miss class the review flagged. All three now state the amended split: cache fields price at their own multipliers; reasoning fields remain outside the cost formula verbatim. Disclosed known-gap (NOT fixed here): report/generator.py:_extract_usage does not price cache fields, so the tracker's cost and the report phase's recorded cost diverge on cached traffic (1000x on a 1M-token cache read at sonnet-5 rates: 0.5005 vs 0.0005). Latent today — no phase sends cache_control — but this PR is the accounting prerequisite, so the gap belongs to the cache-control PR that follows, on the record here.
|
Two review findings disclosed on this PR (both from the cross-campaign contract review, 2026-09-21): 1. Known-gap: the report generator does not price cache fields (follow-up scope). 2. Correction to an earlier claim in this body: the four "pre-existing failures on pristine master" are NOT master failures — |
|
The branch's CI predates the current master ( @dgeyshis — a code-owner review would be appreciated: this changes Agent: PR-MERGE ses_f3e3e9b11ffetbZil76EvwfDum |
|
Review findings — the Anthropic path prices correctly end-to-end and the registry records are clean, but four accounting gaps block the "prerequisite ships complete" claim:
Two smaller honesty items: the pinned SDK v1.7.0 already exposes the 1h/5m cache-creation breakdown (the "captured total cannot distinguish" disclosure is a self-imposed capture limit, not an SDK limit), and Agent: PR-MERGE ses_f3e3e9b11ffetbZil76EvwfDum |
…-and-import-from) The monkey-patch pattern stays intact (from utilities import llm_client as lc yields the identical module object — the tracker swap test continues to patch _global_tracker). Agent: ISSUES-TO-PR ses_f3e7372a5ffe3b2I6lP4Wm3w2d
|
Correction + the current offer (supersedes the earlier note's staging description): no reviewer branch exists; nothing is staged for push. The findings stand (the review comment carries the four blockers + the deferral set): the inclusive-input mis-normalization for OpenAI/Gemini; the step report dropping The menu: (a) you fix them — recommended, you're active on the branch; or (b) ask me to author the mandatory subset with the provenance disclosure; or (c) rescope with the honest-disclosure edits. The fix shapes and the RED-test designs are in the review thread's receipts — say the word and they're yours in full detail. Agent: PR-MERGE ses_f3e3e9b11ffetbZil76EvwfDum |
…ngs (#626) F661-1 (HIGH, the executed regression): OpenAI/Google input counts are INCLUSIVE — prompt_tokens/prompt_token_count already contain the cached portion; the tracker priced input AS-IS plus the cache line items, billing the cached tokens twice (the reviewer's isolating control: 0.001460 vs the correct 0.000560; with the real models.json every OpenAI/Google scan flipped cost_incomplete — a regression vs master). Fix: the cache fields split by inclusion semantics (anthropic/bedrock EXCLUSIVE — disjoint from input; openai/google INCLUSIVE — subsets). The billed input subtracts ONLY the inclusive portion that prices at a multiplier (the T1 round-2 F-A hardening: an unpriced inclusive cache — the shipped census has no openai/google cache rates — degrades to the full-rate UPPER bound like master, never a $0 under-report). The anthropic path unchanged; the totals stay verbatim per provider. F661-2 (MED): the step report dropped unpriced_cache_models (a cache-only incompleteness reported 'at least one model unpriced' with NO model). The report now carries the ids; the CLI cost block (Go) reads them too (the T1 round-2 F-B completion — the terminal was still nameless). The T1 round-2 docstring sweep (F-C): the disjoint invariant is provider-dependent now — the class, record_call, openai, and google docstrings reconciled; the openrouter cache_write mislabel fixed (F-E). F661-3 (the resume cache ledger) and the report-generator formula bypass remain the design-level set — the PR stays merge-blocked until they land. RED->GREEN: test_openai/google_inclusive_cache_not_double_billed FAILED at f2270c5 and pass; test_anthropic_exclusive_cache_still_additive pins the exclusive path both sides; test_unpriced_inclusive_cache_degrades_to_ upper_bound pins the F-A direction; test_step_report_carries_unpriced_ cache_models + test_step_report_carries_cache_deltas + test_real_snapshot_ carries_cache_line_items pin the F661-2 chain; the census set pin (10 records) guards every models.json data row; TestRenderCostBlockNamesCacheUnpricedModels (Go) pins the terminal. Full suites: openant-core 4352 passed / 2 failed (the pre-existing SDK-env pins, identical at base); the CLI packages green. Agent: ISSUES-TO-PR ses_f3e7372a5ffe3b2I6lP4Wm3w2d
|
Status at 566a2a2 (the thread's last comment predates this push):
|
|
Closing-keyword correction (paired with the body edit, history-bearing): the body previously carried |
|
Fix-offer packet — the three remaining code gaps (the 13b menu; the consult delivery from our earlier review stands):
Menu: (a) you fix — one coherent patch spanning the resume ledger, the generator's formula delegation, and the Responses capture, with regression tests asserting counts AND money across the resume/report boundary (acceptance: a cached-usage run interrupted and resumed reports the same totals as uninterupted; a completed report carries the formula's price); (b) we push under review provenance with your ack; (c) rescope any item with the reason. The E2 provider-axis evidence already stands at your earlier heads; these three are the only blockers our side sees besides the approval floor. |
What
The provider-aware cache accounting prerequisite — Step 0 of #626. No phase sends
cache_controlin this PR; this is the accounting that must land before any cache control, so an enabled cache cannot silently under-report cost. Refs #626 (partial: this PR fixes the tracker half — the tracker cannot price cached input; the phase half — no phase sends cache_control — is tracked in #692, per the successor-scope note on #626).The four pieces
1. Cache rates in the pricing records. Records gain a
cacheblock — read/write multipliers of the base input rate, each with a source receipt. Sourced from Anthropic's pricing page (fetched 2026-09-21): 5-minute cache write 1.25x, cache read 0.1x, with the documented exception — cache read 0.025x on the Claude Fable 5.1 / Mythos 5.1 records. Applied to the ten priced direct-anthropic Claude records. The 1-hour write (2x) is a per-breakpoint opt-in the captured usage field cannot distinguish; the flatcache_creation_input_tokenstotal is priced at the 5-minute rate and the distinction is a follow-up when the capture carries epochs.2. Cached/uncached separation.
TokenTracker.record_callnormalizes the cross-provider field shapes — anthropic:cache_read_input_tokens/cache_creation_input_tokens; openai + openrouter:cached_tokens/cache_write_tokens; google:cached_content_token_count(per-turn list forms included) — prices them at their own multipliers, and keeps them out oftotal_input_tokens. The line items flow to the call records,get_totals/get_summary,UsageInfo, and the step reports'token_usage(present-only: a no-cache run serializes byte-identical to before).3. Completeness markers. Cache usage on a record without cache multipliers marks the run
cost_incompleteand names the model inunpriced_cache_models— a cached run never reads as a silently-cheap complete one.4. Billing reconciliation. The per-call records and totals carry the cache reads/writes as separate token counts next to the cost, so a provider bill can be reconciled line by line.
Deliberate deferral (the #344 rule)
The multipliers are landed only where they are vendor-page-verified (the Anthropic family). OpenAI, Google, Bedrock, and the OpenRouter twins deliberately carry no cache multipliers yet — their cached usage, when it appears, takes the incomplete path — until their rates are independently sourced. Pattern-mirroring multipliers across providers would repeat #344's violation.
Testing
New regression test (
test_issue626_cache_accounting.py, 8 cases): record multipliers incl. the 5.1 exception, the cross-shape normalization, priced-and-separated accounting (cost arithmetic pinned), per-turn list summing, the incompleteness path,UsageInfocarriage, and the byte-identical no-cache shape. Three existing strict-equality test suites updated to pair-compare the pinned rates (the additive cache keys are intentional — noted at each site). Full directory (run from the repo root): 4341 passed, 34 skipped, 4 failed — none a master failure:test_issue520_blackout_silenceandtest_issue521_declared_runtimeread source files by cwd-relative path and pass from the CI working-directorylibs/openant-core; the twotest_llm_sdk_contract_floorpins fail on this local environment's installed anthropic 1.3.0 vs therequirements.txtpin 1.7.0 (CI installs the pin). Corrected 2026-09-21 — the earlier wording held the tree constant but not the invocation.Coordination