Skip to content

feat(quota): read the GLM Coding Plan quota for Z.ai and BigModel upstreams - #204

Merged
fylorn merged 1 commit into
mainfrom
claude/core-glm-quota-7028g0
Sep 25, 2026
Merged

fylorn merged 1 commit into
mainfrom
claude/core-glm-quota-7028g0

Conversation

@fylorn

@fylorn fylorn commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Requested by F · project thread

What this changes

Upstreams on api.z.ai or open.bigmodel.cn (written by the Z.ai / BigModel login or added by hand) now show their GLM Coding Plan quota in GET /quota, QuotaSeen and QuotaExhausted: 5-hour, weekly and monthly windows, plus total / used / remaining credits on credit-based plans. A 429 whose body carries a used-up business code marks that window used up immediately.

Why

GLM responses carry no quota headers, so the existing header-based sampling never saw anything for these accounts. The only sources are the account's quota endpoint (/api/monitor/usage/quota/limit, undocumented; shape taken from Z.ai's own open-source tools, logic reimplemented) and the business code in a 429 body.

How it works

  • Detection (tw_gateway::glm::Sites::quota_url): by the host of base_url only; the quota URL is chosen by the matching site. Tests swap the sites for a local fake, like zai::Endpoints.
  • When to ask (glm::Tracker): GET /quota asks each due GLM upstream, coalesced to once per minute, and waits at most 5s (slower answers still land via QuotaSeen). Traffic to a GLM upstream (served, 5xx or 429) asks at most every 5 minutes, in the background. Failures back off 30s / 60s / 120s / 300s. A key with no plan (code: 500, success: false, or no recognizable window) is recorded as "no quota data" (removed from /quota) and re-asked after an hour or when the key changes.
  • Parsing (glm::parse): success is success !== false and code in {absent, 0, 200}; msg is never used. Auth failures (HTTP 401/403, or code 401/1000/1001) are told apart from other failures. Windows come from unit/number (3/5 → 5h, 6/any → weekly, TIME_LIMIT 5/1 → monthly), never from reset order; unrecognized items are dropped. A nextResetTime that is missing, in the past, or later than the window's length is None.
  • 429 (glm::exhausted, AppState::note_glm_429): code from error.type (Anthropic endpoint) or error.code (OpenAI endpoint). 1308, 1310, 1316–1321 count as used up; 1309, 1313, 1302, 1305 do not. The window comes from the message (English or Chinese), then the code, then the tightest known candidate. The reset time is the quota endpoint's when known, otherwise the one in the message, parsed as UTC+8. The body read is capped (16 KiB, 2s) so failover isn't held up.

Contract

  • New tw_api::QuotaCredits { total, used, remaining } and QuotaWindow.credits: Option<QuotaCredits> (also on tw_gateway::quota::Window). remaining is passed through as given, not computed (the real example has 2000 − 23 ≠ 1976).
  • Window vocabulary gains monthly.
  • CONTROL_API_VERSION 22 → 23. Lite needs to regenerate src/generated/tw-api.ts.
  • No msg! sentences and no config changes, so the msg-code table and config manual are unchanged.
  • tw-dialect, tw-guard and tw-breaker are untouched.

How it was verified

  • cargo fmt --all -- --check, cargo clippy --workspace --all-targets -- -D warnings (stable 1.98.1), cargo test --workspace: all green. ./scripts/smoke.sh: 60 passed, 0 failed.
  • 22 unit tests in tw-gateway/src/glm.rs: the real credit-plan example, old plan V1 (5h + monthly MCP) and V2 (+ weekly with number: 7), no plan (Chinese and English msg), auth failures, missing / past / too-late nextResetTime, unnamed windows, every 429 code in both JSON shapes and both languages, and the ask cadence (1-minute coalescing, 5-minute traffic limit, backoff sequence, no-plan hour, key change).
  • 4 tests in state/glm.rs and 4 end-to-end tests in tw-gateway/tests/glm_quota.rs against a local fake GLM: /quota-style refresh with credits; Authorization carries the raw key with no Bearer; no-plan is not re-asked; a 1308 429 through the gateway emits one QuotaExhausted with the message's UTC+8 reset time and triggers a quota ask; a 1302 429 changes nothing.
  • No real key and no real endpoint was used. The key only goes to the quota URL on the upstream's own host, through the upstream's own client and proxy, and is never logged. The tracker keeps a blake3 fingerprint of the key, not the key.

Notes for review

  • Business code 500 with success: false is read as "no plan", because msg is the only other signal and it changes with language. A transient 500 is therefore treated as no plan for up to an hour.
  • If the quota endpoint lags behind a 429 (still reports < 100%), the next ask clears the used-up state. This follows the rule that the endpoint is authoritative.
  • Real-device check still needed: a real GLM Coding Plan key (credit-based and old plan) should show the windows and credits in /quota. This can't be done from CI.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WjXngih1kA6oBMoqCkXDVx


Generated by Claude Code

…treams

GLM Coding Plan responses carry no quota headers, so the quota view stayed
empty for Z.ai / BigModel accounts. The only source is the account's quota
endpoint (/api/monitor/usage/quota/limit), and a used-up window only shows
up as a business code inside a 429 body.

- An upstream whose base_url host is api.z.ai or open.bigmodel.cn is a GLM
  Coding Plan upstream, whether the login wrote it or it was added by hand.
- GET /quota asks each due GLM upstream (at most once a minute) and waits up
  to 5s; traffic to one asks at most every 5 minutes. Failures back off
  30s/60s/120s/300s, and a key without a plan is left alone for an hour.
- Windows are told apart by unit/number, never by reset order: 5h, weekly,
  and monthly (the old plans' MCP calls). A reset time past the window's
  length is dropped rather than shown as a wrong countdown.
- Credit-based plans carry total/used/remaining as the endpoint gives them
  (QuotaWindow.credits, new QuotaCredits type).
- A 429 with 1308, 1310 or 1316-1321 marks the window used up and emits
  QuotaExhausted; the reset time comes from the quota endpoint when known,
  otherwise from the message (read as UTC+8). 1309, 1313, 1302 and 1305 are
  not a used-up quota.

CONTROL_API_VERSION is now 23: QuotaWindow gained `credits` and the window
vocabulary gained `monthly`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WjXngih1kA6oBMoqCkXDVx
@fylorn
fylorn force-pushed the claude/core-glm-quota-7028g0 branch from b7cf0b5 to 3edde89 Compare September 25, 2026 16:33
@fylorn
fylorn merged commit 6f32b1a into main Sep 25, 2026
4 checks passed
@fylorn
fylorn deleted the claude/core-glm-quota-7028g0 branch September 25, 2026 16:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants