feat(quota): read the GLM Coding Plan quota for Z.ai and BigModel upstreams - #204
Merged
Merged
Conversation
…treams GLM Coding Plan responses carry no quota headers, so the quota view stayed empty for Z.ai / BigModel accounts. The only source is the account's quota endpoint (/api/monitor/usage/quota/limit), and a used-up window only shows up as a business code inside a 429 body. - An upstream whose base_url host is api.z.ai or open.bigmodel.cn is a GLM Coding Plan upstream, whether the login wrote it or it was added by hand. - GET /quota asks each due GLM upstream (at most once a minute) and waits up to 5s; traffic to one asks at most every 5 minutes. Failures back off 30s/60s/120s/300s, and a key without a plan is left alone for an hour. - Windows are told apart by unit/number, never by reset order: 5h, weekly, and monthly (the old plans' MCP calls). A reset time past the window's length is dropped rather than shown as a wrong countdown. - Credit-based plans carry total/used/remaining as the endpoint gives them (QuotaWindow.credits, new QuotaCredits type). - A 429 with 1308, 1310 or 1316-1321 marks the window used up and emits QuotaExhausted; the reset time comes from the quota endpoint when known, otherwise from the message (read as UTC+8). 1309, 1313, 1302 and 1305 are not a used-up quota. CONTROL_API_VERSION is now 23: QuotaWindow gained `credits` and the window vocabulary gained `monthly`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WjXngih1kA6oBMoqCkXDVx
fylorn
force-pushed
the
claude/core-glm-quota-7028g0
branch
from
September 25, 2026 16:33
b7cf0b5 to
3edde89
Compare
This was referenced Sep 25, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by F · project thread
What this changes
Upstreams on
api.z.aioropen.bigmodel.cn(written by the Z.ai / BigModel login or added by hand) now show their GLM Coding Plan quota inGET /quota,QuotaSeenandQuotaExhausted: 5-hour, weekly and monthly windows, plus total / used / remaining credits on credit-based plans. A 429 whose body carries a used-up business code marks that window used up immediately.Why
GLM responses carry no quota headers, so the existing header-based sampling never saw anything for these accounts. The only sources are the account's quota endpoint (
/api/monitor/usage/quota/limit, undocumented; shape taken from Z.ai's own open-source tools, logic reimplemented) and the business code in a 429 body.How it works
tw_gateway::glm::Sites::quota_url): by the host ofbase_urlonly; the quota URL is chosen by the matching site. Tests swap the sites for a local fake, likezai::Endpoints.glm::Tracker):GET /quotaasks each due GLM upstream, coalesced to once per minute, and waits at most 5s (slower answers still land viaQuotaSeen). Traffic to a GLM upstream (served, 5xx or 429) asks at most every 5 minutes, in the background. Failures back off 30s / 60s / 120s / 300s. A key with no plan (code: 500, success: false, or no recognizable window) is recorded as "no quota data" (removed from/quota) and re-asked after an hour or when the key changes.glm::parse): success issuccess !== falseandcodein {absent, 0, 200};msgis never used. Auth failures (HTTP 401/403, or code 401/1000/1001) are told apart from other failures. Windows come fromunit/number(3/5 →5h, 6/any →weekly,TIME_LIMIT5/1 →monthly), never from reset order; unrecognized items are dropped. AnextResetTimethat is missing, in the past, or later than the window's length isNone.glm::exhausted,AppState::note_glm_429): code fromerror.type(Anthropic endpoint) orerror.code(OpenAI endpoint). 1308, 1310, 1316–1321 count as used up; 1309, 1313, 1302, 1305 do not. The window comes from the message (English or Chinese), then the code, then the tightest known candidate. The reset time is the quota endpoint's when known, otherwise the one in the message, parsed as UTC+8. The body read is capped (16 KiB, 2s) so failover isn't held up.Contract
tw_api::QuotaCredits { total, used, remaining }andQuotaWindow.credits: Option<QuotaCredits>(also ontw_gateway::quota::Window).remainingis passed through as given, not computed (the real example has 2000 − 23 ≠ 1976).monthly.CONTROL_API_VERSION22 → 23. Lite needs to regeneratesrc/generated/tw-api.ts.msg!sentences and no config changes, so the msg-code table and config manual are unchanged.How it was verified
cargo fmt --all -- --check,cargo clippy --workspace --all-targets -- -D warnings(stable 1.98.1),cargo test --workspace: all green../scripts/smoke.sh: 60 passed, 0 failed.tw-gateway/src/glm.rs: the real credit-plan example, old plan V1 (5h + monthly MCP) and V2 (+ weekly withnumber: 7), no plan (Chinese and English msg), auth failures, missing / past / too-latenextResetTime, unnamed windows, every 429 code in both JSON shapes and both languages, and the ask cadence (1-minute coalescing, 5-minute traffic limit, backoff sequence, no-plan hour, key change).state/glm.rsand 4 end-to-end tests intw-gateway/tests/glm_quota.rsagainst a local fake GLM:/quota-style refresh with credits;Authorizationcarries the raw key with noBearer; no-plan is not re-asked; a 1308 429 through the gateway emits oneQuotaExhaustedwith the message's UTC+8 reset time and triggers a quota ask; a 1302 429 changes nothing.Notes for review
success: falseis read as "no plan", becausemsgis the only other signal and it changes with language. A transient 500 is therefore treated as no plan for up to an hour./quota. This can't be done from CI.🤖 Generated with Claude Code
https://claude.ai/code/session_01WjXngih1kA6oBMoqCkXDVx
Generated by Claude Code