Skip to content

feat(llm): protocol-first BYOK onboarding, live model discovery, and adaptive token budget - #3

Open
Mr-Shaw-Yihan wants to merge 2 commits into
he-yufeng:mainfrom
Mr-Shaw-Yihan:feat/provider-registry
Open

Mr-Shaw-Yihan wants to merge 2 commits into
he-yufeng:mainfrom
Mr-Shaw-Yihan:feat/provider-registry

Conversation

@Mr-Shaw-Yihan

Copy link
Copy Markdown

一句话:沿用RepoWiki的feat/provider-registry前端,优化了llm接入

关于提交者

使用中我发现「连接模型」这一步似乎有一段时间没有更新了:只有一个裸 key 输入框、没法指定网关、报错也只有本来就知道答案的人才看得懂。由于我是在 AI 的协助下完成的优化,因此可能会有诸多问题与规则未能理解,且两个项目的pr描述均由ai完成(我不太看得懂这些术语,因此难以审核优化),还请见谅。

对照你的 roadmap(#2)来读这份 PR

你在 #2 里写:没有大重构的计划,要先想清楚「给谁用、解决什么问题」再谈功能扩张。这份 PR 既不是重构,也不是新功能面。它拆掉的是你点名的那类受众——手里攥着一个别人写的仓库、自己不会写代码的产品经理——够到任何价值之前必经的那一关:过 API key 这一关。今天这一步默认读者知道 litellm 前缀是什么,而这恰恰是 CodeABC 声称要服务的人。除了在一个已存在的对话框里加三个字段,没有给 UI 增加任何新面。

如果你希望这份改动等 demo / 定位工作做完之后再谈,说一声我就关掉——完全不介意。

动机:四个能在当前 main 上复现的问题

  1. 根本没有 api_base 这条通道。 stream_llm / call_llm 只收模型串和 key,别的都不收。拿着 OpenAI 兼容网关 key(比如通义 Token Plan 的 sk-sp-…)的用户无处可说「发到哪」,而带 openai/ 前缀的模型会打到 platform.openai.com 然后鉴权失败。
  2. max_tokens 写死 4096,两个调用 helper 都是。输出窗口更大的模型会被静默截断——长解释说到一半断掉,在非技术用户眼里就是产品 bug。
  3. 能力查询跨前缀 MISS。 litellm 认识 openrouter/qwen/qwen3.8-flash,但不认识网关复用的裸名字,于是上下文窗口和成本元数据全空,token 预算无法自适应。
  4. key 对话框既不帮忙也不反馈。 一个输入框,没有发现、没有校验;base URL 填错只吐一句 httpx 原文报错。

方案:协议优先,而不是厂商优先

协议 请求形态 Base URL 约定 模型列表
chat_completions POST {base}/chat/completions 以 /v1 结尾 GET {base}/models
anthropic_messages POST {base}/v1/messages 裸根地址 GET {base}/v1/models——多数网关没有

第二行正是「测试连接」不能做成「探一下 /models」的原因:Anthropic 兼容网关通常根本没有 models 路由(阿里云百炼的 /apps/anthropic 只文档化了 /v1/messages),探它会得到一个 404,看起来像 key 坏了。所以测试只 ping 用户真正选中的那个模型,用 1 token 的请求——这也是唯一对用户有意义的答复。

改动

  • 协议注册表(backend/services/providers.py,新增):每个协议声明 litellm 前缀、models 路径、Base URL 约定与鉴权头;normalize_base_url() 补缺的或剥多余的 /v1(两类 404 一起消灭);resolve_litellm_model();suggest_max_tokens() 按真实输出窗口推导、上下限夹住、不低于历史上的 4096;能力查询跨前缀回退;litellm 在 helper 内惰性导入。
  • 三个发现 / 校验端点(backend/routers/providers.py,新增;models.py 加 ProtocolCheckRequest;app.py 注册):GET /protocols、POST /protocols/check(只 ping 选中的模型,返回掩码 key 与可读诊断)、GET /models(发现结果与能力元数据合并,非对话模型置灰并给原因;discoverable + status_code 区分「这里没有列表」和「你的 key 被拒了」)。它们不产生任何生成调用,所以不套 _enforce_rate_limit——免费额度预算不受影响。
  • 请求头贯穿(backend/routers/analyze.py):_llm_kwargs(request) 读 x-api-protocol / x-api-key / x-api-base / x-model,overview / annotations / qa / edit 四个端点都走它。不带任何头时四个值全为 None,既有的 env / 默认值解析与今天完全一致。
  • 三步表单(frontend/src/components/ApiKeyModal.tsx 重写、lib/protocols.ts 新增、lib/api.ts 发头):端点 → 模型 → 验证,用既有 t(zh, en) 做双语,视觉与现有对话框一致;默认导出签名未变,所以 Home.tsx 一行没改;完成 按钮在连接测试通过前保持禁用。后端持有的四个串(格式名、线路说明、模型置灰原因、连接诊断)现在像 activity.py / contributing.py 那样同时给英文与 *_zh,因此会跟语言切换走。重写时保留了原对话框对用户的三处承诺:key 只在本地、OpenRouter 注册链接、不填 key 也有每天 20 次免费额度。
  • 文档(README.md、README_CN.md、.env.example):BYOK 一节改为走三步流程,并仍然写明 OpenRouter key 只需要 key 本身;.env.example 文档化 OPENAI_API_BASE,即表单 Base URL 字段的服务端对应物。

我刻意没有动的不变量

  • backend/services/llm.py 保留顶层 eager import litellm 与 drop_params——test_llm.py 锁着;新注册表改为惰性导入 litellm,而不是去改它。
  • litellm 1.98.0–1.100.0 在 Python 3.10 上的排除没有碰,这里也没有新增会踩到它的导入。
  • OpenRouter key 形态自动识别逐字保留,现在由四个测试覆盖(sk-or- 无模型、带前缀模型、裸模型、以及走 CODEABC_MODEL)。新的 protocol 参数只在表单提供时才生效。
  • 免费额度限流、无 key 时的确定性地图、扫描 / 跳过规则均未变。

向后兼容是纯增量的:没有删除或重命名任何端点、环境变量、请求头、函数签名或请求字段。只发 x-api-key 的客户端行为与今天的 main 完全一致。

安全

所有响应里 key 均脱敏(sk-sp-***T123),错误串过 redact_secret();失败时日志行与返回给读者的字符串共用同一份脱敏结果——厂商的错误文本可能把 key 原样引回来,只清返回值不够。测试断言响应体(含传输错误路径)中不出现明文 key。key 仍只存在浏览器 localStorage,与文档一致。

测试

本地在 93a797e 全绿,逐条对齐 CI lane:

CI lane 本地命令 结果
backend ruff check . All checks passed
backend python -m compileall -q backend / python -c "from backend.app import app" exit 0 / Import OK
backend pytest -q 809 passed(原 768;+41 覆盖注册表、端点与请求头路径)
frontend npm run lint / npm run build(tsc -b && vite build) exit 0 / 绿,442ms

新增 / 重写 tests/test_providers.py(注册表、归一化、能力 helper)、tests/test_providers_router.py(三个端点跑在 mock 的 httpx 上,含 key 脱敏纪律),并扩展 tests/test_llm.py 覆盖 BYOK 头、协议解析与保留的 OpenRouter 路由。

吃自己的狗粮: 用 CodeABC 自家的四个扫描器扫本 PR 新增 / 触碰的六个后端文件,long_functions 0 flag(最长 54 行 vs 阈值 60)、too_many_params 0 flag(最宽 5 参 vs 阈值 6);docstring / typing 报告里点名的全是既有上游符号(models.py 的 pydantic 类、app.py 的 lifespan / health),新增的 providers.py 与 llm.py 没有出现在任何缺失清单里。这些数字背后还有一次真实冒烟:一个真实的 Anthropic 兼容网关(无 /v1/models)完整走过「无列表 → 手填模型 → 只 ping 选中模型」,全程 key 脱敏。本地为 Windows / Python 3.12.10,3.10–3.12 × ubuntu/windows 矩阵交给 CI。
image

…adaptive token budget

Describe gateways by the wire protocol they speak (OpenAI-compatible chat
completions vs Anthropic-compatible messages) instead of a vendor brand, add
three discovery/validation endpoints (list protocols, ping the selected model
before saving, discover models merged with litellm capability metadata), and
rewrite the API-key modal into a linear three-step form (endpoint -> model ->
verify).

Fixes concrete integration problems:
- stream_llm/call_llm had no api_base channel, so an OpenAI-compatible gateway
  (e.g. a Qwen Token Plan key, sk-sp-) was routed to platform.openai.com; the
  form's base URL + protocol + model now flow through x-api-base /
  x-api-protocol / x-model headers into analyze._llm_kwargs and on to litellm,
  with the base URL normalised per protocol (/v1 suffix vs bare root).
- Anthropic-compatible gateways often expose no /v1/models (e.g. Alibaba
  Bailian's /apps/anthropic), so connection checks ping the selected model with
  a 1-token request and the form falls back to manual model entry.
- max_tokens was hard-coded to 4096 in both call helpers, truncating models
  with larger output windows; it now derives from the model's real output window
  via suggest_max_tokens (clamped, floored at 4096).
- capability lookups missed when a gateway reuses a model name litellm only
  knows under another prefix; added a cross-prefix catalog fallback.

Everything the form renders from a backend payload (protocol names, wire
descriptions, grey-out reasons, connection diagnosis) ships as `label` plus a
`_zh` twin, following the convention activity/contributing already use, so the
language toggle reaches it. The rewritten modal also keeps the three promises
the previous one made: the key stays in local storage, the OpenRouter signup
link, and the free tier of 20 requests/day.

llm.py keeps its eager `import litellm` + drop_params invariant (locked by
test_llm.py); the registry imports litellm lazily. API keys are masked in every
response and redacted from error strings and log lines alike. Fully backward
compatible: without the new headers, behavior is unchanged.

Tests: 809 passed (test_providers, test_providers_router rewritten for the
protocol design; test_llm extended). ruff clean; new backend files are 0-flag
under the project's own long_functions/too_many_params/docstrings/typing_coverage
scanners. Frontend tsc + eslint + build green.
README (EN + CN) now spell out the endpoint -> model -> verify flow: Base URL
normalisation per API format, model discovery with a manual fallback for
gateways that expose no list, and a connection test that pings only the selected
model. Both still note that an OpenRouter key needs nothing but the key.
.env.example documents OPENAI_API_BASE, the server-side counterpart of the form's
Base URL field.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant