Repository navigation
feat(qwen3-8b): first Ascend variant -- Qwen3-8B on one 910B card - #23
Conversation
huskcass
left a comment
There was a problem hiding this comment.
helm values render check
Validated every variant's values: with helm template against the sglang chart — both the pinned version and the latest published (>=0.7.6). Some variants failed to render — see details below. (Checked with the deploy-time modelRoute.nginx.outputConfigMap set, which swiss injects since the catalog schema rejects per-deploy values by design.)
vllm-tp1-910b@ chart >=0.7.6: `Error: values don't meet the specifications of the schema(s) in the following chart(s):
vllm:- at '/lifecycle': additional properties 'forceShutdown' not allowed`
vllm-tp1-910b@ chart 0.7.6: `Error: values don't meet the specifications of the schema(s) in the following chart(s):
vllm:- at '/lifecycle': additional properties 'forceShutdown' not allowed`
chart.version already uses a semver range — no pinning to flag. 👍
|
Good catch — the failure is real and I should have caught it myself. Root cause. I adapted the values from
Fix (8f68ad0): Along the way Verified the way you did, against the real chart rather than the catalog schema: Putting The gap in my own review: I validated against the catalog schema, which types The flags caveat in the PR description still stands: the |
The catalog has had vendor support in the schema from the start (requires.vendor accepts ascend, and the README maps it to huawei.com/Ascend910), but no variant had ever used it. This is the first non-NVIDIA entry. Qwen3-8B because it fits on a single card, which makes it the cheapest way to prove that a non-NVIDIA node can actually serve: one card, one pod, no tensor-parallel topology to get wrong first. The image is upstream vllm-ascend. The digest is the multi-arch index rather than a platform manifest, so a pull still resolves to arm64 -- which is the one that matters, since Atlas 800T A2 is aarch64 and an amd64-only image would not run there at all. Two NVIDIA-specific values from the other qwen variants are deliberately absent: NVIDIA_DISABLE_REQUIRE is a CUDA-image concern, and PYTORCH_ALLOC_CONF configures the CUDA caching allocator, which torch_npu does not use. No Kubernetes resource strings appear in values: requires.vendor is what decides the extended resource, and that mapping belongs to swiss. extraArgs are derived from the other qwen variants and from what vLLM needs for this model, NOT measured on 910B hardware -- see the PR description.
The values were adapted from qwen3.8-27b-fp8, which is an sglang variant, so
they carried lifecycle.forceShutdown -- a key the vllm chart's values.schema.json
does not have. Rendering failed outright:
at '/lifecycle': additional properties 'forceShutdown' not allowed
vllm expresses the same intent as shutdownTimeout, its --shutdown-timeout flag:
at 0 vLLM aborts in-flight requests on SIGTERM, above 0 it drains them under
that ceiling. Modelled on kimi-k3's vllm-tp8-b300, the only other vllm variant
here that configures a lifecycle.
terminationGracePeriodSeconds 3600 -> 1200, which is drainSeconds 600 plus the
shutdown ceiling plus room. 3600 was inherited from a TP8 variant and is far
more than a single-card 8B needs.
Rendered against the vllm chart to confirm, rather than trusting the catalog
schema: it types values as "any object" and so cannot catch a key the chart
rejects. terminationGracePeriodSeconds 1200, preStop present, --shutdown-timeout
120. Putting forceShutdown back reproduces the original failure, so the check
discriminates.
Four of the six extraArgs did not hold up:
--tool-call-parser=hermes guessed from the folklore that Qwen uses the
Hermes format. Every other vllm variant here
pairs same-family parsers (deepseek_v4/deepseek_v4,
kimi_k3/kimi_k3); this one did not, which was the
tell.
--reasoning-parser=qwen3 copied from qwen3.8-27b-fp8 -- an sglang variant,
whose parser names are its own.
--enable-auto-tool-choice pointless without a tool parser.
--max-model-len=32768 no other variant sets it, and the model's
config.json already says 32768.
Upstream main does register a "qwen3" parser that serves both roles, but the
image here is vllm-ascend v0.23.0.post1, and v0.23's vllm/parser holds only
abstract_parser and parser_manager -- the per-model parsers and their registered
names came in a later refactor. A name that does not exist in the running
version fails at startup, so neither parser is worth asserting until someone has
run this image.
Tags narrow to chat for the same reason: the model can reason and call tools,
but this variant configures no parser for either, so the entry should not
advertise them.
Rendered against the vllm chart: args are --enable-prompt-tokens-details and
--gpu-memory-utilization=0.9 on top of what the chart supplies.
8f68ad0 to
e4508fd
Compare
huskcass
left a comment
There was a problem hiding this comment.
helm values render check ✅
Validated every variant's values: with helm template against the sglang chart — both the pinned version and the latest published (>=0.7.6). All variants render cleanly. (Checked with the deploy-time modelRoute.nginx.outputConfigMap set, which swiss injects since the catalog schema rejects per-deploy values by design.)
chart.version already uses a semver range — no pinning to flag. 👍
|
@huskcass review |
huskcass
left a comment
There was a problem hiding this comment.
Reviewed the fixed head (e4508fd) against the actual chart sources, not just the render check.
Lifecycle fix is correct. Verified against each chart's own values.schema.json:
- vllm chart
lifecycleaccepts exactlypreStop,shutdownTimeout,shutdownReserveSeconds,preStopKill— noforceShutdown. - sglang chart
lifecycleacceptspreStop,shutdownReserveSeconds,forceShutdown,preStopKill— noshutdownTimeout.
So the two charts genuinely have different lifecycle vocabularies, and shutdownTimeout: 120 is the right vllm expression of the same intent. The block matches kimi-k3's vllm-tp8-b300 shape; drain 600 vs 900 scaled to a single 8B card is sensible, and preStopKill: true is an allowed vllm-chart key. terminationGracePeriodSeconds: 1200 comfortably covers drain 600 + shutdown 120.
Rest of the values:
model.contextLength: ''is a documented chart idiom — the vllm chart renders it as--max-model-len, and empty means "let vLLM use the model's own". Fine for Qwen3-8B's 32768.chart.version: ">=0.7.6"— semver range, nothing to flag.image.digestis well-formed (sha256:+ 64 hex). Not registry-verified, but no catalog CI verifies digests today, so that's consistent with the current state, not a blocker here.- The parser call is the right conservative one: a wrong
--reasoning-parsername fails the container at startup, so omitting both until someone has actually run this image is correct.tags: [chat]matches the configured capabilities.
Nothing blocking from me. The automated render check (✅ on this head) plus this human pass.
What
Adds
qwen3-8b, the catalog's first non-NVIDIA variant: Qwen3-8B in bf16 ona single Ascend 910B card, served by vLLM's Ascend backend.
Why this model
One card, one pod, no tensor-parallel topology to get wrong — the cheapest way
to prove a non-NVIDIA node can serve at all.
The entry is deliberately minimal
extraArgsis two flags, and both can be justified:No
--reasoning-parseror--tool-call-parser. Upstream main registers aqwen3parser serving both roles, but this image is vllm-ascendv0.23.0.post1, and v0.23'svllm/parser/holds onlyabstract_parser.pyandparser_manager.py— the per-model parsers and their registered names arrivedin a later refactor. A parser name that does not exist in the running version
fails at startup, so neither is worth asserting before someone has run this
image. They can be added in a follow-up version file.
tagsis thereforechatonly. The model can reason and call tools; thisvariant configures no parser for either, so the entry does not advertise them.
Why no Kubernetes resource strings in
valuesrequires.vendoris what decides the extended resource, and the README isexplicit that the mapping (
ascend→huawei.com/Ascend910) lives in swiss.I checked that the vllm chart does not block this: it hardcodes
nvidia.com/gpuwhen renderingmodel.gpus, but withmodel.gpus: ""it emitsno GPU resource at all, leaving the accelerator resource to the caller.
Confirmed by rendering —
nvidia.com/gpu: False. No chart change needed.Why the digest is an index, not a platform manifest
Atlas 800T A2 is aarch64; an amd64-only image would not run there at all.
v0.23.0.post1is multi-arch (linux/amd64+linux/arm64), and pinning theindex digest keeps a pull resolving to the right architecture. It is also the
newest non-rc tag — the alternatives upstream are
v0.26.0rc*,v0.27.1rc1and nightlies.
lifecycle uses vllm's keys
shutdownTimeout: 120, not sglang'sforceShutdown— caught by the helmvalidator, see the thread below.
terminationGracePeriodSecondsis 1200:drainSeconds600 plus the shutdown ceiling plus room.What is verified, and what is not
Verified: the image exists, is multi-arch with arm64, and its digest is
pinned;
Qwen/Qwen3-8Bexists on HF; the values render against the vllm chartat
>=0.7.6with the expected args, grace period and preStop; schema,naming/layout and index build all pass.
Not verified: nothing in here has been started on a 910B card. The Ascend
machine that proved a pod can get a card (
ascend-docker-runtimemounts/dev/davinci*without the container declaring them, and vllm-ascend servedQwen3-8B on one card over a ClusterIP Service) was removed from the test cluster
on 2026-09-30 and is no longer reachable, so this could not be re-checked.
Merging makes the model visible and schedulable. The first real deploy should be
treated as the measurement — and is also when the parser flags can be settled.