Skip to content

feat(qwen3-8b): first Ascend variant -- Qwen3-8B on one 910B card - #23

Merged
aceforeverd merged 3 commits into
masterfrom
feat/ascend-qwen3-8b
Oct 9, 2026
Merged

aceforeverd merged 3 commits into
masterfrom
feat/ascend-qwen3-8b

Conversation

@zhanghaohit

@zhanghaohit zhanghaohit commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

What

Adds qwen3-8b, the catalog's first non-NVIDIA variant: Qwen3-8B in bf16 on
a single Ascend 910B card, served by vLLM's Ascend backend.

requires:
  gpus: 1
  topology: single-node
  vendor: ascend
image:
  repository: quay.io/ascend/vllm-ascend
  tag: v0.23.0.post1
  digest: sha256:ffe9306186781ddb80812eda574134e601ee8499693ea36e4a97173111ddf7e7

Why this model

One card, one pod, no tensor-parallel topology to get wrong — the cheapest way
to prove a non-NVIDIA node can serve at all.

The entry is deliberately minimal

extraArgs is two flags, and both can be justified:

--enable-prompt-tokens-details    without it cached_tokens is always null
--gpu-memory-utilization=0.9      every variant in this catalog states it

No --reasoning-parser or --tool-call-parser. Upstream main registers a
qwen3 parser serving both roles, but this image is vllm-ascend
v0.23.0.post1, and v0.23's vllm/parser/ holds only abstract_parser.py and
parser_manager.py — the per-model parsers and their registered names arrived
in a later refactor. A parser name that does not exist in the running version
fails at startup, so neither is worth asserting before someone has run this
image. They can be added in a follow-up version file.

tags is therefore chat only. The model can reason and call tools; this
variant configures no parser for either, so the entry does not advertise them.

Why no Kubernetes resource strings in values

requires.vendor is what decides the extended resource, and the README is
explicit that the mapping (ascend → huawei.com/Ascend910) lives in swiss.

I checked that the vllm chart does not block this: it hardcodes
nvidia.com/gpu when rendering model.gpus, but with model.gpus: "" it emits
no GPU resource at all, leaving the accelerator resource to the caller.
Confirmed by rendering — nvidia.com/gpu: False. No chart change needed.

Why the digest is an index, not a platform manifest

Atlas 800T A2 is aarch64; an amd64-only image would not run there at all.
v0.23.0.post1 is multi-arch (linux/amd64 + linux/arm64), and pinning the
index digest keeps a pull resolving to the right architecture. It is also the
newest non-rc tag — the alternatives upstream are v0.26.0rc*, v0.27.1rc1
and nightlies.

lifecycle uses vllm's keys

shutdownTimeout: 120, not sglang's forceShutdown — caught by the helm
validator, see the thread below. terminationGracePeriodSeconds is 1200:
drainSeconds 600 plus the shutdown ceiling plus room.

What is verified, and what is not

Verified: the image exists, is multi-arch with arm64, and its digest is
pinned; Qwen/Qwen3-8B exists on HF; the values render against the vllm chart
at >=0.7.6 with the expected args, grace period and preStop; schema,
naming/layout and index build all pass.

Not verified: nothing in here has been started on a 910B card. The Ascend
machine that proved a pod can get a card (ascend-docker-runtime mounts
/dev/davinci* without the container declaring them, and vllm-ascend served
Qwen3-8B on one card over a ClusterIP Service) was removed from the test cluster
on 2026-09-30 and is no longer reachable, so this could not be re-checked.

Merging makes the model visible and schedulable. The first real deploy should be
treated as the measurement — and is also when the parser flags can be settled.

@huskcass huskcass left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

helm values render check ⚠️

Validated every variant's values: with helm template against the sglang chart — both the pinned version and the latest published (>=0.7.6). Some variants failed to render — see details below. (Checked with the deploy-time modelRoute.nginx.outputConfigMap set, which swiss injects since the catalog schema rejects per-deploy values by design.)

  • vllm-tp1-910b @ chart >=0.7.6: `Error: values don't meet the specifications of the schema(s) in the following chart(s):
    vllm:
  • at '/lifecycle': additional properties 'forceShutdown' not allowed`
  • vllm-tp1-910b @ chart 0.7.6: `Error: values don't meet the specifications of the schema(s) in the following chart(s):
    vllm:
  • at '/lifecycle': additional properties 'forceShutdown' not allowed`

chart.version already uses a semver range — no pinning to flag. 👍

@zhanghaohit

Copy link
Copy Markdown
Contributor Author

Good catch — the failure is real and I should have caught it myself.

Root cause. I adapted the values from qwen3.8-27b-fp8, which is an sglang variant, onto a vllm chart. The two charts have different lifecycle schemas:

key vllm sglang
forceShutdown ✗ ✓
shutdownTimeout ✓ ✗
preStop / preStopKill / shutdownReserveSeconds ✓ ✓

Fix (8f68ad0): forceShutdown: true → shutdownTimeout: 120, which is how vllm expresses the same intent — vLLM aborts in-flight requests on SIGTERM at 0 and drains them under that ceiling above 0. Modelled on kimi-k3/vllm-tp8-b300, the only other vllm variant here that configures a lifecycle. Also added the preStop.enabled / endpointSyncSeconds keys it uses.

Along the way terminationGracePeriodSeconds went 3600 → 1200: 3600 came in with the TP8 variant I copied and is far more than drainSeconds: 600 plus the shutdown ceiling needs on a single-card 8B.

Verified the way you did, against the real chart rather than the catalog schema:

terminationGracePeriodSeconds: 1200
preStop: yes
--shutdown-timeout: 120
resources: ephemeral-storage only, no nvidia.com/gpu

Putting forceShutdown back reproduces the original error, so the check discriminates rather than just passing.

The gap in my own review: I validated against the catalog schema, which types values as "any object" and therefore cannot catch a key the chart rejects. I did render the chart earlier — but with a hand-written probe file, not with the variant's final values. Rendering the actual values against the actual chart is the only check that would have caught this, and it is exactly what this bot does.

The flags caveat in the PR description still stands: the extraArgs are derived, not measured on 910B hardware.

The catalog has had vendor support in the schema from the start (requires.vendor
accepts ascend, and the README maps it to huawei.com/Ascend910), but no variant
had ever used it. This is the first non-NVIDIA entry.

Qwen3-8B because it fits on a single card, which makes it the cheapest way to
prove that a non-NVIDIA node can actually serve: one card, one pod, no
tensor-parallel topology to get wrong first.

The image is upstream vllm-ascend. The digest is the multi-arch index rather
than a platform manifest, so a pull still resolves to arm64 -- which is the one
that matters, since Atlas 800T A2 is aarch64 and an amd64-only image would not
run there at all.

Two NVIDIA-specific values from the other qwen variants are deliberately absent:
NVIDIA_DISABLE_REQUIRE is a CUDA-image concern, and PYTORCH_ALLOC_CONF
configures the CUDA caching allocator, which torch_npu does not use.

No Kubernetes resource strings appear in values: requires.vendor is what
decides the extended resource, and that mapping belongs to swiss.

extraArgs are derived from the other qwen variants and from what vLLM needs for
this model, NOT measured on 910B hardware -- see the PR description.
The values were adapted from qwen3.8-27b-fp8, which is an sglang variant, so
they carried lifecycle.forceShutdown -- a key the vllm chart's values.schema.json
does not have. Rendering failed outright:

    at '/lifecycle': additional properties 'forceShutdown' not allowed

vllm expresses the same intent as shutdownTimeout, its --shutdown-timeout flag:
at 0 vLLM aborts in-flight requests on SIGTERM, above 0 it drains them under
that ceiling. Modelled on kimi-k3's vllm-tp8-b300, the only other vllm variant
here that configures a lifecycle.

terminationGracePeriodSeconds 3600 -> 1200, which is drainSeconds 600 plus the
shutdown ceiling plus room. 3600 was inherited from a TP8 variant and is far
more than a single-card 8B needs.

Rendered against the vllm chart to confirm, rather than trusting the catalog
schema: it types values as "any object" and so cannot catch a key the chart
rejects. terminationGracePeriodSeconds 1200, preStop present, --shutdown-timeout
120. Putting forceShutdown back reproduces the original failure, so the check
discriminates.
Four of the six extraArgs did not hold up:

  --tool-call-parser=hermes   guessed from the folklore that Qwen uses the
                              Hermes format. Every other vllm variant here
                              pairs same-family parsers (deepseek_v4/deepseek_v4,
                              kimi_k3/kimi_k3); this one did not, which was the
                              tell.
  --reasoning-parser=qwen3    copied from qwen3.8-27b-fp8 -- an sglang variant,
                              whose parser names are its own.
  --enable-auto-tool-choice   pointless without a tool parser.
  --max-model-len=32768       no other variant sets it, and the model's
                              config.json already says 32768.

Upstream main does register a "qwen3" parser that serves both roles, but the
image here is vllm-ascend v0.23.0.post1, and v0.23's vllm/parser holds only
abstract_parser and parser_manager -- the per-model parsers and their registered
names came in a later refactor. A name that does not exist in the running
version fails at startup, so neither parser is worth asserting until someone has
run this image.

Tags narrow to chat for the same reason: the model can reason and call tools,
but this variant configures no parser for either, so the entry should not
advertise them.

Rendered against the vllm chart: args are --enable-prompt-tokens-details and
--gpu-memory-utilization=0.9 on top of what the chart supplies.
@zhanghaohit
zhanghaohit force-pushed the feat/ascend-qwen3-8b branch from 8f68ad0 to e4508fd Compare October 9, 2026 01:59

@huskcass huskcass left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

helm values render check ✅

Validated every variant's values: with helm template against the sglang chart — both the pinned version and the latest published (>=0.7.6). All variants render cleanly. (Checked with the deploy-time modelRoute.nginx.outputConfigMap set, which swiss injects since the catalog schema rejects per-deploy values by design.)

chart.version already uses a semver range — no pinning to flag. 👍

@zhanghaohit

Copy link
Copy Markdown
Contributor Author

@huskcass review

@huskcass huskcass left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the fixed head (e4508fd) against the actual chart sources, not just the render check.

Lifecycle fix is correct. Verified against each chart's own values.schema.json:

  • vllm chart lifecycle accepts exactly preStop, shutdownTimeout, shutdownReserveSeconds, preStopKill — no forceShutdown.
  • sglang chart lifecycle accepts preStop, shutdownReserveSeconds, forceShutdown, preStopKill — no shutdownTimeout.

So the two charts genuinely have different lifecycle vocabularies, and shutdownTimeout: 120 is the right vllm expression of the same intent. The block matches kimi-k3's vllm-tp8-b300 shape; drain 600 vs 900 scaled to a single 8B card is sensible, and preStopKill: true is an allowed vllm-chart key. terminationGracePeriodSeconds: 1200 comfortably covers drain 600 + shutdown 120.

Rest of the values:

  • model.contextLength: '' is a documented chart idiom — the vllm chart renders it as --max-model-len, and empty means "let vLLM use the model's own". Fine for Qwen3-8B's 32768.
  • chart.version: ">=0.7.6" — semver range, nothing to flag.
  • image.digest is well-formed (sha256: + 64 hex). Not registry-verified, but no catalog CI verifies digests today, so that's consistent with the current state, not a blocker here.
  • The parser call is the right conservative one: a wrong --reasoning-parser name fails the container at startup, so omitting both until someone has actually run this image is correct. tags: [chat] matches the configured capabilities.

Nothing blocking from me. The automated render check (✅ on this head) plus this human pass.

@aceforeverd
aceforeverd merged commit e502e5f into master Oct 9, 2026
4 checks passed
@aceforeverd
aceforeverd deleted the feat/ascend-qwen3-8b branch October 9, 2026 05:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants