Skip to content

feat(lambda): adopt GA awslambdaric 4.1.0, drop the preview variants - #6809

Merged
Eren-Jeager123 merged 1 commit into
lambdafrom
feat/lambda-ga-ric-4.1.0
Sep 29, 2026
Merged

Eren-Jeager123 merged 1 commit into
lambdafrom
feat/lambda-ga-ric-4.1.0

Conversation

@Eren-Jeager123

@Eren-Jeager123 Eren-Jeager123 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Why

awslambdaric 4.1.0 ships @register_pre_fork in the public RIC (upstream PR #226), so the separate preview image line is no longer needed. 10 images become 5.

Changes

  • RIC from PyPI — awslambdaric 4.0.0 → 4.1.0 in the three dependency closures (base, cupy, pytorch) + relock. 4.1.0 publishes manylinux/musllinux wheels with a prebuilt runtime_client.so for cp310–cp315, so nothing compiles: the RIC builder stage, its gcc/cmake/autotools install, and the pre-build tarball fetch are all removed. sglang and vllm inherit the RIC from the pytorch builder, so no change is needed there.

  • Preview images removed — the 5 *-preview configs, ARG AWSLAMBDARIC_VERSION, the builder-ric-preview stage and the 5 *-preview-py3 targets.

  • Scope: process workers + pre_fork only — the GA RIC has no AWS_LAMBDA_CONCURRENCY_MODE; multi-concurrency is entered by AWS_LAMBDA_MAX_CONCURRENCY alone and workers are forked processes, so thread and hybrid go away with the preview images. That removes the run-ric-topology-test input, the lambdaric-topology-test job, the RIE-topology and mode-override validation steps, rie_topology_check.sh, and test/lambda/platform/ (whose only entry point was the topology runner).

  • Handlers simplified — both serving handlers drop the try/except ImportError and mode check:

    if os.environ.get("AWS_LAMBDA_MAX_CONCURRENCY"):
        register_pre_fork(_start_server)
    else:
        _start_server()

Testing

test_lambda_runtime.py drops its concurrency-mode tests — the knob they asserted no longer exists — and gains test_pre_fork_hook_available, asserting register_pre_fork is importable. That is now a contract of every image rather than of the preview variants, so it runs in all 5 pipelines.

One caveat: this also removes the end-to-end pre-fork coverage (hook runs once in the parent, N workers fork, all proxy to a single engine). Happy to re-add a trimmed process-only local check in a follow-up if that signal is wanted.

register_pre_fork ships in the public RIC as of 4.1.0, so the separate
preview images exist for nothing. Pin awslambdaric==4.1.0 in the base,
cupy and pytorch closures and install it from PyPI — 4.1.0 publishes
manylinux/musllinux wheels for cp310-cp315 with a prebuilt
runtime_client.so, so the RIC builder stage and its gcc/cmake/autotools
install go away along with the S3 tarball fetch in pre_build.sh.

Removes the 5 *-preview configs, the builder-ric-preview stage, the 5
*-preview-py3 targets and ARG AWSLAMBDARIC_VERSION.

The GA RIC has no AWS_LAMBDA_CONCURRENCY_MODE: multi-concurrency is
entered by AWS_LAMBDA_MAX_CONCURRENCY alone and workers are forked
processes, so thread and hybrid go with the preview images. That drops
the run-ric-topology-test input, the lambdaric-topology-test job, the RIE
topology + mode-override validation steps, rie_topology_check.sh and
test/lambda/platform/ (run_topology.py, test_handler.py, Dockerfile.test).

Both serving handlers lose the try/except ImportError and mode check:
register_pre_fork when AWS_LAMBDA_MAX_CONCURRENCY is set, module-level
otherwise. test_lambda_runtime.py drops its concurrency-mode tests and
asserts register_pre_fork is importable in every image instead.
@Eren-Jeager123
Eren-Jeager123 merged commit 3963ec8 into lambda Sep 29, 2026
94 of 95 checks passed
@Eren-Jeager123
Eren-Jeager123 deleted the feat/lambda-ga-ric-4.1.0 branch September 29, 2026 23:31
Eren-Jeager123 added a commit that referenced this pull request Oct 2, 2026
The try/except around the lambda_concurrency_hooks import let the probe degrade
silently on a pre-4.1.0 RIC, reporting a failed hook assertion instead of the real
cause. 4.1.0 is pinned in every image closure, so an older RIC is a defect to fail
on, not to tolerate — #6809 removed the same guard from the serving handlers.

Importing register_pre_fork unconditionally now surfaces a pre-4.1.0 RIC as a
Runtime.ImportModuleError at init. The baked-handler import is gated on the file
existing instead of a broad except, so core images (which ship no handler.py) are
still handled while an engine handler that fails to import now propagates.

wait_ready greps the container logs for the init error before tailing, since RIE
chatter otherwise pushes it out of the tail window.

Verified on a scratch image pinned to awslambdaric 4.0.0: both shapes fail, exit 1,
and the output names the cause —
  Unable to import module 'ric_probe_handler': No module named
  'awslambdaric.lambda_concurrency_hooks'
4.1.0 image still passes 9/9.

Signed-off-by: Kevin Wang <kwanggg@amazon.com>
Eren-Jeager123 added a commit that referenced this pull request Oct 7, 2026
* feat: add Lambda GPU DLC images (base, cupy, pytorch)

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* rename workflows

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add tmp torch as torch home

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add oss install

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix oss install link

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* bump uvlock

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix cve

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* convert to dispatch workflow for release

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* change image path

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* release to gamma

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* release torch_home to prod

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add s3 torch connector for lambda pytorch container (#6211)

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add lambda ric patch package (#6217)

* add lambda ric patch

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* run single line copy

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* deploy gamma with new RIC

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* revert to prod deployment

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

---------

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* merge from main

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* upgrade lambda ric version (#6402)

* upgrade lambda ric version

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix paths

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add additional tests

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* update pytorch

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

---------

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* remove customer type for lambda

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add cupy type

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* release to prod

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* change lambda docker path

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* feat: Lambda SGLang/vLLM serving engine (#6455)

* re add sglang

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add vllm

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add vllm cache

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add timeout

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix telemetry

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

---------

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* release sglang and vllm lambda

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* Add lambda ric test for LMI platform (#6585)

* dockerfmt

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add lambda ric test

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix json jq

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* use jq

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add warmup

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix vllm test

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix fasterlock

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

---------

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* disable ric test

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add lambda vllm hf

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* bump hydra-core

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix handlers

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* update sglang and vllm (#6648)

* update sglang and vllm

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add allowlist

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

---------

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* add s3 streaming for lambda inference engine (#6665)

* add s3 streaming for lambda inference engine

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* wire s3 test

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* forward credentials

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* server without rie

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix curl

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

---------

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>

* fix(lambda): report the real framework version, bump vLLM to 0.30.0, add modelscope (#6764)

* fix(lambda): report the real framework version for pytorch and vllm images

pytorch configs claimed 2.11.0 while pyproject pins torch==2.13.0, and the
vllm configs carried the CUDA version (13.0.3) instead of vllm 0.27.1. The
value is baked into bash_telemetry.sh at build time, so usage metrics for
these four images were labelled with the wrong version.

Release tags are unaffected: the lambda entry in DLContainersReleaseLogicPython
frameworks.yml renders release_tag as "{container_type}" and never references
{version}, and its version_pattern is '.*'.

Signed-off-by: Kevin Wang <kwanggg@amazon.com>

* fix(lambda): bump transformers to 5.10.4 for CVE-2026-9856

Path traversal via save_pretrained() in transformers <=5.8.0.dev0, fixed
in 5.10.0. 5.10.0 itself is yanked upstream, so pin 5.10.4.

* fix(lambda): bump vllm to 0.29.0 for CVE-2026-90553

RCE in the LlavaOnevision2 processor loader (ignores trust_remote_code),
fixed in 0.28.0; take 0.29.0 as current. flashinfer moves to 0.6.18 to
match 0.29.0's requirements/cuda.txt. The torch trio pin is unchanged
(2.13.0 / 2.11.0 / 0.28.0), so torch-constraints.txt still applies, and
the build-system requires are identical between the two refs.

* feat(lambda): bump vllm to 0.30.0, align transformers, add modelscope

vllm 0.30.0 (ref ced6857a) + flashinfer 0.6.18.post1 per its cuda.txt.

vllm asks for an unbounded transformers>=5.10.4 while its upstream test
suite hard-pins transformers==5.16.1, so the test env was swapping out
the transformers the image ships. Pin 5.16.1 in the pytorch pyproject
(the base every engine builder inherits) and add it to the vllm
constraint file so the runtime resolve can't drift. safetensors moves
0.8.0rc1 -> 0.8.0 because transformers 5.16.1 needs the final release.

modelscope<1.38 gives the vllm image ModelScope hub loading, matching
the al2023 image's serving extras; sglang already gets it via
sglang[all]. The cap is vllm#47325.

The spec-decode examples now pass --gpu-memory-utilization 0.95: at the
0.9 default vllm 0.29+ leaves ~0.09 GiB for KV on the 1xL4 runner, which
OOM'd the eagle case at max-model-len 2048.

* fix(lambda): align vllm ref and transformers with main's AL2023 image

Use the same vllm_ref as main's AL2023 vllm config (ec4a3a53, upstream's
'[CI] Bump Transformers version to 5.17.0') instead of the v0.30.0 tag,
so the upstream test suite checks out the ref the image is built from and
its transformers==5.17.0 test pin matches what the image ships.

Revert the hand-edit of the shared vllm_ec2_examples_test.sh and take
main's copy, which already carries the gpu-memory-utilization fix.

* fix(lambda): drop redundant transformers constraint

The pytorch pyproject pin already fixes the version the vllm image ships:
the vllm builder inherits /var/lang from the pytorch builder and uv leaves
an already-satisfied requirement alone.

* fix(lambda): install DeepGEMM/DeepSelect build deps, floor anyio for CVE-2026-63374

The vllm wheel build failed compiling deepgemm_C: DeepGEMM's vendored
deep_jit needs elfutils/libdwfl.h. Install elfutils-devel and gcc14 and
build with gcc14, matching docker/vllm/Dockerfile.amzn2023, which builds
the same vllm ref green.

anyio 4.13.0 (transitive via httpx) is CVE-2026-63374 CRITICAL, failing
the pytorch and sglang scans; floor it to 4.14.2.

* fix(lambda): drop added comments

---------

Signed-off-by: Kevin Wang <kwanggg@amazon.com>

* fix(lambda): point runtime caches at /tmp in base, cupy and pytorch images (#6791)

* fix(lambda): point runtime caches at /tmp in base, cupy and pytorch images

Lambda mounts everything except /tmp read-only, but only the sglang and
vllm stages redirected their caches there. CuPy JIT-compiles kernels and
writes them to $HOME/.cupy by default, so a customer's first custom
kernel fails on a read-only HOME; numba, huggingface_hub, triton and
inductor have the same problem in the cupy and pytorch images.

Sets HOME, USER and the XDG dirs on all three stages, CUPY_CACHE_DIR and
NUMBA_CACHE_DIR on cupy, and HF_HOME, TRITON_CACHE_DIR and
TORCHINDUCTOR_CACHE_DIR on pytorch, matching what the engine stages
already do. Preview stages inherit via FROM.

Adds env assertions plus a cupy GPU test that compiles a RawKernel and
checks the cache landed under /tmp — the case that was reported.

* fix(lambda): source telemetry from $HOME/.bashrc in base, cupy and pytorch

Setting HOME=/tmp moved the interactive-shell rc file off /root/.bashrc,
which is where the telemetry hook was written — amzn2023 bash only reads
/etc/bashrc via ~/.bashrc, so the hook stopped firing and four telemetry
tests failed. Write the hook to $HOME/.bashrc too, as the sglang and vllm
stages already do.

* fix(lambda): ship libNVVM in the cupy image so numba.cuda works (#6807)

The cupy image advertises Numba, but numba.cuda JIT-compiles kernels through
libNVVM, which the CUDA runtime base omits: cuda.is_available() returned False
and every @cuda.jit raised NvvmSupportError. CuPy is unaffected because it
compiles via NVRTC, which the runtime base does ship.

Copies nvvm/lib64 and nvvm/libdevice from the devel base and points CUDA_HOME
at them. nvvm/bin (cicc, 73 MB) is left out — numba loads libnvvm.so directly
and never invokes the nvcc driver, so the image grows by ~62 MB rather than 134.

Covered by a GPU-free compile_ptx unit test and a @cuda.jit launch test, both
verified to fail against an image without libNVVM.

* feat(lambda): adopt GA awslambdaric 4.1.0, drop the preview variants (#6809)

register_pre_fork ships in the public RIC as of 4.1.0, so the separate
preview images exist for nothing. Pin awslambdaric==4.1.0 in the base,
cupy and pytorch closures and install it from PyPI — 4.1.0 publishes
manylinux/musllinux wheels for cp310-cp315 with a prebuilt
runtime_client.so, so the RIC builder stage and its gcc/cmake/autotools
install go away along with the S3 tarball fetch in pre_build.sh.

Removes the 5 *-preview configs, the builder-ric-preview stage, the 5
*-preview-py3 targets and ARG AWSLAMBDARIC_VERSION.

The GA RIC has no AWS_LAMBDA_CONCURRENCY_MODE: multi-concurrency is
entered by AWS_LAMBDA_MAX_CONCURRENCY alone and workers are forked
processes, so thread and hybrid go with the preview images. That drops
the run-ric-topology-test input, the lambdaric-topology-test job, the RIE
topology + mode-override validation steps, rie_topology_check.sh and
test/lambda/platform/ (run_topology.py, test_handler.py, Dockerfile.test).

Both serving handlers lose the try/except ImportError and mode check:
register_pre_fork when AWS_LAMBDA_MAX_CONCURRENCY is set, module-level
otherwise. test_lambda_runtime.py drops its concurrency-mode tests and
asserts register_pre_fork is importable in every image instead.

* chore(lambda): add checksum and signature verification for third-party artifacts (#6825)

Implements the checksum/signature rows of the dependency verifiability
inventory. Each change is annotated in the Dockerfile with its row number.

Row 10 - FFmpeg: sha256 + GPG signature. Source moved from the GitHub archive
  tag to ffmpeg.org, which is the only publisher of the two that ships a
  detached signature; GitHub's archive tarballs are generated per request and
  are not byte-stable. Signing public key vendored at docker/lambda/keys/.
Row 11 - rustup: fetch rustup-init and verify its published sha256 instead of
  piping sh.rustup.rs to a shell.
Row 12 - sccache: verify the published sha256.
Row 16 - nvidia/cuda runtime and devel: digest-pinned (13 FROM lines).
Row 17 - uv: digest-pinned once as a uv-source stage, since COPY --from= does
  not expand ARGs.

Not included, and why:
Rows 4, 5, 7, 8 - hash enforcement for the sglang and vllm Python closures.
  Those stages have no lockfile, so this needs uv.lock + --require-hashes,
  exact pins in place of the >= constraints, and per-version hashes for
  sglang-kernel and flashinfer-cubin carried in the image configs. Needs build
  iteration to settle known resolution conflicts.
Row 9 - git commit signature verification. Upstream commits are GPG-verified,
  but enforcing it in the build requires a decision on whose keys to trust.
Rows 6, 13, 14 - no checksum or signature published at source.

* test(security): refresh past-due lambda ECR scan allowlist review dates (#6832)

All 19 dated entries hit their 2026-10-01 review date, failing every lambda
ecr-vulnerability-scan job. Bumped to 2026-12-31, and corrected two reasons
that no longer hold:

- the 13 aws-lambda-rie Go stdlib / go-chi entries now cite the concrete
  state rather than 'no new RIE release yet': latest is v1.37 (2026-08-31),
  still built with Go 1.25.8, and we download releases/latest so they
  auto-resolve when a newer RIE ships.
- CVE-2025-23308 / CVE-2025-23339 claimed the fix 'requires CUDA 13.x base
  image which is not yet available'. These images are on CUDA 13.0.3, so that
  premise is obsolete; the standing justification is that neither is
  exploitable at Lambda runtime (both need nvdisasm/cuobjdump run against a
  malicious ELF).

No entries added or removed; only review_by and reason changed.

* chore(lambda): point CI at main ahead of the branch merge

PR workflows triggered on branches: [lambda], so a PR targeting main matched
nothing and no lambda CI ran (branches: filters the base branch). Repointed
all three to [main], matching every other framework on main.

Fixes four path filters that silently matched nothing, in both the
on.pull_request.paths and the paths-filter blocks:
- scripts/lambda/** does not exist; core's entrypoint is
  scripts/docker/lambda/lambda_entrypoint.sh
- scripts/common/** and scripts/telemetry/** are scripts/docker/common/**
  and scripts/docker/telemetry/**
- the engine workflows watched scripts/docker/lambda/<engine>/** but the
  entrypoints sit beside that dir, so {vllm,sglang}_entrypoint.sh changes
  never triggered a build

release-gate no longer accepts refs/heads/lambda. Release stays dispatch-only;
no autorelease workflow, so the release-schedule prcheck does not apply.

* fix(lambda): bump urllib3 to 2.8.0 for CVE-2026-97687 and CVE-2026-97689

Two HIGH CVEs in urllib3 2.7.0, failing the ecr-vulnerability-scan on all
three core images: proxy/TLS options silently ignored on some code paths
(CVE-2026-97687) and unbounded memory in the chunked streaming parser
(CVE-2026-97689). Both fixed in 2.8.0.

urllib3 is pinned explicitly in base, cupy and pytorch, so all three
closures move; lock churn is urllib3 only. sglang and vllm inherit from the
pytorch closure.

* fix(lambda): bind torch_cuda_arch_list from image config

The vllm builder declared `ARG TORCH_CUDA_ARCH_LIST`, but
resolve_build_args.py keeps `torch_cuda_arch_list` lower-case
(PRESERVE_CASE_KEYS) so the config value bound to nothing and the
Dockerfile default was what compiled. Switch to the lower-case
ARG + upper-case ENV pattern used by docker/vllm/Dockerfile.amzn2023.

Also drop the duplicated arch-list fallbacks in the lambda build hooks:
the value feeds the wheel cache key, so a local default diverging from
the config would hand out a mismatched wheel.

* test(lambda): smoke-test entrypoint dispatch and handler concurrency mode

Adds test/lambda/unit/test_entrypoint.py (all images) covering the
RIE-vs-exec branch, exec semantics, and argv forwarding, and
test/lambda/unit/test_handler_concurrency.py (engine images) covering the
AWS_LAMBDA_MAX_CONCURRENCY dispatch, server startup failure, and payload
normalisation. Both run CPU-only with no engine launched.

Also fail fast when the engine server dies during startup: the Popen handle
was discarded, so a bad MODEL_ID, OOM or bad --load-format polled /health for
the full timeout before raising, with the real error only in stderr.

And narrow the sglang/vllm runtime stages to `dnf upgrade --security`,
matching base/cupy/pytorch and the repo-wide convention; an unqualified
upgrade is only used for named packages elsewhere.

* fix(lambda): floor fsspec to 2026.6.0 for CVE-2026-104851

HIGH, ReferenceFileSystem evaluates Kerchunk reference fields unsafely.
fsspec is transitive (torch / s3torchconnector / transformers), so the floor
goes in the pytorch closure's uv constraints; the lock resolves to 2026.9.0.
Also floored in the engine constraint files so the sglang and vLLM
runtime-deps resolves cannot pull an older fsspec back in.

Fixes the ecr-vulnerability-scan failures on lambda-pytorch, lambda-sglang
and lambda-vllm; base and cupy do not ship fsspec and already passed.

---------

Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>
Signed-off-by: Kevin Wang <kwanggg@amazon.com>
Co-authored-by: sirutBuasai <sirutbuasai27@outlook.com>
Co-authored-by: Sirut Buasai <73297481+sirutBuasai@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant