Repository navigation
feat(lambda): adopt GA awslambdaric 4.1.0, drop the preview variants - #6809
Merged
Merged
Conversation
register_pre_fork ships in the public RIC as of 4.1.0, so the separate preview images exist for nothing. Pin awslambdaric==4.1.0 in the base, cupy and pytorch closures and install it from PyPI — 4.1.0 publishes manylinux/musllinux wheels for cp310-cp315 with a prebuilt runtime_client.so, so the RIC builder stage and its gcc/cmake/autotools install go away along with the S3 tarball fetch in pre_build.sh. Removes the 5 *-preview configs, the builder-ric-preview stage, the 5 *-preview-py3 targets and ARG AWSLAMBDARIC_VERSION. The GA RIC has no AWS_LAMBDA_CONCURRENCY_MODE: multi-concurrency is entered by AWS_LAMBDA_MAX_CONCURRENCY alone and workers are forked processes, so thread and hybrid go with the preview images. That drops the run-ric-topology-test input, the lambdaric-topology-test job, the RIE topology + mode-override validation steps, rie_topology_check.sh and test/lambda/platform/ (run_topology.py, test_handler.py, Dockerfile.test). Both serving handlers lose the try/except ImportError and mode check: register_pre_fork when AWS_LAMBDA_MAX_CONCURRENCY is set, module-level otherwise. test_lambda_runtime.py drops its concurrency-mode tests and asserts register_pre_fork is importable in every image instead.
Eren-Jeager123
added a commit
that referenced
this pull request
Oct 2, 2026
The try/except around the lambda_concurrency_hooks import let the probe degrade silently on a pre-4.1.0 RIC, reporting a failed hook assertion instead of the real cause. 4.1.0 is pinned in every image closure, so an older RIC is a defect to fail on, not to tolerate — #6809 removed the same guard from the serving handlers. Importing register_pre_fork unconditionally now surfaces a pre-4.1.0 RIC as a Runtime.ImportModuleError at init. The baked-handler import is gated on the file existing instead of a broad except, so core images (which ship no handler.py) are still handled while an engine handler that fails to import now propagates. wait_ready greps the container logs for the init error before tailing, since RIE chatter otherwise pushes it out of the tail window. Verified on a scratch image pinned to awslambdaric 4.0.0: both shapes fail, exit 1, and the output names the cause — Unable to import module 'ric_probe_handler': No module named 'awslambdaric.lambda_concurrency_hooks' 4.1.0 image still passes 9/9. Signed-off-by: Kevin Wang <kwanggg@amazon.com>
1 task done
Eren-Jeager123
added a commit
that referenced
this pull request
Oct 7, 2026
* feat: add Lambda GPU DLC images (base, cupy, pytorch) Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * rename workflows Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add tmp torch as torch home Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add oss install Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix oss install link Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * bump uvlock Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix cve Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * convert to dispatch workflow for release Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change image path Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * release to gamma Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * release torch_home to prod Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add s3 torch connector for lambda pytorch container (#6211) Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add lambda ric patch package (#6217) * add lambda ric patch Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * run single line copy Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * deploy gamma with new RIC Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * revert to prod deployment Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> --------- Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * merge from main Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * upgrade lambda ric version (#6402) * upgrade lambda ric version Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix paths Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add additional tests Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * update pytorch Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> --------- Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove customer type for lambda Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add cupy type Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * release to prod Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change lambda docker path Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * feat: Lambda SGLang/vLLM serving engine (#6455) * re add sglang Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add vllm Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add vllm cache Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add timeout Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix telemetry Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> --------- Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * release sglang and vllm lambda Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * Add lambda ric test for LMI platform (#6585) * dockerfmt Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add lambda ric test Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix json jq Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use jq Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add warmup Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix vllm test Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix fasterlock Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> --------- Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * disable ric test Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add lambda vllm hf Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * bump hydra-core Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix handlers Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * update sglang and vllm (#6648) * update sglang and vllm Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add allowlist Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> --------- Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add s3 streaming for lambda inference engine (#6665) * add s3 streaming for lambda inference engine Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * wire s3 test Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * forward credentials Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * server without rie Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix curl Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> --------- Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix(lambda): report the real framework version, bump vLLM to 0.30.0, add modelscope (#6764) * fix(lambda): report the real framework version for pytorch and vllm images pytorch configs claimed 2.11.0 while pyproject pins torch==2.13.0, and the vllm configs carried the CUDA version (13.0.3) instead of vllm 0.27.1. The value is baked into bash_telemetry.sh at build time, so usage metrics for these four images were labelled with the wrong version. Release tags are unaffected: the lambda entry in DLContainersReleaseLogicPython frameworks.yml renders release_tag as "{container_type}" and never references {version}, and its version_pattern is '.*'. Signed-off-by: Kevin Wang <kwanggg@amazon.com> * fix(lambda): bump transformers to 5.10.4 for CVE-2026-9856 Path traversal via save_pretrained() in transformers <=5.8.0.dev0, fixed in 5.10.0. 5.10.0 itself is yanked upstream, so pin 5.10.4. * fix(lambda): bump vllm to 0.29.0 for CVE-2026-90553 RCE in the LlavaOnevision2 processor loader (ignores trust_remote_code), fixed in 0.28.0; take 0.29.0 as current. flashinfer moves to 0.6.18 to match 0.29.0's requirements/cuda.txt. The torch trio pin is unchanged (2.13.0 / 2.11.0 / 0.28.0), so torch-constraints.txt still applies, and the build-system requires are identical between the two refs. * feat(lambda): bump vllm to 0.30.0, align transformers, add modelscope vllm 0.30.0 (ref ced6857a) + flashinfer 0.6.18.post1 per its cuda.txt. vllm asks for an unbounded transformers>=5.10.4 while its upstream test suite hard-pins transformers==5.16.1, so the test env was swapping out the transformers the image ships. Pin 5.16.1 in the pytorch pyproject (the base every engine builder inherits) and add it to the vllm constraint file so the runtime resolve can't drift. safetensors moves 0.8.0rc1 -> 0.8.0 because transformers 5.16.1 needs the final release. modelscope<1.38 gives the vllm image ModelScope hub loading, matching the al2023 image's serving extras; sglang already gets it via sglang[all]. The cap is vllm#47325. The spec-decode examples now pass --gpu-memory-utilization 0.95: at the 0.9 default vllm 0.29+ leaves ~0.09 GiB for KV on the 1xL4 runner, which OOM'd the eagle case at max-model-len 2048. * fix(lambda): align vllm ref and transformers with main's AL2023 image Use the same vllm_ref as main's AL2023 vllm config (ec4a3a53, upstream's '[CI] Bump Transformers version to 5.17.0') instead of the v0.30.0 tag, so the upstream test suite checks out the ref the image is built from and its transformers==5.17.0 test pin matches what the image ships. Revert the hand-edit of the shared vllm_ec2_examples_test.sh and take main's copy, which already carries the gpu-memory-utilization fix. * fix(lambda): drop redundant transformers constraint The pytorch pyproject pin already fixes the version the vllm image ships: the vllm builder inherits /var/lang from the pytorch builder and uv leaves an already-satisfied requirement alone. * fix(lambda): install DeepGEMM/DeepSelect build deps, floor anyio for CVE-2026-63374 The vllm wheel build failed compiling deepgemm_C: DeepGEMM's vendored deep_jit needs elfutils/libdwfl.h. Install elfutils-devel and gcc14 and build with gcc14, matching docker/vllm/Dockerfile.amzn2023, which builds the same vllm ref green. anyio 4.13.0 (transitive via httpx) is CVE-2026-63374 CRITICAL, failing the pytorch and sglang scans; floor it to 4.14.2. * fix(lambda): drop added comments --------- Signed-off-by: Kevin Wang <kwanggg@amazon.com> * fix(lambda): point runtime caches at /tmp in base, cupy and pytorch images (#6791) * fix(lambda): point runtime caches at /tmp in base, cupy and pytorch images Lambda mounts everything except /tmp read-only, but only the sglang and vllm stages redirected their caches there. CuPy JIT-compiles kernels and writes them to $HOME/.cupy by default, so a customer's first custom kernel fails on a read-only HOME; numba, huggingface_hub, triton and inductor have the same problem in the cupy and pytorch images. Sets HOME, USER and the XDG dirs on all three stages, CUPY_CACHE_DIR and NUMBA_CACHE_DIR on cupy, and HF_HOME, TRITON_CACHE_DIR and TORCHINDUCTOR_CACHE_DIR on pytorch, matching what the engine stages already do. Preview stages inherit via FROM. Adds env assertions plus a cupy GPU test that compiles a RawKernel and checks the cache landed under /tmp — the case that was reported. * fix(lambda): source telemetry from $HOME/.bashrc in base, cupy and pytorch Setting HOME=/tmp moved the interactive-shell rc file off /root/.bashrc, which is where the telemetry hook was written — amzn2023 bash only reads /etc/bashrc via ~/.bashrc, so the hook stopped firing and four telemetry tests failed. Write the hook to $HOME/.bashrc too, as the sglang and vllm stages already do. * fix(lambda): ship libNVVM in the cupy image so numba.cuda works (#6807) The cupy image advertises Numba, but numba.cuda JIT-compiles kernels through libNVVM, which the CUDA runtime base omits: cuda.is_available() returned False and every @cuda.jit raised NvvmSupportError. CuPy is unaffected because it compiles via NVRTC, which the runtime base does ship. Copies nvvm/lib64 and nvvm/libdevice from the devel base and points CUDA_HOME at them. nvvm/bin (cicc, 73 MB) is left out — numba loads libnvvm.so directly and never invokes the nvcc driver, so the image grows by ~62 MB rather than 134. Covered by a GPU-free compile_ptx unit test and a @cuda.jit launch test, both verified to fail against an image without libNVVM. * feat(lambda): adopt GA awslambdaric 4.1.0, drop the preview variants (#6809) register_pre_fork ships in the public RIC as of 4.1.0, so the separate preview images exist for nothing. Pin awslambdaric==4.1.0 in the base, cupy and pytorch closures and install it from PyPI — 4.1.0 publishes manylinux/musllinux wheels for cp310-cp315 with a prebuilt runtime_client.so, so the RIC builder stage and its gcc/cmake/autotools install go away along with the S3 tarball fetch in pre_build.sh. Removes the 5 *-preview configs, the builder-ric-preview stage, the 5 *-preview-py3 targets and ARG AWSLAMBDARIC_VERSION. The GA RIC has no AWS_LAMBDA_CONCURRENCY_MODE: multi-concurrency is entered by AWS_LAMBDA_MAX_CONCURRENCY alone and workers are forked processes, so thread and hybrid go with the preview images. That drops the run-ric-topology-test input, the lambdaric-topology-test job, the RIE topology + mode-override validation steps, rie_topology_check.sh and test/lambda/platform/ (run_topology.py, test_handler.py, Dockerfile.test). Both serving handlers lose the try/except ImportError and mode check: register_pre_fork when AWS_LAMBDA_MAX_CONCURRENCY is set, module-level otherwise. test_lambda_runtime.py drops its concurrency-mode tests and asserts register_pre_fork is importable in every image instead. * chore(lambda): add checksum and signature verification for third-party artifacts (#6825) Implements the checksum/signature rows of the dependency verifiability inventory. Each change is annotated in the Dockerfile with its row number. Row 10 - FFmpeg: sha256 + GPG signature. Source moved from the GitHub archive tag to ffmpeg.org, which is the only publisher of the two that ships a detached signature; GitHub's archive tarballs are generated per request and are not byte-stable. Signing public key vendored at docker/lambda/keys/. Row 11 - rustup: fetch rustup-init and verify its published sha256 instead of piping sh.rustup.rs to a shell. Row 12 - sccache: verify the published sha256. Row 16 - nvidia/cuda runtime and devel: digest-pinned (13 FROM lines). Row 17 - uv: digest-pinned once as a uv-source stage, since COPY --from= does not expand ARGs. Not included, and why: Rows 4, 5, 7, 8 - hash enforcement for the sglang and vllm Python closures. Those stages have no lockfile, so this needs uv.lock + --require-hashes, exact pins in place of the >= constraints, and per-version hashes for sglang-kernel and flashinfer-cubin carried in the image configs. Needs build iteration to settle known resolution conflicts. Row 9 - git commit signature verification. Upstream commits are GPG-verified, but enforcing it in the build requires a decision on whose keys to trust. Rows 6, 13, 14 - no checksum or signature published at source. * test(security): refresh past-due lambda ECR scan allowlist review dates (#6832) All 19 dated entries hit their 2026-10-01 review date, failing every lambda ecr-vulnerability-scan job. Bumped to 2026-12-31, and corrected two reasons that no longer hold: - the 13 aws-lambda-rie Go stdlib / go-chi entries now cite the concrete state rather than 'no new RIE release yet': latest is v1.37 (2026-08-31), still built with Go 1.25.8, and we download releases/latest so they auto-resolve when a newer RIE ships. - CVE-2025-23308 / CVE-2025-23339 claimed the fix 'requires CUDA 13.x base image which is not yet available'. These images are on CUDA 13.0.3, so that premise is obsolete; the standing justification is that neither is exploitable at Lambda runtime (both need nvdisasm/cuobjdump run against a malicious ELF). No entries added or removed; only review_by and reason changed. * chore(lambda): point CI at main ahead of the branch merge PR workflows triggered on branches: [lambda], so a PR targeting main matched nothing and no lambda CI ran (branches: filters the base branch). Repointed all three to [main], matching every other framework on main. Fixes four path filters that silently matched nothing, in both the on.pull_request.paths and the paths-filter blocks: - scripts/lambda/** does not exist; core's entrypoint is scripts/docker/lambda/lambda_entrypoint.sh - scripts/common/** and scripts/telemetry/** are scripts/docker/common/** and scripts/docker/telemetry/** - the engine workflows watched scripts/docker/lambda/<engine>/** but the entrypoints sit beside that dir, so {vllm,sglang}_entrypoint.sh changes never triggered a build release-gate no longer accepts refs/heads/lambda. Release stays dispatch-only; no autorelease workflow, so the release-schedule prcheck does not apply. * fix(lambda): bump urllib3 to 2.8.0 for CVE-2026-97687 and CVE-2026-97689 Two HIGH CVEs in urllib3 2.7.0, failing the ecr-vulnerability-scan on all three core images: proxy/TLS options silently ignored on some code paths (CVE-2026-97687) and unbounded memory in the chunked streaming parser (CVE-2026-97689). Both fixed in 2.8.0. urllib3 is pinned explicitly in base, cupy and pytorch, so all three closures move; lock churn is urllib3 only. sglang and vllm inherit from the pytorch closure. * fix(lambda): bind torch_cuda_arch_list from image config The vllm builder declared `ARG TORCH_CUDA_ARCH_LIST`, but resolve_build_args.py keeps `torch_cuda_arch_list` lower-case (PRESERVE_CASE_KEYS) so the config value bound to nothing and the Dockerfile default was what compiled. Switch to the lower-case ARG + upper-case ENV pattern used by docker/vllm/Dockerfile.amzn2023. Also drop the duplicated arch-list fallbacks in the lambda build hooks: the value feeds the wheel cache key, so a local default diverging from the config would hand out a mismatched wheel. * test(lambda): smoke-test entrypoint dispatch and handler concurrency mode Adds test/lambda/unit/test_entrypoint.py (all images) covering the RIE-vs-exec branch, exec semantics, and argv forwarding, and test/lambda/unit/test_handler_concurrency.py (engine images) covering the AWS_LAMBDA_MAX_CONCURRENCY dispatch, server startup failure, and payload normalisation. Both run CPU-only with no engine launched. Also fail fast when the engine server dies during startup: the Popen handle was discarded, so a bad MODEL_ID, OOM or bad --load-format polled /health for the full timeout before raising, with the real error only in stderr. And narrow the sglang/vllm runtime stages to `dnf upgrade --security`, matching base/cupy/pytorch and the repo-wide convention; an unqualified upgrade is only used for named packages elsewhere. * fix(lambda): floor fsspec to 2026.6.0 for CVE-2026-104851 HIGH, ReferenceFileSystem evaluates Kerchunk reference fields unsafely. fsspec is transitive (torch / s3torchconnector / transformers), so the floor goes in the pytorch closure's uv constraints; the lock resolves to 2026.9.0. Also floored in the engine constraint files so the sglang and vLLM runtime-deps resolves cannot pull an older fsspec back in. Fixes the ecr-vulnerability-scan failures on lambda-pytorch, lambda-sglang and lambda-vllm; base and cupy do not ship fsspec and already passed. --------- Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> Signed-off-by: Kevin Wang <kwanggg@amazon.com> Co-authored-by: sirutBuasai <sirutbuasai27@outlook.com> Co-authored-by: Sirut Buasai <73297481+sirutBuasai@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
awslambdaric4.1.0 ships@register_pre_forkin the public RIC (upstream PR #226), so the separate preview image line is no longer needed. 10 images become 5.Changes
RIC from PyPI —
awslambdaric4.0.0 → 4.1.0 in the three dependency closures (base,cupy,pytorch) + relock. 4.1.0 publishes manylinux/musllinux wheels with a prebuiltruntime_client.sofor cp310–cp315, so nothing compiles: the RIC builder stage, itsgcc/cmake/autotools install, and the pre-build tarball fetch are all removed. sglang and vllm inherit the RIC from the pytorch builder, so no change is needed there.Preview images removed — the 5
*-previewconfigs,ARG AWSLAMBDARIC_VERSION, thebuilder-ric-previewstage and the 5*-preview-py3targets.Scope: process workers + pre_fork only — the GA RIC has no
AWS_LAMBDA_CONCURRENCY_MODE; multi-concurrency is entered byAWS_LAMBDA_MAX_CONCURRENCYalone and workers are forked processes, sothreadandhybridgo away with the preview images. That removes therun-ric-topology-testinput, thelambdaric-topology-testjob, the RIE-topology and mode-override validation steps,rie_topology_check.sh, andtest/lambda/platform/(whose only entry point was the topology runner).Handlers simplified — both serving handlers drop the
try/except ImportErrorand mode check:Testing
test_lambda_runtime.pydrops its concurrency-mode tests — the knob they asserted no longer exists — and gainstest_pre_fork_hook_available, assertingregister_pre_forkis importable. That is now a contract of every image rather than of the preview variants, so it runs in all 5 pipelines.One caveat: this also removes the end-to-end pre-fork coverage (hook runs once in the parent, N workers fork, all proxy to a single engine). Happy to re-add a trimmed process-only local check in a follow-up if that signal is wanted.