Status update - September 21, 2026: The intermittent boot stall remains
unresolved. Draft #173 fixes missing failure-diagnostic retention only;
it does not change OpenVMM runtime behavior or close this issue.
The detailed investigation update
includes 28 focused native-KVM reproductions, six verified focused nested-KVM
reproductions, both complete KVM suites, exact boot-stage timings, artifact
identities, environment differences, and next discriminating checks.
A later CI run also passed on the original failing runner name and ID.
The original incident report is preserved below. Its references to the
"current dev head" describe the state at the incident, not the latest branch tip.
Summary
Linux/KVM CI intermittently stalls while booting the Ubuntu 25.04 UEFI guest for
multiarch::ic::openvmm_uefi_x64_ubuntu_2504_server_x64_timesync_ic.
The failing VM reached VM ready, began waiting for the firmware boot event, saved
one framebuffer screenshot, and then reported an unchanged framebuffer until Petri's
10-minute watchdog aborted the test. The test body never logged a guest time, so the
stall occurred before the time-sync assertions ran.
An immediate failed-jobs rerun passed on a different virtual-machine runner. The
same OpenVMM revision also passed this test in several earlier virtual-machine runs
and in both focused and full-suite bare-metal reproductions.
Current incident
- Workflow:
CI, run 35651853023, attempt 1
- NVX head:
caed7fad98837ac64310b5835370a47161e98058
- OpenVMM source:
64054d7a7275c2c6d66c2bbb10f77ec777e11889
- Failed job: OpenVMM vmm-tests / Linux / KVM
- Failed step:
Run OpenVMM tests on KVM
- Runner:
azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000004
- Labels:
self-hosted, linux, kvm, virtual-machine
- Test result:
57 passed, 1 failed, 37 skipped
- Failed test duration:
602.044s
- Full test phase:
941.838s
- Classification: current
dev head; not cancellation or supersession
The downstream Required status check failure only reflected the failed
openvmm-vmm-tests result.
Decisive trace
[petri] VM ready
[petri] Waiting for boot event...
[petri] Screenshot saved.
[petri] No change in framebuffer, skipping screenshot.
...
FAIL [602.044s] multiarch::ic::openvmm_uefi_x64_ubuntu_2504_server_x64_timesync_ic
Test timed out
The last substantive OpenVMM activity was approximately 4.85 seconds after launch.
Petri then emitted unchanged-framebuffer messages every two seconds until the global
watchdog fired. There were no guest time messages from the test body.
The same job successfully booted other Ubuntu 25.04 UEFI variants, including the
plain boot test, secure boot, VMGS, VLAN, and battery tests.
Immediate rerun
The rerun used the same NVX and OpenVMM revisions but moved from
azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000004 to azure-kvm-3.
Recent successful virtual-machine runs
All entries below used the same OpenVMM pin.
| Run |
Runner |
Affected test |
Full test phase |
| 35603041650 |
azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005 |
277.662s |
611.783s |
| 35607488114 |
azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005 |
273.927s |
607.584s |
| 35609983784 |
azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005 |
282.185s |
612.922s |
| 35622360833 |
azure-kvm-2 |
283.047s |
613.207s |
| 35628432072 |
azure-kvm-1 |
303.827s |
658.900s |
These successful runs show that this test is consistently the longest selected KVM
test on virtual-machine runners, normally taking approximately 4.6-5.1 minutes.
Bare-metal reproduction
An isolated exact-SHA checkout was tested on prometheus32
(linux-kvm-baremetal):
cargo xflowey vmm-tests-run --release --ci-profile --skip-vhd-prompt \
--filter "test(/^multiarch::ic::openvmm_uefi_x64_ubuntu_2504_server_x64_timesync_ic$/)"
- Focused test: passed in 15.263s.
- Original
python3 scripts/nvx.py test-openvmm --backend kvm command: 58/58
passed.
- Affected test under the full suite: 19.635s.
- Full test phase: 36.284s.
- No OpenVMM process remained afterward.
The host-type difference means this does not reproduce the nested/virtual-machine
condition, but it helps rule out a deterministic failure in the pinned source,
guest image, or test body.
Relevant control flow
timesync_ic first awaits PetriVmBuilder::run(). That builder waits for the
expected firmware boot event before returning the VM and agent to the test. UEFI
guests do not have the PCAT-only flaky_boot timeout/reset workaround, so this wait
can continue until the global 10-minute watchdog aborts the test.
Impact and investigation requests
- A single intermittent boot stall fails the required
dev CI status despite all
other selected KVM tests passing.
- The job log references
petri.log and screenshot.png, but the workflow does not
upload the failed test-results directory, so the screenshot and complete
diagnostics are unavailable after the runner workspace is recycled.
Suggested next checks:
- Preserve
target/vmm_tests/test_results as a workflow-attempt-scoped artifact
when a VMM test fails.
- Inspect the frozen screenshot and Petri/OpenVMM diagnostics from the next
occurrence.
- Determine why this UEFI guest can stop producing boot progress after device and
framebuffer initialization on nested KVM.
- Decide whether UEFI boot-event waiting needs a bounded diagnostic/reset path
analogous to the existing PCAT flaky_boot handling, without merely extending
or retrying the overall test timeout.
Summary
Linux/KVM CI intermittently stalls while booting the Ubuntu 25.04 UEFI guest for
multiarch::ic::openvmm_uefi_x64_ubuntu_2504_server_x64_timesync_ic.The failing VM reached
VM ready, began waiting for the firmware boot event, savedone framebuffer screenshot, and then reported an unchanged framebuffer until Petri's
10-minute watchdog aborted the test. The test body never logged a guest time, so the
stall occurred before the time-sync assertions ran.
An immediate failed-jobs rerun passed on a different virtual-machine runner. The
same OpenVMM revision also passed this test in several earlier virtual-machine runs
and in both focused and full-suite bare-metal reproductions.
Current incident
CI, run 35651853023, attempt 1caed7fad98837ac64310b5835370a47161e9805864054d7a7275c2c6d66c2bbb10f77ec777e11889Run OpenVMM tests on KVMazure-kvm-esaurez-nvx-ci-linux-kvm-vmss000004self-hosted,linux,kvm,virtual-machine57 passed, 1 failed, 37 skipped602.044s941.838sdevhead; not cancellation or supersessionThe downstream
Required status checkfailure only reflected the failedopenvmm-vmm-testsresult.Decisive trace
The last substantive OpenVMM activity was approximately 4.85 seconds after launch.
Petri then emitted unchanged-framebuffer messages every two seconds until the global
watchdog fired. There were no
guest timemessages from the test body.The same job successfully booted other Ubuntu 25.04 UEFI variants, including the
plain boot test, secure boot, VMGS, VLAN, and battery tests.
Immediate rerun
azure-kvm-3The rerun used the same NVX and OpenVMM revisions but moved from
azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000004toazure-kvm-3.Recent successful virtual-machine runs
All entries below used the same OpenVMM pin.
azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005azure-kvm-2azure-kvm-1These successful runs show that this test is consistently the longest selected KVM
test on virtual-machine runners, normally taking approximately 4.6-5.1 minutes.
Bare-metal reproduction
An isolated exact-SHA checkout was tested on
prometheus32(
linux-kvm-baremetal):python3 scripts/nvx.py test-openvmm --backend kvmcommand: 58/58passed.
The host-type difference means this does not reproduce the nested/virtual-machine
condition, but it helps rule out a deterministic failure in the pinned source,
guest image, or test body.
Relevant control flow
timesync_icfirst awaitsPetriVmBuilder::run(). That builder waits for theexpected firmware boot event before returning the VM and agent to the test. UEFI
guests do not have the PCAT-only
flaky_boottimeout/reset workaround, so this waitcan continue until the global 10-minute watchdog aborts the test.
Impact and investigation requests
devCI status despite allother selected KVM tests passing.
petri.logandscreenshot.png, but the workflow does notupload the failed test-results directory, so the screenshot and complete
diagnostics are unavailable after the runner workspace is recycled.
Suggested next checks:
target/vmm_tests/test_resultsas a workflow-attempt-scoped artifactwhen a VMM test fails.
occurrence.
framebuffer initialization on nested KVM.
analogous to the existing PCAT
flaky_boothandling, without merely extendingor retrying the overall test timeout.