Skip to content

ci: investigate intermittent Linux/KVM Ubuntu timesync boot stalls #170

Description

Status update - September 21, 2026: The intermittent boot stall remains
unresolved. Draft #173 fixes missing failure-diagnostic retention only;
it does not change OpenVMM runtime behavior or close this issue.

The detailed investigation update
includes 28 focused native-KVM reproductions, six verified focused nested-KVM
reproductions, both complete KVM suites, exact boot-stage timings, artifact
identities, environment differences, and next discriminating checks.
A later CI run also passed on the original failing runner name and ID.

The original incident report is preserved below. Its references to the
"current dev head" describe the state at the incident, not the latest branch tip.


Summary

Linux/KVM CI intermittently stalls while booting the Ubuntu 25.04 UEFI guest for
multiarch::ic::openvmm_uefi_x64_ubuntu_2504_server_x64_timesync_ic.

The failing VM reached VM ready, began waiting for the firmware boot event, saved
one framebuffer screenshot, and then reported an unchanged framebuffer until Petri's
10-minute watchdog aborted the test. The test body never logged a guest time, so the
stall occurred before the time-sync assertions ran.

An immediate failed-jobs rerun passed on a different virtual-machine runner. The
same OpenVMM revision also passed this test in several earlier virtual-machine runs
and in both focused and full-suite bare-metal reproductions.

Current incident

  • Workflow: CI, run 35651853023, attempt 1
  • NVX head: caed7fad98837ac64310b5835370a47161e98058
  • OpenVMM source: 64054d7a7275c2c6d66c2bbb10f77ec777e11889
  • Failed job: OpenVMM vmm-tests / Linux / KVM
  • Failed step: Run OpenVMM tests on KVM
  • Runner: azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000004
  • Labels: self-hosted, linux, kvm, virtual-machine
  • Test result: 57 passed, 1 failed, 37 skipped
  • Failed test duration: 602.044s
  • Full test phase: 941.838s
  • Classification: current dev head; not cancellation or supersession

The downstream Required status check failure only reflected the failed
openvmm-vmm-tests result.

Decisive trace

[petri] VM ready
[petri] Waiting for boot event...
[petri] Screenshot saved.
[petri] No change in framebuffer, skipping screenshot.
...
FAIL [602.044s] multiarch::ic::openvmm_uefi_x64_ubuntu_2504_server_x64_timesync_ic
Test timed out

The last substantive OpenVMM activity was approximately 4.85 seconds after launch.
Petri then emitted unchanged-framebuffer messages every two seconds until the global
watchdog fired. There were no guest time messages from the test body.

The same job successfully booted other Ubuntu 25.04 UEFI variants, including the
plain boot test, secure boot, VMGS, VLAN, and battery tests.

Immediate rerun

The rerun used the same NVX and OpenVMM revisions but moved from
azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000004 to azure-kvm-3.

Recent successful virtual-machine runs

All entries below used the same OpenVMM pin.

Run Runner Affected test Full test phase
35603041650 azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005 277.662s 611.783s
35607488114 azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005 273.927s 607.584s
35609983784 azure-kvm-esaurez-nvx-ci-linux-kvm-vmss000005 282.185s 612.922s
35622360833 azure-kvm-2 283.047s 613.207s
35628432072 azure-kvm-1 303.827s 658.900s

These successful runs show that this test is consistently the longest selected KVM
test on virtual-machine runners, normally taking approximately 4.6-5.1 minutes.

Bare-metal reproduction

An isolated exact-SHA checkout was tested on prometheus32
(linux-kvm-baremetal):

cargo xflowey vmm-tests-run --release --ci-profile --skip-vhd-prompt \
  --filter "test(/^multiarch::ic::openvmm_uefi_x64_ubuntu_2504_server_x64_timesync_ic$/)"
  • Focused test: passed in 15.263s.
  • Original python3 scripts/nvx.py test-openvmm --backend kvm command: 58/58
    passed.
  • Affected test under the full suite: 19.635s.
  • Full test phase: 36.284s.
  • No OpenVMM process remained afterward.

The host-type difference means this does not reproduce the nested/virtual-machine
condition, but it helps rule out a deterministic failure in the pinned source,
guest image, or test body.

Relevant control flow

timesync_ic first awaits PetriVmBuilder::run(). That builder waits for the
expected firmware boot event before returning the VM and agent to the test. UEFI
guests do not have the PCAT-only flaky_boot timeout/reset workaround, so this wait
can continue until the global 10-minute watchdog aborts the test.

Impact and investigation requests

  • A single intermittent boot stall fails the required dev CI status despite all
    other selected KVM tests passing.
  • The job log references petri.log and screenshot.png, but the workflow does not
    upload the failed test-results directory, so the screenshot and complete
    diagnostics are unavailable after the runner workspace is recycled.

Suggested next checks:

  1. Preserve target/vmm_tests/test_results as a workflow-attempt-scoped artifact
    when a VMM test fails.
  2. Inspect the frozen screenshot and Petri/OpenVMM diagnostics from the next
    occurrence.
  3. Determine why this UEFI guest can stop producing boot progress after device and
    framebuffer initialization on nested KVM.
  4. Decide whether UEFI boot-event waiting needs a bounded diagnostic/reset path
    analogous to the existing PCAT flaky_boot handling, without merely extending
    or retrying the overall test timeout.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

confirmedIssue affects multiple people.

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions