Skip to content

perf(memtrack): reduce benchmark noise and resolve workload symbols - #555

Open
not-matthias wants to merge 2 commits into
mainfrom
cod-3690-investigate-flaky-ls-and-tar-benchmarks
Open

not-matthias wants to merge 2 commits into
mainfrom
cod-3690-investigate-flaky-ls-and-tar-benchmarks

Conversation

@not-matthias

@not-matthias not-matthias commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Changes

  • Remove the short ls benchmark, whose measured time and variation are dominated by allocator-uprobe teardown.
  • Increase tar warmup from 5s to 30s while retaining real archive output and the 60s measurement budget.
  • Install exact-version, architecture-matched tar/coreutils debug packages from Launchpad over HTTPS before profiling. Do not upgrade workload binaries or depend on mutable ddebs indexes.

Missing symbols

The Ubuntu workload binaries lack full symbols, and the pinned Samply fork explicitly blocks Ubuntu debuginfod. Existing memtrack Rust names were already present. Installing local matching debug files resolves workload names and source locations before presymbolication.

All 20 hosted flamegraphs from the five executions have named tar/dd workload nodes. Tar graphs contain 181–212 workload nodes each, all named; dd graphs contain 23–35, all named. Tar's named hot frames consistently include flush_archive and sys_write_archive_buffer/flush_write. dd's leading named costs consistently include BPF CO-RE candidate matching and BTF initialization.

High-address frames with no mapped object remain unresolved (18.5–21.6% of sampled self-CPU weight in tar). Kernel-side attribution is plausible, but its cause is not established by these graphs. This change does not claim to resolve those frames.

Five-run validation

Exact revision: a2c9b21. All five tar and dd benchmark jobs passed. Four complete workflows passed; one failed only because macOS clippy could not resolve index.crates.io.

Benchmark Range (seconds) Sample CV
tar 9.299–9.403 0.45%
tar, physical 9.427–9.497 0.30%
dd 1.139–1.251 3.67%
dd, physical 1.228–1.326 2.92%

Each tar variant completed four warmup rounds and seven measured rounds. The first measured workload span was within 0.65% of the following-round median in every execution; event counts stayed at 24,902 normally and 26,403 with physical tracking. These five runs show low tar variation, not a guarantee of future stability. dd remains noisier.

CI executions: 36602936493, 36602947014, 36602957193, 36602966329, 36602976080.

  • Relevant prek hooks passed.
  • Ubuntu 22.04 container smoke verified unchanged tar/dd executable checksums and build-ID-matched debug files containing symbols and DWARF.
  • An earlier debug-package approach hit inconsistent Ubuntu ARM64 mirror indexes and implicitly upgraded workload packages. Those runs are excluded from the statistics above; exact-version Launchpad downloads avoid both problems.

The ls workload finishes quickly, while allocator-uprobe teardown accounts for most of the measured command time and its variation. Remove it from the walltime matrix instead of treating kernel grace-period latency as memtrack throughput.
The previous five-second budget allowed only one warmup round. An unusually fast first measured round could then determine the reported minimum. Increase warmup to thirty seconds while retaining the real archive and the existing sixty-second measurement budget.
@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

✅ 31 untouched benchmarks
⏩ 6 skipped benchmarks1


Comparing cod-3690-investigate-flaky-ls-and-tar-benchmarks (fcf3921) with main (080ed5f)

Open in CodSpeed

Footnotes

  1. 6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@not-matthias
not-matthias marked this pull request as ready for review September 30, 2026 07:57
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Low risk] Adjusts benchmark configuration and removes a benchmark workload.

The PR does not appear safe to merge while tar/dd flamegraphs lose the workload symbols this change aims to resolve.

Summary

The PR removes the short ls benchmark and its CI matrix entry, increases tar warmup to 30 seconds, and removes installation of matching tar/coreutils debug packages.

  • The longer tar warmup retains the 60-second measurement budget.
  • The debug-package removal undermines the stated workload-symbol resolution goal.

Reviews (2) · Last reviewed commit: "perf(memtrack): warm up tar for several ..."

Comment thread .github/workflows/ci.yml Outdated
Comment thread .github/workflows/ci.yml Outdated
@not-matthias
not-matthias removed the request for review from GuillaumeLagrange September 30, 2026 08:02
@not-matthias
not-matthias force-pushed the cod-3690-investigate-flaky-ls-and-tar-benchmarks branch from a68b18a to fcf3921 Compare September 30, 2026 08:58
@greptile-apps

greptile-apps Bot commented Sep 30, 2026

Copy link
Copy Markdown

Comments Outside Diff

These findings could not be posted inline.

  • P1 Workload symbols are lost .github/workflows/ci.yml:175 ▶

    The tar and dd jobs still run under Samply with presymbolication, but this workflow no longer installs their matching debug files. The Ubuntu workload binaries lack full symbols, and the pinned Samply fork does not use Ubuntu debuginfod, so the resulting flamegraphs lose tar/dd function names and source locations that this PR aims to provide.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants