Skip to content

Unbounded concurrent Monitoring API requests exhaust memory and OOM-kill the container #537 - #536

Open
ntmspavan wants to merge 1 commit into
prometheus-community:masterfrom
ntmspavan:fix/monitoring-max-concurrency-oom
Open

ntmspavan wants to merge 1 commit into
prometheus-community:masterfrom
ntmspavan:fix/monitoring-max-concurrency-oom

Conversation

@ntmspavan

@ntmspavan ntmspavan commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes #537.

Problem

reportMonitoringMetrics starts one goroutine per metric descriptor per project, with no cap. When google.projects.filter (or a long google.project-ids list) resolves to many projects, or a project has many metric descriptors, the number of in-flight TimeSeries.List requests is unbounded. Each one holds a fully decoded JSON response in memory. In a memory-constrained pod the heap can pass the container's memory limit and the pod is OOM-killed.

Fix

  1. Cap concurrent TimeSeries.List requests process-wide at a fixed internal limit (maxConcurrentTimeSeriesRequests = 20 in collectors/monitoring_collector.go), shared by every project's collector. This bounds how many decoded responses a scrape holds in memory at once, however many projects or descriptors it covers. It is an internal constant, not a new flag. 20 is a conservative starting value, and I'm happy to tune it based on feedback or profiling.

  2. At startup, read the container's cgroup memory limit and set GOMEMLIMIT to 90% of it via automemlimit (https://github.com/KimMachineGun/automemlimit), so the garbage collector reclaims more aggressively as usage nears the real limit. It applies only when a container memory limit exists (for example resources.limits.memory). Otherwise it changes nothing. The GOMEMLIMIT and AUTOMEMLIMIT environment variables still override it.

Trade-offs

  • The cap bounds in-flight API responses, not the total size of the scrape output, which the Prometheus registry still builds in memory.
  • Scrapes that cover many descriptors take longer, because requests are now processed at most 20 at a time. Narrowing monitoring.metrics-prefixes or raising scrape_timeout helps in large setups.

Testing

  • go build ./..., go vet ./..., gofmt -l . and golangci-lint run ./... are clean.
  • go test -race ./... passes.
  • Added TestTimeSeriesRequestLimiterBoundsConcurrency, which fails if concurrency exceeds the limit or never reaches it.

@ntmspavan ntmspavan changed the title Add monitoring.max-concurrency to bound per-scrape Monitoring API concurrency Add monitoring.max-concurrency to bound per-scrape Monitoring API concurrency #537 Aug 7, 2026
@ntmspavan ntmspavan changed the title Add monitoring.max-concurrency to bound per-scrape Monitoring API concurrency #537 Unbounded concurrent Monitoring API requests exhaust memory and OOM-kill the container #537 Aug 7, 2026
@ntmspavan
ntmspavan force-pushed the fix/monitoring-max-concurrency-oom branch 4 times, most recently from 38ec6f0 to 8debe5d Compare October 7, 2026 19:21
reportMonitoringMetrics starts one goroutine per metric descriptor per
project with no cap, and each holds a fully decoded TimeSeries.List
response while in flight. When google.projects.filter (or a long
google.project-ids list) resolves to many projects, or a project has many
metric descriptors, the number of in-flight responses is unbounded and
can push a memory-limited container past its limit and get it OOM killed.

Fix this without adding any configuration:

- Cap concurrent TimeSeries.List requests process-wide at a fixed
  internal limit (maxConcurrentTimeSeriesRequests), shared by every
  project's collector, so a scrape never holds more than that many
  decoded responses in memory at once.
- Read the container's cgroup memory limit at startup and set GOMEMLIMIT
  to 90% of it via automemlimit, so the garbage collector reclaims more
  aggressively as usage approaches the real limit. GOMEMLIMIT and
  AUTOMEMLIMIT still override it.

Signed-off-by: Pavan Nalam <pavan.nalam3693@gmail.com>
@ntmspavan
ntmspavan force-pushed the fix/monitoring-max-concurrency-oom branch from 8debe5d to 5ac415f Compare October 7, 2026 19:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Unbounded concurrent Monitoring API requests exhaust memory and OOM-kill the container

1 participant