🤖 Generated by the Agentic Engineer
Evidence (measured on prod, 2026-09-24): the DaemonSets plus Longhorn's per-node instance manager request about 19% CPU and 21% memory of a cx43 node before any workload lands on it. The descheduler's HighNodeUtilization strategy only drains a node below 15% of both, so no node can ever qualify: consolidation is inert. The one autoscaled node at the time sat at 20% CPU / 22% memory with three movable pods on it.
Audience / problem: the platform pays for an autoscaled server that the consolidation rule was built to hand back.
Hypothesis: once kube-scheduler scores with MostAllocated (#4166, for #2471), raising the threshold to 25% lets the descheduler empty a spare autoscaled node without pods bouncing back onto it.
Why a separate step: a deploy publishes and reconciles manifests before it updates Talos, so shipping both together could run the higher threshold against the old scoring for a while, and indefinitely if the Talos update fails. This change must only land after MostAllocated is confirmed live on every control plane.
Success signal: within a week of rollout, an autoscaled node that only runs DaemonSets plus a few small pods is drained and removed; descheduler evictions stay low (single digits per day); Coroot shows no HA tier collapsing onto one node.
Acceptance criteria
Size: small. Follow-up to #2471 (whose scheduler half ships in #4166).
Evidence (measured on prod, 2026-09-24): the DaemonSets plus Longhorn's per-node instance manager request about 19% CPU and 21% memory of a cx43 node before any workload lands on it. The descheduler's
HighNodeUtilizationstrategy only drains a node below 15% of both, so no node can ever qualify: consolidation is inert. The one autoscaled node at the time sat at 20% CPU / 22% memory with three movable pods on it.Audience / problem: the platform pays for an autoscaled server that the consolidation rule was built to hand back.
Hypothesis: once kube-scheduler scores with
MostAllocated(#4166, for #2471), raising the threshold to 25% lets the descheduler empty a spare autoscaled node without pods bouncing back onto it.Why a separate step: a deploy publishes and reconciles manifests before it updates Talos, so shipping both together could run the higher threshold against the old scoring for a while, and indefinitely if the Talos update fails. This change must only land after
MostAllocatedis confirmed live on every control plane.Success signal: within a week of rollout, an autoscaled node that only runs DaemonSets plus a few small pods is drained and removed; descheduler evictions stay low (single digits per day); Coroot shows no HA tier collapsing onto one node.
Acceptance criteria
MostAllocated.HighNodeUtilizationthresholds raised from 15% to 25% (cpu and memory), with the floor measurement recorded next to them.docs/node-autoscaling.mdand the HelmRelease comment beside the thresholds state the new threshold and the MostAllocated pairing (the comment still points at feat(scheduling): adopt MostAllocated kube-scheduler scoring so descheduler consolidation converges #2471; editing it moves the authorization fingerprint, hence here).Size: small. Follow-up to #2471 (whose scheduler half ships in #4166).