From d498750b87664ff93eb899aab24aab110ad5bffc Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Fri, 11 Sep 2026 10:35:28 -0400 Subject: [PATCH 01/17] ci: fork automation (upstream sync + OCR AI review) Squashes the four CI commits this fork carried on top of upstream into one, and cuts the carry from 3776 lines to ~1000. Removed: - .github/{README,QUICKSTART}.md and .github/docs/* (2273 lines). Beyond being bulk we rebase hourly, they had drifted into being wrong: they described a daily (not hourly) sync, Claude 3.5 Sonnet via the direct Anthropic API (we run Opus on Bedrock), a static IAM user (we use OIDC), and a Windows dependency-builder workflow that the same commit which documented it had already deleted. 16 referenced paths did not exist. An AI reviewer reading this repo ingests that as truth. - sync-upstream-manual.yml (252 lines). sync-upstream.yml already has workflow_dispatch; the sole behavioral difference was a force_push toggle whose "false" setting cannot work after a rebase anyway. - .github/.gitignore, which only ignored scripts/ai-review/, a directory that does not exist here. Rewritten: - sync-upstream.yml -> fork-sync-upstream.yml, 260 lines to 47. git rebase already handles the fast-forward, no-op and diverged cases that the ahead/behind arithmetic hand-rolled, so all that goes away, as do the issue open/comment/close machinery (a scheduled-run failure is already emailed to the one consumer, and the run log carries what the issue body would) and the "dev setup|dev v[0-9]" commit allowlist, which matched zero commits on master and had since it was written. Sync drops to every three hours: hourly meant 24 full-history clones of an 840MB repo per day, each force-push also re-triggering pg-ci.yml on master, to collect a handful of upstream commits. Renamed: - ocr-review.yml, ocr-model-check.yml -> fork-*.yml. Upstream owns .github/ too (it has touched pg-ci.yml three times this year), so fork-owned workflows now live under a name prefix upstream will not collide with, making the rebase structurally conflict-free and letting the sync guard match fork-owned paths exactly rather than allowing all of .github/. Unchanged: the OCR review workflows' logic and .github/ocr/* (rule.json, context.md, litellm.yaml, pg-history.py) - the part that carries value. --- .github/ocr/context.md | 126 ++++++ .github/ocr/litellm.yaml | 41 ++ .github/ocr/pg-history.py | 225 +++++++++++ .github/ocr/rule.json | 65 ++++ .github/workflows/fork-ocr-model-check.yml | 89 +++++ .github/workflows/fork-ocr-review.yml | 427 +++++++++++++++++++++ .github/workflows/fork-sync-upstream.yml | 47 +++ 7 files changed, 1020 insertions(+) create mode 100644 .github/ocr/context.md create mode 100644 .github/ocr/litellm.yaml create mode 100644 .github/ocr/pg-history.py create mode 100644 .github/ocr/rule.json create mode 100644 .github/workflows/fork-ocr-model-check.yml create mode 100644 .github/workflows/fork-ocr-review.yml create mode 100644 .github/workflows/fork-sync-upstream.yml diff --git a/.github/ocr/context.md b/.github/ocr/context.md new file mode 100644 index 0000000000000..c4a83b85b124e --- /dev/null +++ b/.github/ocr/context.md @@ -0,0 +1,126 @@ +# OCR review context — PostgreSQL contribution standards + +You are reviewing a change to a **PostgreSQL** fork. Every PR here is destined to +become a patch posted to the **pgsql-hackers** mailing list and tracked in a +**commitfest**. Review with the combined rigor, taste, and attention to detail of +the PostgreSQL committers. This context applies to the *whole* change, on top of +the per-file rules. + +## Review discipline +- Be precise and blunt; lead with the most serious problem. No praise, no + validation of the author, no disclaimers — accuracy is the only metric. +- Verify every claim against the actual diff. Confirm names, signatures, line + numbers, and APIs before asserting. Never invent behavior or cite code not in + the change. If unsure, say so, and tag each finding **high / moderate / low** + confidence. +- Judge the change on its merits regardless of how the PR frames it. A draft PR + is WIP: weight design/approach feedback over style nits. + +## Patch hygiene (top rejection reasons on -hackers) +1. **Minimal diff.** The fastest way to get a patch rejected is unrelated + changes: reformatting untouched lines, rewording unrelated comments, touching + code not required by the change. Flag any hunk not needed for the stated + purpose. After the patch, the code should read as if it had always been + written that way. +2. **Atomic, bisectable commits.** Each commit must build and pass tests on its + own — a broken intermediate commit breaks `git bisect`, revert, and + cherry-pick. Flag a commit that only compiles once a later commit lands. + Prefer one focused patch, or a clearly-ordered series of + independently-committable pieces. +3. **Tests + docs are mandatory.** A user-visible change without regression/TAP + tests **and** documentation is WIP, not commit-ready. New behavior needs + tests that cover edge and error paths, not just the happy path. +4. **DRY / reuse.** Prefer existing infrastructure (`List` in `pg_list.h`, + `StringInfo`, `dynahash`/`simplehash`, `palloc`/`MemoryContext`, `foreach`) + over reinventing it. Flag copy-paste and speculative abstraction alike — the + community wants minimal, targeted changes that fit the subsystem's existing + patterns. +5. **Whitespace.** No trailing whitespace; tabs (width 4) for C indentation; + `git diff --check` must be clean. Whitespace-only churn on untouched lines is + a defect. + +## Committer-owned files — do NOT touch in a patch (flag if present) +These are the committer's job at push time; including them causes needless +merge conflicts and is a mistake: +- **`src/include/catalog/catversion.h`** — the `CATALOG_VERSION_NO` bump is done + by the **committer** when pushing. A catversion bump in the PR is **wrong** — + flag it. (This is the single most common author mistake in catalog patches.) +- **Release notes** (`doc/src/sgml/release-*.sgml`) and version strings + (`configure.ac` `AC_INIT` version, `meson.build` `version`, `PG_VERSION`). + +## Generated files — never hand-edit; edit the source +Flag direct edits to generated output; point the author at the source instead: +- Catalog headers `src/include/catalog/*_d.h`, `postgres.bki`, `schemapg.h`, + `system_constraints.sql` → edit the `pg_*.dat` files. +- `src/backend/nodes/{copy,equal,out,read}funcs.c` and other + `gen_node_support.pl` output → annotate the `Node` struct in its header. +- `fmgroids.h`, `fmgrprotos.h`, `fmgrtab.c` → edit `pg_proc.dat`. +- `utils/errcodes.h` → `errcodes.txt`; wait-event headers → + `wait_event_names.txt`; `lwlocknames.h` → `lwlocknames.txt`. +- `configure` → `configure.ac`; `*.po` translations are handled separately; + generated Unicode tables come from their source scripts. + +## Portability is a hard gate +PostgreSQL runs on Linux, Windows (MSVC), macOS, the BSDs and Solaris, across +**x86_64, ARM64, RISC-V, PPC64, s390x**, both endiannesses and 32/64-bit. Any +change must be portable across all of them: +- No unaligned memory access; no dependence on `char` signedness, integer/pointer + width, endianness, or struct padding for on-disk/wire formats. +- Use `int16/int32/int64`, `Size`, and `INT64_FORMAT`/`UINT64_FORMAT` (never + `%ld` for `int64`). +- Atomics/barriers only via `port/atomics` (`pg_atomic_*`, `pg_read/write_barrier`). +- **Windows/MSVC:** any `extern` variable used from another module or an + extension needs `PGDLLIMPORT` in its header; no VLAs or compiler-specific + extensions beyond the tree's C99 baseline. + +## Backward compatibility — the strongest constraint +Do not break SQL behavior, the libpq wire protocol, the logical-replication +protocol, dump/restore, `pg_upgrade`, or exported/`PGDLLIMPORT` APIs without +extraordinary justification. **ABI** matters for back-branches: changing the +size/layout of an exported struct or the signature of an exported function +breaks installed extensions. + +## Mailing-list context & etiquette +Because each PR becomes a pgsql-hackers email read by a busy, expert, opinionated +audience, also flag what reliably wastes reviewer time or draws rejection: +- A patch that **does more than one thing** or bundles unrelated cleanup — split it. +- **Footguns**: easy-to-misuse APIs, silent data-loss/corruption hazards, unsafe + defaults — name them explicitly. +- **Performance claims without a reproducible benchmark.** +- No reference to the **design discussion / prior -hackers thread** (Message-Id) + for a non-trivial change. +- **Do not bikeshed:** keep style nits proportionate and clearly separated from + substantive correctness findings. + +## Minimalism — the "ponytail" discipline +The best code is the code you never wrote (YAGNI). Before accepting new code, +apply the ladder: (1) Does this need to exist at all? (2) Can existing +code/infrastructure already do it? (3) Is this the simplest thing that works? +Flag: speculative scaffolding and config for a path that isn't wired yet; dead +code and unused "flexibility" (fields, params, abstractions, options with no +caller); premature abstraction (a helper used exactly once); knobs/GUCs/flags +nobody asked for. Minimal, targeted changes that fit the existing patterns beat +clever or general-purpose ones. + +## Comment & identity accuracy +- Comments must describe what the code does **now**. Flag aspirational/ + future-tense comments for behavior that already shipped ("will be", "for now", + "not yet", "future", and stale "TODO/FIXME/XXX/HACK"); comments that drifted + from the code they sit above; and incomplete/trailing comments. Comments + explain **why**, not what. No commented-out code. +- **ASCII only** in source and diffs — no smart quotes, em-dashes, or ellipsis + characters. + +## Commit & versioning discipline +- Conventional-commit style, imperative subject, one logical change per commit, + each commit building on its own. +- Do **not** bump version numbers or generated version stamps (including + `catversion.h`) — that is the maintainer's job at commit/release time. + +Understand common list shorthand so your comments are precise and not +miscommunicated: WIP (work in progress), GUC (config variable), WAL, LSN, OID, +TOAST, FSM, TAM (table access method), RLS, DSM, 2PC, PITR, CIC (concurrent index +creation), SAOP, ABI/API, backpatch (apply to supported back-branches), HEAD +(master tip), catversion (catalog version), pgindent, buildfarm, cfbot, +`s/x/y/` (suggested text substitution), footgun, bikeshedding, POLA (principle of +least astonishment). diff --git a/.github/ocr/litellm.yaml b/.github/ocr/litellm.yaml new file mode 100644 index 0000000000000..e23cc4eee6fe2 --- /dev/null +++ b/.github/ocr/litellm.yaml @@ -0,0 +1,41 @@ +# LiteLLM proxy config — bridges Open Code Review (OpenAI protocol) to AWS Bedrock. +# +# This proxy is NOT a hosted service. The ocr-review.yml workflow installs it +# (`pip install 'litellm[proxy]'`) and runs it as a background process bound to +# 127.0.0.1:4000 for the duration of a single GitHub Actions job, then it exits. +# +# Auth to Bedrock: LiteLLM uses boto3's default credential chain, which reads +# the temporary AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_SESSION_TOKEN +# minted by the workflow's OIDC "Configure AWS credentials" step; region from +# AWS_REGION. + +model_list: + - model_name: ocr-bedrock + litellm_params: + # Set the repo variable OCR_BEDROCK_MODEL to an Opus inference-profile id + # your account has access to, e.g.: + # bedrock/converse/us.anthropic.claude-opus-4-8 + # The 'converse/' prefix uses Bedrock's Converse API, which is the most + # reliable path for Claude tool-use (what OCR relies on). + model: os.environ/OCR_BEDROCK_MODEL + aws_region_name: os.environ/AWS_REGION + + # "High effort" review. Claude Opus 4.8 on Bedrock uses *adaptive* thinking + # controlled by output_config.effort. Set it DIRECTLY here — NOT via + # reasoning_effort, which LiteLLM still maps to the legacy + # thinking.type.enabled that Opus 4.8 rejects. LiteLLM forwards + # output_config into additionalModelRequestFields for Anthropic models; if + # the build doesn't recognize the effort param it is dropped with a warning + # (no error) and the model reviews at its default effort. + # Valid: low|medium|high|max|xhigh (auto-clamped to the model ceiling). + output_config: + effort: xhigh + max_tokens: 32000 + +litellm_settings: + drop_params: true # silently drop params a model doesn't support + modify_params: true # auto-fix minor request incompatibilities + request_timeout: 600 + +general_settings: + master_key: os.environ/LITELLM_MASTER_KEY diff --git a/.github/ocr/pg-history.py b/.github/ocr/pg-history.py new file mode 100644 index 0000000000000..5794f8a920bd7 --- /dev/null +++ b/.github/ocr/pg-history.py @@ -0,0 +1,225 @@ +#!/usr/bin/env python3 +""" +pg-history: tie a PR's changes to PostgreSQL git + pgsql-hackers email history. + +OCR (the code reviewer) cannot call MCP servers, so this is a separate agent: +it runs a Bedrock (Claude Opus) tool-use loop wired to the Agora MCP server at +https://pg.ddx.io/mcp, lets the model search the mailing-list archives / commit +history / commitfest data, and emits a Markdown summary linking the changes to +the relevant threads (https://pg.ddx.io/m/pgsql-hackers/). + +Env: + PG_HISTORY_MCP_URL MCP endpoint (default https://pg.ddx.io/mcp) + PG_HISTORY_MODEL Bedrock model id (e.g. us.anthropic.claude-opus-4-8) + AWS_REGION region (creds come from the OIDC step's env) + BASE_REF, HEAD_SHA PR base ref and head sha (for the git diff context) + GH_PR_TITLE PR title (optional, adds context) + PG_HISTORY_OUT output markdown path (default /tmp/pg-history.md) +Writes the markdown to PG_HISTORY_OUT; exits 0 even on soft failures (writes a note). +""" +import json, os, subprocess, sys, urllib.request + +MCP_URL = os.environ.get("PG_HISTORY_MCP_URL", "https://pg.ddx.io/mcp") +MODEL = os.environ.get("PG_HISTORY_MODEL", "us.anthropic.claude-opus-4-8").replace("bedrock/converse/", "").replace("bedrock/", "") +REGION = os.environ.get("AWS_REGION", "us-east-1") +BASE_REF = os.environ.get("BASE_REF", "") +HEAD_SHA = os.environ.get("HEAD_SHA", "") +PR_TITLE = os.environ.get("GH_PR_TITLE", "") +OUT = os.environ.get("PG_HISTORY_OUT", "/tmp/pg-history.md") +UA = "pg-history/0.1 (+github-actions)" + +# Curated subset of the 108 Agora tools — the ones useful for connecting a +# change to its discussion/commit history. Intersected with what the server +# actually exposes, so unknown names are harmless. +TOOL_WHITELIST = { + "find_related_discussions", "find_similar_messages", "get_thread", + "discussion_links", "get_author_messages", "browse_by_date", + "blame_symbol", "check_upstream_status", "find_related", + "find_entries_for_thread", "find_entries_for_author", "get_commit", + "search", "hybrid_search", "get_callers", "get_callees", "find_pattern", +} +MAX_ROUNDS = 14 +TOOL_RESULT_CAP = 8000 # chars per tool result fed back to the model + + +def _mcp_post(body, sid=None): + headers = {"Content-Type": "application/json", + "Accept": "application/json, text/event-stream", "User-Agent": UA} + if sid: + headers["Mcp-Session-Id"] = sid + req = urllib.request.Request(MCP_URL, data=json.dumps(body).encode(), headers=headers, method="POST") + resp = urllib.request.urlopen(req, timeout=60) + sid_out = resp.headers.get("Mcp-Session-Id") + result = None + for line in resp.read().decode().splitlines(): + line = line.strip() + if line.startswith("data:"): + line = line[5:].strip() + if not line or line.startswith("event:"): + continue + try: + obj = json.loads(line) + except Exception: + continue + if isinstance(obj, dict) and ("result" in obj or "error" in obj): + result = obj + return result, sid_out + + +class MCP: + def __init__(self): + init, self.sid = _mcp_post({"jsonrpc": "2.0", "id": 1, "method": "initialize", + "params": {"protocolVersion": "2025-06-18", "capabilities": {}, + "clientInfo": {"name": "pg-history", "version": "0.1"}}}) + if not init or "result" not in init: + raise RuntimeError(f"MCP initialize failed: {init}") + try: + _mcp_post({"jsonrpc": "2.0", "method": "notifications/initialized", "params": {}}, self.sid) + except Exception: + pass + self._id = 1 + + def list_tools(self): + self._id += 1 + res, _ = _mcp_post({"jsonrpc": "2.0", "id": self._id, "method": "tools/list", "params": {}}, self.sid) + return (res or {}).get("result", {}).get("tools", []) + + def call(self, name, args): + self._id += 1 + res, _ = _mcp_post({"jsonrpc": "2.0", "id": self._id, "method": "tools/call", + "params": {"name": name, "arguments": args or {}}}, self.sid) + if not res: + return "(no response)" + if "error" in res: + return f"ERROR: {json.dumps(res['error'])[:500]}" + parts = [] + for c in res.get("result", {}).get("content", []): + if c.get("type") == "text": + parts.append(c["text"]) + return ("\n".join(parts) or "(empty)")[:TOOL_RESULT_CAP] + + +def git(*args): + try: + return subprocess.check_output(["git", *args], text=True, stderr=subprocess.DEVNULL).strip() + except Exception: + return "" + + +def pr_context(): + base = f"origin/{BASE_REF}" if BASE_REF else "" + rng = f"{base}..{HEAD_SHA}" if base and HEAD_SHA else HEAD_SHA + commits = git("log", "--no-merges", "--format=%h %s", f"{rng}") if rng else "" + stat = git("diff", "--stat", rng) if rng else "" + files = git("diff", "--name-only", rng) if rng else "" + return commits[:4000], stat[:3000], files[:2000] + + +SYSTEM = """You are a PostgreSQL community research assistant. Given a pull request's +commits and changed files, use the available tools (backed by the Agora index of +pgsql-hackers mail, commit history, and commitfest data) to connect the change to +its history. Your goal: + +- Find the mailing-list thread(s) and prior discussion behind this change. +- Identify related/superseded prior commits and any commitfest entry. +- Note relevant prior art, rejected approaches, or design rationale. + +Rules (voice & rigor): +- Be precise and blunt. No praise, no filler, no hedging, no disclaimers. Accuracy is + the only success metric — not the author's approval. Lead with the most important finding. +- NEVER hallucinate. Verify every Message-ID, thread subject, commit hash, author name, + and date against an actual tool result before citing it. If a search returns nothing, + say so plainly — do not guess or fabricate a plausible-looking link. +- Assess the change on its merits, independent of how the PR frames it. +- Tag any inferred (not tool-confirmed) linkage with an explicit confidence level: + high / moderate / low. +- Be decisive and efficient: a handful of targeted tool calls, not exhaustive search. +- Cite every mailing-list message as a Markdown link: [subject](https://pg.ddx.io/m/pgsql-hackers/MESSAGE_ID). +- If you find nothing relevant, say so in one line — do not pad. + +When done, output ONLY Markdown (no preamble) with these sections, omitting any that are empty: +## 🧵 Related discussion +## 🔗 Related commits / prior art +## 📋 Commitfest +## 🧭 Context for reviewers +Keep it tight (use bullets; link generously).""" + + +def to_toolspec(t): + schema = t.get("inputSchema") or {"type": "object", "properties": {}} + return {"toolSpec": {"name": t["name"], + "description": (t.get("description") or "")[:600], + "inputSchema": {"json": schema}}} + + +def main(): + commits, stat, files = pr_context() + if not commits and not files: + open(OUT, "w").write("") # nothing to do + print("No PR diff context; skipping.") + return + user = (f"PR title: {PR_TITLE}\n\n" if PR_TITLE else "") + \ + f"Commits:\n{commits or '(none)'}\n\nChanged files:\n{files or '(none)'}\n\nDiffstat:\n{stat or '(none)'}\n" + + try: + mcp = MCP() + tools = [to_toolspec(t) for t in mcp.list_tools() if t.get("name") in TOOL_WHITELIST] + except Exception as e: + open(OUT, "w").write(f"_pg-history: could not reach the Agora MCP server ({MCP_URL}): {e}_\n") + print(f"MCP unavailable: {e}") + return + if not tools: + open(OUT, "w").write("_pg-history: no usable MCP tools available._\n") + return + + import boto3 + from botocore.config import Config + + # botocore's default read timeout (60s) is too short for a multi-round + # (MAX_ROUNDS) tool-use loop against a large PR diff on a reasoning model; + # each converse() call alone can take several minutes. Bump it well past + # what a single round needs; connect_timeout stays short since a stuck + # TCP handshake is a different (and much cheaper to detect) failure mode. + brt = boto3.client("bedrock-runtime", region_name=REGION, + config=Config(read_timeout=900, connect_timeout=10)) + messages = [{"role": "user", "content": [{"text": user}]}] + final_text = "" + try: + for _ in range(MAX_ROUNDS): + resp = brt.converse( + modelId=MODEL, + system=[{"text": SYSTEM}], + messages=messages, + toolConfig={"tools": tools}, + inferenceConfig={"maxTokens": 4096}, + ) + out = resp["output"]["message"] + messages.append(out) + if resp.get("stopReason") == "tool_use": + results = [] + for blk in out["content"]: + tu = blk.get("toolUse") + if not tu: + continue + res_text = mcp.call(tu["name"], tu.get("input") or {}) + results.append({"toolResult": {"toolUseId": tu["toolUseId"], + "content": [{"text": res_text}]}}) + messages.append({"role": "user", "content": results}) + continue + final_text = "".join(b.get("text", "") for b in out["content"]).strip() + break + except Exception as e: + open(OUT, "w").write(f"_pg-history: Bedrock call failed: {e}_\n") + print(f"Bedrock error: {e}") + return + + if not final_text: + final_text = "_pg-history: no related history found._" + body = "## 📜 Change history & discussion (Agora / pg.ddx.io)\n\n" + final_text + \ + "\n\nGenerated by pg-history via the Agora MCP server (pg.ddx.io).\n" + open(OUT, "w").write(body) + print(body) + + +if __name__ == "__main__": + main() diff --git a/.github/ocr/rule.json b/.github/ocr/rule.json new file mode 100644 index 0000000000000..60e13e73dcbe0 --- /dev/null +++ b/.github/ocr/rule.json @@ -0,0 +1,65 @@ +{ + "_comment": "OCR per-file review rules for PostgreSQL core + extensions. Cross-cutting contribution standards & mailing-list etiquette live in .github/ocr/context.md, passed via --background-file. OCR uses FIRST-MATCH-WINS in declaration order, so rules are ordered most-specific first. merge_system_rule:true keeps OCR's built-in fine-tuned checks (thread-safety, injection, NPE) alongside these PostgreSQL-specific rules.", + "rules": [ + { + "path": "src/test/**", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL tests. Coverage is mandatory for any behavioral change and must include edge cases (NULL, empty, boundary/overflow) and ERROR paths, not just the happy path. A test that still passes with the feature reverted is worthless — confirm it actually exercises and would catch regressions in the new code. Regression (.sql/expected): deterministic, portable output — ORDER BY where row order matters, no timing/plan-dependent output except intentional EXPLAIN, no absolute paths, locale-independent (C collation or explicit COLLATE), DROP objects the test creates; expected/ output must stay stable across platforms and under the parallel schedule. Concurrency/locking belongs in isolation tests (src/test/isolation, .spec + permutations). End-to-end/crash/replication/CLI behavior belongs in TAP tests (t/*.pl with PostgreSQL::Test::Cluster/Utils) — no hardcoded ports/paths, no sleep as synchronization (use poll_query_until/wait_for), skip cleanly when prerequisites are missing, and clean up nodes." + }, + { + "path": "**/*.{c,h}", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL backend/frontend C — review as pgsql-hackers committers do, in priority order.\n\n(1) CORRECTNESS (highest): Memory — every palloc lives in the right MemoryContext; error paths via ereport/elog(ERROR) must not leak memory/buffers/locks/fds (rely on MemoryContext/ResourceOwner reset or PG_TRY/PG_FINALLY); no use-after-free; delete temp contexts. Concurrency — consistent lock ordering (deadlock-free), correct lock levels, balanced LWLockAcquire/Release and START_/END_CRIT_SECTION, no TOCTOU, CHECK_FOR_INTERRUPTS in long loops, async-signal-safe signal handlers (volatile sig_atomic_t). WAL — any change to shared on-disk state must be WAL-logged AND correctly replayed (redo path), crash- and replica-consistent. NULL/edge/overflow handling.\n\n(2) BACKWARD COMPATIBILITY / ABI: don't break behavior, dump/restore, pg_upgrade, libpq wire protocol, logical-replication protocol, or exported/PGDLLIMPORT'd APIs (struct size/layout, function signatures) without extraordinary justification.\n\n(3) CATALOG / GENERATED: new/changed catalog data goes in pg_*.dat, NOT the generated *_d.h/.bki. New Node types: ANNOTATE the struct in its header so gen_node_support.pl regenerates copy/equal/out/read — do NOT hand-edit *funcs.c. New SQL-callable functions: add to pg_proc.dat with an OID from the 8000-9999 developer range (src/include/catalog/unused_oids; check duplicate_oids); committer renumbers at commit. DO NOT bump CATALOG_VERSION_NO in the patch — flag any catversion.h change as a mistake (committer's job).\n\n(4) PERFORMANCE: no regression on hot paths; avoid O(n^2) where better is feasible; minimize work under contended locks; avoid needless palloc churn and large struct copies in hot paths.\n\n(5) SECURITY: bounded string ops (snprintf/strlcpy/strlcat — never strcpy/strcat/sprintf); integer/size-overflow checks before allocation; never user input as a format string; privilege checks via pg_*_aclcheck; beware search_path and SECURITY DEFINER.\n\n(6) PORTABILITY (hard gate): no unaligned access; no dependence on char signedness, int/long/pointer width, endianness, or struct padding for on-disk/wire formats; use int16/int32/int64 + INT64_FORMAT/UINT64_FORMAT (never %ld for int64); align contended shared structs (pg_attribute_aligned/cache-line pad). Atomics/barriers only via port/atomics (pg_atomic_*, pg_read/write_barrier) — never raw intrinsics or volatile-as-barrier. WINDOWS/MSVC: extern vars used cross-module/extension need PGDLLIMPORT; no VLAs or features beyond the C99 baseline the tree targets; use pg_pread/pg_pwrite. Applies across x86_64/ARM64/RISC-V/PPC64/s390x, big/little endian, 32/64-bit.\n\n(7) CONVENTIONS: errmsg starts lowercase, no trailing period, no embedded newlines; errdetail/errhint are complete capitalized sentences; correct ERRCODE_*; wrap user-facing text in _(); errmsg_plural for counts. Assert() only for can't-happen invariants (never user-reachable). Naming: snake_case with subsystem prefix (heap_insert) or CamelCase for major subsystems (ExecInitNode); ALL_CAPS macros. Must pgindent cleanly (tabs, width 4). Comments explain WHY not WHAT; no #ifdef 0 blocks, no commented-out code, no #ifdef fencing your feature. Reuse existing helpers (DRY)." + }, + { + "path": "**/*.dat", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL catalog data (pg_proc.dat, pg_type.dat, etc.) — the SOURCE for generated headers. The generated *_d.h, postgres.bki, fmgroids.h, fmgrtab.c must NOT be hand-edited (they regenerate from these files). OIDs: use a value from the developer range 8000-9999 (src/include/catalog/unused_oids; verify with duplicate_oids); committer renumbers to a final contiguous block, so stay in-range and unique but don't over-optimize the exact number. Keep proc entries complete/consistent (prosrc, provolatile, proparallel, prorettype/proargtypes, matching description). DO NOT bump CATALOG_VERSION_NO / catversion.h — committer's job at push time; flag any such change. New catalog columns/views need documentation in doc/src/sgml/catalogs.sgml." + }, + { + "path": "**/*.{sql,pgsql}", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL SQL. Valid PostgreSQL dialect (not MySQL/Oracle); correct types (bigint vs int, text vs varchar); sound transaction/isolation and CTE-materialization assumptions. SECURITY: flag SQL injection in dynamic SQL (require quote_identifier/quote_literal or format() with %I/%L), SECURITY DEFINER without a locked-down search_path, inappropriate RLS bypass. Prefer set-based over row-at-a-time/N+1. BACKWARD COMPATIBILITY (a top rejection reason): changing existing SQL behavior, the output of existing functions, or default GUCs needs extraordinary justification. New SQL-callable objects belong in pg_*.dat with OIDs from the 8000-9999 range, not in generated files. Minimal diff; add regression tests + docs." + }, + { + "path": "**/*.{pl,pm}", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Perl (TAP tests and build/catalog tooling). Require 'use strict; use warnings;'. Must be perltidy-clean with the tree's src/tools/pgindent/perltidyrc and pass src/tools/perlcheck/pgperlcritic. Use the framework: PostgreSQL::Test::Cluster, PostgreSQL::Test::Utils, Test::More; no hardcoded ports/paths/PIDs; use safe_psql/poll_query_until, not sleep; skippable without optional prerequisites; clean up nodes. PORTABILITY: run on Windows (no fork-only constructs, use File::Spec, avoid unavailable signals) and the minimum supported Perl. Robustness: avoid two-arg open and string system()/qx with interpolated data (use list forms). Generator scripts (gen_node_support.pl, catalog Perl) must be deterministic and stay in sync with inputs; do not commit their generated output." + }, + { + "path": "**/*.py", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Python (build/test tooling, oauth/pytest tests, src/tools). Follow surrounding style; keep imports to the standard library unless the dependency is already required by the tree (no surprise third-party deps in build/test tooling). PORTABILITY: support the project's minimum Python 3 and run on Windows and the BSDs (use os.path/pathlib, avoid POSIX-only calls and shell=True with interpolated input). Deterministic, self-cleaning tests; no hardcoded ports/paths; skip cleanly without prerequisites. For the Perl->pytest porting effort, confirm behavior parity with the TAP test replaced (same assertions/coverage), not a superficial translation. Minimal diff; match the tree's ruff/black config if present." + }, + { + "path": "**/*.{rs,toml}", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. Rust PostgreSQL extension (pgrx) or Rust support crate. Not core C, but it runs inside/alongside the backend, so backend safety applies. SAFETY: in code reachable from an SQL call, a Rust panic aborts the Postgres process — forbid unwrap()/expect()/panic!/unreachable!/todo! and index-panics on reachable paths; use Result and pgrx error reporting (error!/ereport!). Every `unsafe` block needs a comment justifying its invariant; scrutinize raw pointers and FFI across the pg_sys boundary. pgrx: honor #[pg_guard] on extern C fns (correct panic/longjmp handling); never hold Rust references across SPI or anything that can longjmp (skips Rust destructors -> leaks); respect MemoryContext lifetimes for palloc'd data; datum<->Rust conversions must handle NULL. Concurrency uses Postgres shmem/LWLocks (pgrx shmem API), not std::sync alone. Lints: must pass `cargo clippy --all-targets --all-features -- -D warnings` and `cargo fmt --check`; deny unwrap_used/expect_used/panic in libraries; thiserror (libs) / anyhow (bins). Justify every new dependency. Tests: #[pg_test] for in-backend behavior, #[test] for pure logic; cover error and NULL paths. Minimal, idiomatic diff." + }, + { + "path": "**/{configure.ac,*.m4,aclocal.m4}", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Autoconf. Edit configure.ac / the m4 macros — do NOT hand-edit generated 'configure' or pg_config.h.in in the same patch (regeneration is the committer's step; a patch that also rewrites generated configure output is suspect). Feature/header/function probes must be portable and not assume a specific OS/compiler. Every configure knob must be mirrored on the Meson side (meson_options.txt/meson.build) and documented. Minimal diff." + }, + { + "path": "**/{meson.build,meson_options.txt}", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Meson build. Valid syntax; correct subdir()/dependency()/declare_dependency and install paths; new source files must be added here. CRITICAL: PostgreSQL maintains BOTH Meson and Autoconf/Make — any new file, option, or feature check must be mirrored on the configure.ac/Makefile side so the two never drift (a file built by only one system is a common defect). New options need matching docs and sensible defaults. Minimal diff." + }, + { + "path": "**/{Makefile,GNUmakefile,*.mk}", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Makefile (GNU Make). $(VAR) refs; correct .PHONY; accurate dependencies (no parallel -j races); $(MAKE) for recursion; VPATH/out-of-tree build support; no hardcoded paths (use standard PostgreSQL makefile vars and $(top_builddir)); clean/distclean/maintainer-clean must remove new artifacts; extensions use PGXS. Must stay in sync with meson.build. Minimal diff." + }, + { + "path": "doc/**/*.sgml", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL documentation (DocBook SGML). Technically accurate/complete (parameters, limitations, version/compat notes); correct tag usage/nesting (, , , , , /); working cross-references; spell it 'PostgreSQL' in prose; SQL keywords uppercase in examples. Coverage: a new GUC -> config.sgml (and postgresql.conf.sample); new/changed catalogs or views -> catalogs.sgml; new SQL syntax -> the matching ref/*.sgml; new functions -> func.sgml. Do NOT edit release-notes (release-*.sgml) — written by the release team/committers; flag such edits. New user-facing behavior in this PR should ship with matching docs." + }, + { + "path": "**/*.md", + "merge_system_rule": true, + "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. Markdown docs. Clear heading hierarchy; fenced code blocks with language hints; accurate instructions/prerequisites; consistent PostgreSQL terminology; no broken relative links or stale claims. Minimal diff." + } + ] +} diff --git a/.github/workflows/fork-ocr-model-check.yml b/.github/workflows/fork-ocr-model-check.yml new file mode 100644 index 0000000000000..10d250528cf7c --- /dev/null +++ b/.github/workflows/fork-ocr-model-check.yml @@ -0,0 +1,89 @@ +# Checks AWS Bedrock weekly for a newer Claude Opus inference profile than the +# one OCR currently uses (vars.OCR_BEDROCK_MODEL) and, if found, opens/updates a +# single GitHub issue telling the maintainer to bump the variable. It does NOT +# change the model automatically: GITHUB_TOKEN cannot write Actions *variables* +# (that needs a PAT with admin), so this is a notify-only mechanism by design. +name: OCR model self-check + +on: + schedule: + - cron: '0 12 * * 1' # Mondays 12:00 UTC + workflow_dispatch: + +permissions: + id-token: write + contents: read + issues: write + +jobs: + check-model: + runs-on: ubuntu-latest + steps: + - name: Configure AWS credentials (OIDC) + uses: aws-actions/configure-aws-credentials@v6 + with: + role-to-assume: ${{ vars.AWS_ROLE_ARN }} + aws-region: ${{ vars.AWS_REGION }} + role-session-name: ocr-model-check-${{ github.run_id }} + + - name: Find newest Opus vs configured + id: check + env: + CURRENT: ${{ vars.OCR_BEDROCK_MODEL }} + AWS_REGION: ${{ vars.AWS_REGION }} + run: | + python3 - <<'PY' >> "$GITHUB_OUTPUT" + import os, re, subprocess, json + region = os.environ.get("AWS_REGION", "us-east-1") + current = os.environ.get("CURRENT", "") + out = subprocess.run( + ["aws", "bedrock", "list-inference-profiles", "--region", region, + "--query", "inferenceProfileSummaries[].inferenceProfileId", "--output", "json"], + capture_output=True, text=True) + ids = json.loads(out.stdout or "[]") + # Parse claude-opus-- from any profile id (prefix us./global. ok). + def ver(s): + m = re.search(r"claude-opus-(\d+)-(\d+)", s) + return (int(m.group(1)), int(m.group(2))) if m else None + opus = [(ver(i), i) for i in ids if ver(i) and i.startswith(("us.", "global."))] + if not opus: + print("newer=false"); raise SystemExit(0) + best_ver, best_id = max(opus, key=lambda x: x[0]) + cur = ver(current) + newer = (cur is None) or (best_ver > cur) + print(f"newer={'true' if newer else 'false'}") + print(f"best_id={best_id}") + print(f"best_ver={best_ver[0]}.{best_ver[1]}") + print(f"cur_ver={'unknown' if cur is None else f'{cur[0]}.{cur[1]}'}") + PY + + - name: Open/update issue if a newer model exists + if: steps.check.outputs.newer == 'true' + uses: actions/github-script@v9 + with: + script: | + const best = '${{ steps.check.outputs.best_id }}'; + const bestVer = '${{ steps.check.outputs.best_ver }}'; + const curVer = '${{ steps.check.outputs.cur_ver }}'; + const marker = ''; + const title = `OCR: newer Claude Opus available (${bestVer} > ${curVer})`; + const body = `${marker}\n` + + `A newer Claude Opus inference profile is available on Bedrock.\n\n` + + `- **Configured** (\`vars.OCR_BEDROCK_MODEL\`): Opus ${curVer}\n` + + `- **Newest on Bedrock**: \`${best}\` (Opus ${bestVer})\n\n` + + `To upgrade, set the repo variable:\n\n` + + '```\n' + + `gh variable set OCR_BEDROCK_MODEL -R ${context.repo.owner}/${context.repo.repo} \\\n` + + ` -b "bedrock/converse/${best}"\n` + + '```\n\n' + + `Also confirm the \`ocr-bedrock-ci\` IAM inline policy allows invoking the new model ` + + `(the resource is scoped to \`anthropic.claude-opus-*\`), then re-run OCR.\n\n` + + `_Automated by \`.github/workflows/ocr-model-check.yml\`; this issue is upserted._`; + const q = `repo:${context.repo.owner}/${context.repo.repo} in:body "${marker}" state:open`; + const found = await github.rest.search.issuesAndPullRequests({ q, per_page: 1 }); + if (found.data.total_count > 0) { + const n = found.data.items[0].number; + await github.rest.issues.update({ owner: context.repo.owner, repo: context.repo.repo, issue_number: n, title, body }); + } else { + await github.rest.issues.create({ owner: context.repo.owner, repo: context.repo.repo, title, body }); + } diff --git a/.github/workflows/fork-ocr-review.yml b/.github/workflows/fork-ocr-review.yml new file mode 100644 index 0000000000000..0828af429b57c --- /dev/null +++ b/.github/workflows/fork-ocr-review.yml @@ -0,0 +1,427 @@ +# Open Code Review (OCR) — AI PR review backed by AWS Bedrock via a LiteLLM proxy. +# +# Flow: +# PR opened/updated (incl. DRAFTS) ─┐ +# /open-code-review PR comment ─┼─► start LiteLLM (127.0.0.1:4000 → Bedrock) +# manual workflow_dispatch ─┘ └► ocr review --format json +# └► post inline PR review comments +# +# Required (repo settings — all repo *variables*, no secrets; auth is via GitHub OIDC): +# vars.AWS_ROLE_ARN - IAM role to assume via OIDC (granting bedrock:InvokeModel*) +# vars.AWS_REGION - e.g. us-east-1 +# vars.OCR_BEDROCK_MODEL - LiteLLM model string for the Opus inference profile, e.g. +# bedrock/converse/us.anthropic.claude-opus-4-8 +# +# No static AWS keys are stored. GITHUB_TOKEN (auto) posts the review comments. + +name: OCR AI Review + +on: + pull_request: + # Note: no draft filter — drafts are reviewed too. + types: [opened, synchronize, reopened, ready_for_review] + issue_comment: + types: [created] + workflow_dispatch: + inputs: + pr_number: + description: 'PR number to review' + required: true + type: number + +# One review per PR; cancel superseded runs to save Bedrock spend. +concurrency: + group: ocr-review-${{ github.event.pull_request.number || github.event.issue.number || github.event.inputs.pr_number }} + cancel-in-progress: true + +permissions: + id-token: write # required to mint the GitHub OIDC token for AWS role assumption + contents: read + pull-requests: write + +jobs: + ocr-review: + runs-on: ubuntu-latest + # PR events always; comment events only when the comment is on a PR and + # starts with the trigger keyword; manual dispatch always. + if: | + github.event_name == 'pull_request' || + github.event_name == 'workflow_dispatch' || + (github.event_name == 'issue_comment' && github.event.issue.pull_request && + (startsWith(github.event.comment.body, '/open-code-review') || + startsWith(github.event.comment.body, '@open-code-review'))) + + env: + # LiteLLM listens on localhost only; this key never leaves the runner. + LITELLM_MASTER_KEY: sk-ocr-ci-local + OCR_BEDROCK_MODEL: ${{ vars.OCR_BEDROCK_MODEL }} + # Region is a static var (safe at job level). AWS credentials are NOT set + # here — they're minted by the OIDC "Configure AWS credentials" step below + # and exported to the environment for the LiteLLM/boto3 Bedrock calls. + AWS_REGION: ${{ vars.AWS_REGION }} + + steps: + - name: Resolve PR context + id: pr + uses: actions/github-script@v9 + with: + script: | + let prNumber; + if (context.eventName === 'pull_request') { + prNumber = context.payload.pull_request.number; + } else if (context.eventName === 'issue_comment') { + prNumber = context.issue.number; + } else { + prNumber = parseInt('${{ github.event.inputs.pr_number }}', 10); + } + const { data: pr } = await github.rest.pulls.get({ + owner: context.repo.owner, + repo: context.repo.repo, + pull_number: prNumber, + }); + const { data: repo } = await github.rest.repos.get({ + owner: context.repo.owner, + repo: context.repo.repo, + }); + core.setOutput('number', String(prNumber)); + core.setOutput('base_ref', pr.base.ref); + core.setOutput('head_ref', pr.head.ref); + core.setOutput('head_sha', pr.head.sha); + core.setOutput('default_branch', repo.default_branch); + core.setOutput('cross_repo', String(pr.head.repo.full_name !== pr.base.repo.full_name)); + + # NOTE: do NOT checkout the PR head. OCR reads the diff and file contents + # straight from git refs (git diff , git show :path, + # git grep ), so the working tree is irrelevant — but our OCR config + # lives on the default branch, not on the PR branch. We check out the repo + # (default ref), fetch the base/head objects, and materialize the config + # from origin/. + - name: Checkout + uses: actions/checkout@v6 + with: + fetch-depth: 0 + + - name: Prepare git refs and OCR config + env: + BASE_REF: ${{ steps.pr.outputs.base_ref }} + HEAD_REF: ${{ steps.pr.outputs.head_ref }} + HEAD_SHA: ${{ steps.pr.outputs.head_sha }} + DEFAULT_BRANCH: ${{ steps.pr.outputs.default_branch }} + run: | + git fetch --no-tags origin "+refs/heads/${DEFAULT_BRANCH}:refs/remotes/origin/${DEFAULT_BRANCH}" || true + git fetch --no-tags origin "+refs/heads/${BASE_REF}:refs/remotes/origin/${BASE_REF}" || true + git fetch --no-tags origin "+refs/heads/${HEAD_REF}:refs/remotes/origin/${HEAD_REF}" || true + git fetch --no-tags origin "${HEAD_SHA}" || true + + # OCR config lives on the default branch; materialize it independently + # of whatever ref is checked out. + mkdir -p "$RUNNER_TEMP/ocr" + git show "origin/${DEFAULT_BRANCH}:.github/ocr/litellm.yaml" > "$RUNNER_TEMP/ocr/litellm.yaml" + git show "origin/${DEFAULT_BRANCH}:.github/ocr/rule.json" > "$RUNNER_TEMP/ocr/rule.json" + git show "origin/${DEFAULT_BRANCH}:.github/ocr/context.md" > "$RUNNER_TEMP/ocr/context.md" + echo "Config materialized:"; ls -l "$RUNNER_TEMP/ocr" + + - name: Setup Python + uses: actions/setup-python@v6 + with: + python-version: '3.12' + + - name: Setup Node.js + uses: actions/setup-node@v6 + with: + node-version: '20' + + - name: Install LiteLLM proxy + Open Code Review + run: | + python -m pip install --upgrade pip + # Pin LiteLLM to a main commit that supports Claude Opus 4.8 adaptive + # thinking (maps reasoning_effort -> output_config.effort, incl. xhigh). + # Not in any tagged release yet (PyPI latest 1.87.1 lacks the Opus + # normalizer). Bump this SHA once a release ships the feature. + pip install "litellm[proxy] @ git+https://github.com/BerriAI/litellm.git@5be0797d24a2f26eb2123e13788f90055a59d91d" + npm install -g @alibaba-group/open-code-review + + - name: Configure AWS credentials (OIDC) + uses: aws-actions/configure-aws-credentials@v6 + with: + role-to-assume: ${{ vars.AWS_ROLE_ARN }} + aws-region: ${{ vars.AWS_REGION }} + role-session-name: ocr-review-${{ github.run_id }} + + - name: Start LiteLLM proxy (Bedrock bridge) + run: | + if [ -z "$OCR_BEDROCK_MODEL" ]; then + echo "::error::vars.OCR_BEDROCK_MODEL is not set (e.g. bedrock/converse/us.anthropic.claude-opus-4-1-20250805-v1:0)" + exit 1 + fi + nohup litellm --config "$RUNNER_TEMP/ocr/litellm.yaml" --host 127.0.0.1 --port 4000 \ + > /tmp/litellm.log 2>&1 & + echo "Waiting for LiteLLM to become ready..." + for i in $(seq 1 60); do + if curl -sf http://127.0.0.1:4000/health/readiness >/dev/null; then + echo "LiteLLM ready."; exit 0 + fi + sleep 2 + done + echo "::error::LiteLLM did not become ready in time"; cat /tmp/litellm.log; exit 1 + + - name: Configure OCR + run: | + ocr config set llm.url http://127.0.0.1:4000/v1/chat/completions + ocr config set llm.auth_token "$LITELLM_MASTER_KEY" + ocr config set llm.model ocr-bedrock + ocr config set llm.use_anthropic false + ocr config set language English + + - name: Run OCR review + run: | + ocr review \ + --from "origin/${{ steps.pr.outputs.base_ref }}" \ + --to "${{ steps.pr.outputs.head_sha }}" \ + --rule "$RUNNER_TEMP/ocr/rule.json" \ + --background-file "$RUNNER_TEMP/ocr/context.md" \ + --concurrency 3 \ + --timeout 20 \ + --format json \ + > /tmp/ocr-result.json 2>/tmp/ocr-stderr.log || true + echo "----- OCR stdout -----"; cat /tmp/ocr-result.json || true + echo "----- OCR stderr -----"; cat /tmp/ocr-stderr.log || true + echo "----- LiteLLM log (tail) -----"; tail -n 50 /tmp/litellm.log || true + + - name: Post review to PR + uses: actions/github-script@v9 + with: + github-token: ${{ secrets.GITHUB_TOKEN }} + script: | + const fs = require('fs'); + const prNumber = parseInt('${{ steps.pr.outputs.number }}', 10); + const commitSha = '${{ steps.pr.outputs.head_sha }}'; + + // Opus at high effort can emit dozens of findings. Posting them all + // one-by-one trips GitHub's SECONDARY rate limit (403 "content + // creation"), which is what made every run fail after the review + // was already generated. We (a) cap inline comments and overflow + // the rest into the summary, (b) prefer a single bulk createReview, + // and (c) throttle + back off with Retry-After on any fallback. + const MAX_INLINE = 25; + const sleep = (ms) => new Promise(r => setTimeout(r, ms)); + + async function withRetry(fn, label) { + for (let attempt = 1; attempt <= 5; attempt++) { + try { return await fn(); } + catch (e) { + const status = e.status || (e.response && e.response.status); + const h = (e.response && e.response.headers) || {}; + const isRate = status === 403 || status === 429; + if (!isRate || attempt === 5) throw e; + let waitMs = 0; + if (h['retry-after']) waitMs = parseInt(h['retry-after'], 10) * 1000; + else if (h['x-ratelimit-reset']) waitMs = parseInt(h['x-ratelimit-reset'], 10) * 1000 - Date.now(); + if (!waitMs || Number.isNaN(waitMs) || waitMs < 0) waitMs = 1000 * Math.pow(2, attempt); + waitMs = Math.min(waitMs, 60000) + 500; + core.warning(`${label}: rate-limited (status ${status}); waiting ${Math.round(waitMs / 1000)}s (attempt ${attempt}/5)`); + await sleep(waitMs); + } + } + } + + let result; + try { + result = JSON.parse(fs.readFileSync('/tmp/ocr-result.json', 'utf8')); + } catch (e) { + const stderr = (() => { try { return fs.readFileSync('/tmp/ocr-stderr.log', 'utf8').trim(); } catch { return ''; } })(); + await withRetry(() => github.rest.issues.createComment({ + owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, + body: `⚠️ **OCR** could not produce a review.\n\n\`\`\`\n${(stderr || e.message).slice(0, 8000)}\n\`\`\``, + }), 'error-comment'); + return; + } + + const comments = result.comments || []; + const warnings = result.warnings || []; + + const formatComment = (c) => { + let body = c.content || ''; + if (c.suggestion_code && c.existing_code) { + body += '\n\n```suggestion\n' + c.suggestion_code + (c.suggestion_code.endsWith('\n') ? '' : '\n') + '```'; + } + return body; + }; + const formatMarkdown = (c) => { + let md = `### 📄 \`${c.path}\``; + if (c.start_line && c.end_line) md += ` (L${c.start_line}-L${c.end_line})`; + md += '\n\n' + (c.content || ''); + if (c.suggestion_code && c.existing_code) { + md += '\n\n
💡 Suggested change\n\n'; + md += '**Before:**\n```\n' + c.existing_code + '\n```\n\n**After:**\n```\n' + c.suggestion_code + '\n```\n\n
'; + } + return md; + }; + + if (comments.length === 0) { + await withRetry(() => github.rest.issues.createComment({ + owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, + body: `✅ **OCR**: ${result.message || 'No issues found.'}`, + }), 'no-issues-comment'); + return; + } + + const inlineAll = []; + const noLine = []; + for (const c of comments) { + const body = formatComment(c); + const hasLine = (c.start_line >= 1) || (c.end_line >= 1); + if (!hasLine) { noLine.push(c); continue; } + const rc = { path: c.path, body, side: 'RIGHT' }; + if (c.start_line >= 1 && c.end_line >= 1 && c.start_line !== c.end_line) { + rc.start_line = c.start_line; rc.line = c.end_line; rc.start_side = 'RIGHT'; + } else { + rc.line = c.end_line >= 1 ? c.end_line : c.start_line; + } + inlineAll.push({ rc, c }); + } + + const inline = inlineAll.slice(0, MAX_INLINE).map(x => x.rc); + const overflow = inlineAll.slice(MAX_INLINE).map(x => x.c); + + let summary = `🔍 **OCR** found **${comments.length}** issue(s).`; + summary += `\n- ${inline.length} inline, ${noLine.length + overflow.length} in summary`; + if (overflow.length) summary += ` (inline capped at ${MAX_INLINE})`; + if (warnings.length) summary += `\n- ⚠️ ${warnings.length} warning(s) during review`; + for (const c of noLine.concat(overflow)) summary += '\n\n---\n\n' + formatMarkdown(c); + + // Preferred path: ONE createReview carrying every inline comment. + try { + await withRetry(() => github.rest.pulls.createReview({ + owner: context.repo.owner, repo: context.repo.repo, pull_number: prNumber, + commit_id: commitSha, body: summary, event: 'COMMENT', comments: inline, + }), 'bulk-review'); + return; + } catch (e) { + core.warning(`bulk createReview failed (${e.status || '?'}: ${e.message}); falling back to throttled per-comment posting`); + } + + // Fallback: an invalid inline position (line not in the diff -> 422) + // rejects the whole bulk review. Post the summary, then each comment + // individually with a delay + backoff, skipping ones GitHub rejects. + let ok = 0; const failed = []; + try { + await withRetry(() => github.rest.pulls.createReview({ + owner: context.repo.owner, repo: context.repo.repo, pull_number: prNumber, + commit_id: commitSha, body: summary, event: 'COMMENT', + }), 'summary-review'); + } catch (err) { failed.push(`summary: ${err.message}`); } + + for (const rc of inline) { + try { + await withRetry(() => github.rest.pulls.createReviewComment({ + owner: context.repo.owner, repo: context.repo.repo, pull_number: prNumber, + commit_id: commitSha, path: rc.path, body: rc.body, + ...(rc.start_line ? { start_line: rc.start_line, start_side: rc.start_side } : {}), + line: rc.line, side: rc.side, + }), `comment ${rc.path}:${rc.line}`); + ok++; + } catch (inner) { + failed.push(`\`${rc.path}\` L${rc.line}: ${inner.message}`); + } + await sleep(1200); // stay under the secondary content-creation limit + } + + if (failed.length) { + await withRetry(() => github.rest.issues.createComment({ + owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, + body: `📊 OCR posted ${ok}/${inline.length} inline comment(s).\n\n
${failed.length} could not be posted\n\n${failed.join('\n')}\n
`, + }), 'summary-failures'); + } + + # Companion job: OCR can't call MCP, so this separate agent ties the PR's + # changes to PostgreSQL git + pgsql-hackers history via the Agora MCP server + # (pg.ddx.io) and posts a single, upserted "history & discussion" comment. + pg-history: + runs-on: ubuntu-latest + if: | + github.event_name == 'pull_request' || + github.event_name == 'workflow_dispatch' || + (github.event_name == 'issue_comment' && github.event.issue.pull_request && + (startsWith(github.event.comment.body, '/open-code-review') || + startsWith(github.event.comment.body, '@open-code-review') || + startsWith(github.event.comment.body, '/pg-history'))) + steps: + - name: Resolve PR context + id: pr + uses: actions/github-script@v9 + with: + script: | + let prNumber; + if (context.eventName === 'pull_request') prNumber = context.payload.pull_request.number; + else if (context.eventName === 'issue_comment') prNumber = context.issue.number; + else prNumber = parseInt('${{ github.event.inputs.pr_number }}', 10); + const { data: pr } = await github.rest.pulls.get({ + owner: context.repo.owner, repo: context.repo.repo, pull_number: prNumber }); + core.setOutput('number', String(prNumber)); + core.setOutput('base_ref', pr.base.ref); + core.setOutput('head_sha', pr.head.sha); + core.setOutput('title', pr.title || ''); + + - name: Checkout + uses: actions/checkout@v6 + with: + fetch-depth: 0 + + - name: Make base/head refs available + env: + BASE_REF: ${{ steps.pr.outputs.base_ref }} + HEAD_SHA: ${{ steps.pr.outputs.head_sha }} + run: | + git fetch --no-tags origin "+refs/heads/${BASE_REF}:refs/remotes/origin/${BASE_REF}" || true + git fetch --no-tags origin "${HEAD_SHA}" || true + + - name: Setup Python + uses: actions/setup-python@v6 + with: + python-version: '3.12' + + - name: Configure AWS credentials (OIDC) + uses: aws-actions/configure-aws-credentials@v6 + with: + role-to-assume: ${{ vars.AWS_ROLE_ARN }} + aws-region: ${{ vars.AWS_REGION }} + role-session-name: pg-history-${{ github.run_id }} + + - name: Install deps + run: pip install boto3 + + - name: Run pg-history (Agora MCP) + env: + PG_HISTORY_MODEL: ${{ vars.OCR_BEDROCK_MODEL }} + AWS_REGION: ${{ vars.AWS_REGION }} + BASE_REF: ${{ steps.pr.outputs.base_ref }} + HEAD_SHA: ${{ steps.pr.outputs.head_sha }} + GH_PR_TITLE: ${{ steps.pr.outputs.title }} + PG_HISTORY_OUT: ${{ runner.temp }}/pg-history.md + run: | + python .github/ocr/pg-history.py || true + echo "----- output -----"; cat "${{ runner.temp }}/pg-history.md" 2>/dev/null || echo "(no output)" + + - name: Upsert PR comment + uses: actions/github-script@v9 + with: + script: | + const fs = require('fs'); + const path = process.env.RUNNER_TEMP + '/pg-history.md'; + let body = ''; + try { body = fs.readFileSync(path, 'utf8').trim(); } catch (e) {} + if (!body) { console.log('pg-history: empty output, nothing to post'); return; } + const prNumber = parseInt('${{ steps.pr.outputs.number }}', 10); + const marker = ''; + body = marker + '\n' + body; + const { data: comments } = await github.rest.issues.listComments({ + owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, per_page: 100 }); + const mine = comments.find(c => c.user.type === 'Bot' && c.body && c.body.includes(marker)); + if (mine) { + await github.rest.issues.updateComment({ + owner: context.repo.owner, repo: context.repo.repo, comment_id: mine.id, body }); + } else { + await github.rest.issues.createComment({ + owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, body }); + } diff --git a/.github/workflows/fork-sync-upstream.yml b/.github/workflows/fork-sync-upstream.yml new file mode 100644 index 0000000000000..e950821239968 --- /dev/null +++ b/.github/workflows/fork-sync-upstream.yml @@ -0,0 +1,47 @@ +# Keep this fork's master rebased on postgres/postgres master, carrying only +# the fork-owned CI files (.github/workflows/fork-*.yml and .github/ocr/). +# +# Requires the SYNC_TOKEN secret: a PAT with repo+workflow scope. The default +# GITHUB_TOKEN is not allowed to push commits that touch .github/workflows/, +# and every commit we carry does. + +name: Sync from upstream + +on: + schedule: + - cron: '17 */3 * * *' + workflow_dispatch: + +concurrency: + group: fork-sync-upstream + +permissions: + contents: read + +jobs: + sync: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v6 + with: + fetch-depth: 0 + token: ${{ secrets.SYNC_TOKEN }} + + - name: Rebase master onto upstream + run: | + set -euo pipefail + git config user.name "github-actions[bot]" + git config user.email "github-actions[bot]@users.noreply.github.com" + git remote add upstream https://github.com/postgres/postgres.git + git fetch --no-tags upstream master + + # Master carries fork CI only. Anything else means work landed here + # that belongs on a branch; force-rebasing it on a timer would be wrong. + if git diff --name-only upstream/master...HEAD | + grep -vE '^(\.github/workflows/fork-|\.github/ocr/)'; then + echo "::error::master carries non-CI changes (listed above); move them to a branch" + exit 1 + fi + + git rebase upstream/master + git push --force-with-lease From a83f4d233603806b989eda78dcb3f2544c45b41d Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Fri, 11 Sep 2026 10:48:09 -0400 Subject: [PATCH 02/17] ci: move fork CI logic to gburd/ci-workflows The OCR review workflows and their config were the last bulk this fork carried on top of upstream: ~970 lines rebased onto postgres/postgres every few hours, editable only by amending a commit that gets force- pushed. They are now reusable (workflow_call) workflows in the sidecar repo gburd/ci-workflows, and master carries two caller stubs instead. Carry after this: 3 files, ~110 lines, one commit. The sidecar is also where the value is: ocr/rule.json and ocr/context.md (the PostgreSQL committer-grade review corpus) become normal reviewable files with normal PRs and history, rather than a blob inside a commit that is rewritten on every sync. Mechanics worth knowing: - The review jobs check ci-workflows out at github.job_workflow_sha, so the rules always come from the same commit as the logic using them, and pinning @v1 (or a sha) in the stub pins both. This replaces the old "git show origin/:.github/ocr/..." dance, which existed only because the config had to be read off a branch other than the PR's. - AWS settings are passed as explicit inputs rather than read from vars inside the reusable workflow. vars does resolve across the boundary, but the contract stays visible at the call site. - A reusable workflow does not change the OIDC subject: it stays repo:gburd/postgres:*, so the ocr-bedrock-ci trust policy is unchanged. - ci-workflows must stay public; callers read it with GITHUB_TOKEN. The sync guard narrows to .github/workflows/fork-* accordingly, and sync-upstream keeps needing SYNC_TOKEN: the stubs are still workflow files, which GITHUB_TOKEN may not push. --- .github/ocr/context.md | 126 ------- .github/ocr/litellm.yaml | 41 --- .github/ocr/pg-history.py | 225 ----------- .github/ocr/rule.json | 65 ---- .github/workflows/fork-ocr-model-check.yml | 91 +---- .github/workflows/fork-ocr-review.yml | 410 +-------------------- .github/workflows/fork-sync-upstream.yml | 5 +- 7 files changed, 26 insertions(+), 937 deletions(-) delete mode 100644 .github/ocr/context.md delete mode 100644 .github/ocr/litellm.yaml delete mode 100644 .github/ocr/pg-history.py delete mode 100644 .github/ocr/rule.json diff --git a/.github/ocr/context.md b/.github/ocr/context.md deleted file mode 100644 index c4a83b85b124e..0000000000000 --- a/.github/ocr/context.md +++ /dev/null @@ -1,126 +0,0 @@ -# OCR review context — PostgreSQL contribution standards - -You are reviewing a change to a **PostgreSQL** fork. Every PR here is destined to -become a patch posted to the **pgsql-hackers** mailing list and tracked in a -**commitfest**. Review with the combined rigor, taste, and attention to detail of -the PostgreSQL committers. This context applies to the *whole* change, on top of -the per-file rules. - -## Review discipline -- Be precise and blunt; lead with the most serious problem. No praise, no - validation of the author, no disclaimers — accuracy is the only metric. -- Verify every claim against the actual diff. Confirm names, signatures, line - numbers, and APIs before asserting. Never invent behavior or cite code not in - the change. If unsure, say so, and tag each finding **high / moderate / low** - confidence. -- Judge the change on its merits regardless of how the PR frames it. A draft PR - is WIP: weight design/approach feedback over style nits. - -## Patch hygiene (top rejection reasons on -hackers) -1. **Minimal diff.** The fastest way to get a patch rejected is unrelated - changes: reformatting untouched lines, rewording unrelated comments, touching - code not required by the change. Flag any hunk not needed for the stated - purpose. After the patch, the code should read as if it had always been - written that way. -2. **Atomic, bisectable commits.** Each commit must build and pass tests on its - own — a broken intermediate commit breaks `git bisect`, revert, and - cherry-pick. Flag a commit that only compiles once a later commit lands. - Prefer one focused patch, or a clearly-ordered series of - independently-committable pieces. -3. **Tests + docs are mandatory.** A user-visible change without regression/TAP - tests **and** documentation is WIP, not commit-ready. New behavior needs - tests that cover edge and error paths, not just the happy path. -4. **DRY / reuse.** Prefer existing infrastructure (`List` in `pg_list.h`, - `StringInfo`, `dynahash`/`simplehash`, `palloc`/`MemoryContext`, `foreach`) - over reinventing it. Flag copy-paste and speculative abstraction alike — the - community wants minimal, targeted changes that fit the subsystem's existing - patterns. -5. **Whitespace.** No trailing whitespace; tabs (width 4) for C indentation; - `git diff --check` must be clean. Whitespace-only churn on untouched lines is - a defect. - -## Committer-owned files — do NOT touch in a patch (flag if present) -These are the committer's job at push time; including them causes needless -merge conflicts and is a mistake: -- **`src/include/catalog/catversion.h`** — the `CATALOG_VERSION_NO` bump is done - by the **committer** when pushing. A catversion bump in the PR is **wrong** — - flag it. (This is the single most common author mistake in catalog patches.) -- **Release notes** (`doc/src/sgml/release-*.sgml`) and version strings - (`configure.ac` `AC_INIT` version, `meson.build` `version`, `PG_VERSION`). - -## Generated files — never hand-edit; edit the source -Flag direct edits to generated output; point the author at the source instead: -- Catalog headers `src/include/catalog/*_d.h`, `postgres.bki`, `schemapg.h`, - `system_constraints.sql` → edit the `pg_*.dat` files. -- `src/backend/nodes/{copy,equal,out,read}funcs.c` and other - `gen_node_support.pl` output → annotate the `Node` struct in its header. -- `fmgroids.h`, `fmgrprotos.h`, `fmgrtab.c` → edit `pg_proc.dat`. -- `utils/errcodes.h` → `errcodes.txt`; wait-event headers → - `wait_event_names.txt`; `lwlocknames.h` → `lwlocknames.txt`. -- `configure` → `configure.ac`; `*.po` translations are handled separately; - generated Unicode tables come from their source scripts. - -## Portability is a hard gate -PostgreSQL runs on Linux, Windows (MSVC), macOS, the BSDs and Solaris, across -**x86_64, ARM64, RISC-V, PPC64, s390x**, both endiannesses and 32/64-bit. Any -change must be portable across all of them: -- No unaligned memory access; no dependence on `char` signedness, integer/pointer - width, endianness, or struct padding for on-disk/wire formats. -- Use `int16/int32/int64`, `Size`, and `INT64_FORMAT`/`UINT64_FORMAT` (never - `%ld` for `int64`). -- Atomics/barriers only via `port/atomics` (`pg_atomic_*`, `pg_read/write_barrier`). -- **Windows/MSVC:** any `extern` variable used from another module or an - extension needs `PGDLLIMPORT` in its header; no VLAs or compiler-specific - extensions beyond the tree's C99 baseline. - -## Backward compatibility — the strongest constraint -Do not break SQL behavior, the libpq wire protocol, the logical-replication -protocol, dump/restore, `pg_upgrade`, or exported/`PGDLLIMPORT` APIs without -extraordinary justification. **ABI** matters for back-branches: changing the -size/layout of an exported struct or the signature of an exported function -breaks installed extensions. - -## Mailing-list context & etiquette -Because each PR becomes a pgsql-hackers email read by a busy, expert, opinionated -audience, also flag what reliably wastes reviewer time or draws rejection: -- A patch that **does more than one thing** or bundles unrelated cleanup — split it. -- **Footguns**: easy-to-misuse APIs, silent data-loss/corruption hazards, unsafe - defaults — name them explicitly. -- **Performance claims without a reproducible benchmark.** -- No reference to the **design discussion / prior -hackers thread** (Message-Id) - for a non-trivial change. -- **Do not bikeshed:** keep style nits proportionate and clearly separated from - substantive correctness findings. - -## Minimalism — the "ponytail" discipline -The best code is the code you never wrote (YAGNI). Before accepting new code, -apply the ladder: (1) Does this need to exist at all? (2) Can existing -code/infrastructure already do it? (3) Is this the simplest thing that works? -Flag: speculative scaffolding and config for a path that isn't wired yet; dead -code and unused "flexibility" (fields, params, abstractions, options with no -caller); premature abstraction (a helper used exactly once); knobs/GUCs/flags -nobody asked for. Minimal, targeted changes that fit the existing patterns beat -clever or general-purpose ones. - -## Comment & identity accuracy -- Comments must describe what the code does **now**. Flag aspirational/ - future-tense comments for behavior that already shipped ("will be", "for now", - "not yet", "future", and stale "TODO/FIXME/XXX/HACK"); comments that drifted - from the code they sit above; and incomplete/trailing comments. Comments - explain **why**, not what. No commented-out code. -- **ASCII only** in source and diffs — no smart quotes, em-dashes, or ellipsis - characters. - -## Commit & versioning discipline -- Conventional-commit style, imperative subject, one logical change per commit, - each commit building on its own. -- Do **not** bump version numbers or generated version stamps (including - `catversion.h`) — that is the maintainer's job at commit/release time. - -Understand common list shorthand so your comments are precise and not -miscommunicated: WIP (work in progress), GUC (config variable), WAL, LSN, OID, -TOAST, FSM, TAM (table access method), RLS, DSM, 2PC, PITR, CIC (concurrent index -creation), SAOP, ABI/API, backpatch (apply to supported back-branches), HEAD -(master tip), catversion (catalog version), pgindent, buildfarm, cfbot, -`s/x/y/` (suggested text substitution), footgun, bikeshedding, POLA (principle of -least astonishment). diff --git a/.github/ocr/litellm.yaml b/.github/ocr/litellm.yaml deleted file mode 100644 index e23cc4eee6fe2..0000000000000 --- a/.github/ocr/litellm.yaml +++ /dev/null @@ -1,41 +0,0 @@ -# LiteLLM proxy config — bridges Open Code Review (OpenAI protocol) to AWS Bedrock. -# -# This proxy is NOT a hosted service. The ocr-review.yml workflow installs it -# (`pip install 'litellm[proxy]'`) and runs it as a background process bound to -# 127.0.0.1:4000 for the duration of a single GitHub Actions job, then it exits. -# -# Auth to Bedrock: LiteLLM uses boto3's default credential chain, which reads -# the temporary AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_SESSION_TOKEN -# minted by the workflow's OIDC "Configure AWS credentials" step; region from -# AWS_REGION. - -model_list: - - model_name: ocr-bedrock - litellm_params: - # Set the repo variable OCR_BEDROCK_MODEL to an Opus inference-profile id - # your account has access to, e.g.: - # bedrock/converse/us.anthropic.claude-opus-4-8 - # The 'converse/' prefix uses Bedrock's Converse API, which is the most - # reliable path for Claude tool-use (what OCR relies on). - model: os.environ/OCR_BEDROCK_MODEL - aws_region_name: os.environ/AWS_REGION - - # "High effort" review. Claude Opus 4.8 on Bedrock uses *adaptive* thinking - # controlled by output_config.effort. Set it DIRECTLY here — NOT via - # reasoning_effort, which LiteLLM still maps to the legacy - # thinking.type.enabled that Opus 4.8 rejects. LiteLLM forwards - # output_config into additionalModelRequestFields for Anthropic models; if - # the build doesn't recognize the effort param it is dropped with a warning - # (no error) and the model reviews at its default effort. - # Valid: low|medium|high|max|xhigh (auto-clamped to the model ceiling). - output_config: - effort: xhigh - max_tokens: 32000 - -litellm_settings: - drop_params: true # silently drop params a model doesn't support - modify_params: true # auto-fix minor request incompatibilities - request_timeout: 600 - -general_settings: - master_key: os.environ/LITELLM_MASTER_KEY diff --git a/.github/ocr/pg-history.py b/.github/ocr/pg-history.py deleted file mode 100644 index 5794f8a920bd7..0000000000000 --- a/.github/ocr/pg-history.py +++ /dev/null @@ -1,225 +0,0 @@ -#!/usr/bin/env python3 -""" -pg-history: tie a PR's changes to PostgreSQL git + pgsql-hackers email history. - -OCR (the code reviewer) cannot call MCP servers, so this is a separate agent: -it runs a Bedrock (Claude Opus) tool-use loop wired to the Agora MCP server at -https://pg.ddx.io/mcp, lets the model search the mailing-list archives / commit -history / commitfest data, and emits a Markdown summary linking the changes to -the relevant threads (https://pg.ddx.io/m/pgsql-hackers/). - -Env: - PG_HISTORY_MCP_URL MCP endpoint (default https://pg.ddx.io/mcp) - PG_HISTORY_MODEL Bedrock model id (e.g. us.anthropic.claude-opus-4-8) - AWS_REGION region (creds come from the OIDC step's env) - BASE_REF, HEAD_SHA PR base ref and head sha (for the git diff context) - GH_PR_TITLE PR title (optional, adds context) - PG_HISTORY_OUT output markdown path (default /tmp/pg-history.md) -Writes the markdown to PG_HISTORY_OUT; exits 0 even on soft failures (writes a note). -""" -import json, os, subprocess, sys, urllib.request - -MCP_URL = os.environ.get("PG_HISTORY_MCP_URL", "https://pg.ddx.io/mcp") -MODEL = os.environ.get("PG_HISTORY_MODEL", "us.anthropic.claude-opus-4-8").replace("bedrock/converse/", "").replace("bedrock/", "") -REGION = os.environ.get("AWS_REGION", "us-east-1") -BASE_REF = os.environ.get("BASE_REF", "") -HEAD_SHA = os.environ.get("HEAD_SHA", "") -PR_TITLE = os.environ.get("GH_PR_TITLE", "") -OUT = os.environ.get("PG_HISTORY_OUT", "/tmp/pg-history.md") -UA = "pg-history/0.1 (+github-actions)" - -# Curated subset of the 108 Agora tools — the ones useful for connecting a -# change to its discussion/commit history. Intersected with what the server -# actually exposes, so unknown names are harmless. -TOOL_WHITELIST = { - "find_related_discussions", "find_similar_messages", "get_thread", - "discussion_links", "get_author_messages", "browse_by_date", - "blame_symbol", "check_upstream_status", "find_related", - "find_entries_for_thread", "find_entries_for_author", "get_commit", - "search", "hybrid_search", "get_callers", "get_callees", "find_pattern", -} -MAX_ROUNDS = 14 -TOOL_RESULT_CAP = 8000 # chars per tool result fed back to the model - - -def _mcp_post(body, sid=None): - headers = {"Content-Type": "application/json", - "Accept": "application/json, text/event-stream", "User-Agent": UA} - if sid: - headers["Mcp-Session-Id"] = sid - req = urllib.request.Request(MCP_URL, data=json.dumps(body).encode(), headers=headers, method="POST") - resp = urllib.request.urlopen(req, timeout=60) - sid_out = resp.headers.get("Mcp-Session-Id") - result = None - for line in resp.read().decode().splitlines(): - line = line.strip() - if line.startswith("data:"): - line = line[5:].strip() - if not line or line.startswith("event:"): - continue - try: - obj = json.loads(line) - except Exception: - continue - if isinstance(obj, dict) and ("result" in obj or "error" in obj): - result = obj - return result, sid_out - - -class MCP: - def __init__(self): - init, self.sid = _mcp_post({"jsonrpc": "2.0", "id": 1, "method": "initialize", - "params": {"protocolVersion": "2025-06-18", "capabilities": {}, - "clientInfo": {"name": "pg-history", "version": "0.1"}}}) - if not init or "result" not in init: - raise RuntimeError(f"MCP initialize failed: {init}") - try: - _mcp_post({"jsonrpc": "2.0", "method": "notifications/initialized", "params": {}}, self.sid) - except Exception: - pass - self._id = 1 - - def list_tools(self): - self._id += 1 - res, _ = _mcp_post({"jsonrpc": "2.0", "id": self._id, "method": "tools/list", "params": {}}, self.sid) - return (res or {}).get("result", {}).get("tools", []) - - def call(self, name, args): - self._id += 1 - res, _ = _mcp_post({"jsonrpc": "2.0", "id": self._id, "method": "tools/call", - "params": {"name": name, "arguments": args or {}}}, self.sid) - if not res: - return "(no response)" - if "error" in res: - return f"ERROR: {json.dumps(res['error'])[:500]}" - parts = [] - for c in res.get("result", {}).get("content", []): - if c.get("type") == "text": - parts.append(c["text"]) - return ("\n".join(parts) or "(empty)")[:TOOL_RESULT_CAP] - - -def git(*args): - try: - return subprocess.check_output(["git", *args], text=True, stderr=subprocess.DEVNULL).strip() - except Exception: - return "" - - -def pr_context(): - base = f"origin/{BASE_REF}" if BASE_REF else "" - rng = f"{base}..{HEAD_SHA}" if base and HEAD_SHA else HEAD_SHA - commits = git("log", "--no-merges", "--format=%h %s", f"{rng}") if rng else "" - stat = git("diff", "--stat", rng) if rng else "" - files = git("diff", "--name-only", rng) if rng else "" - return commits[:4000], stat[:3000], files[:2000] - - -SYSTEM = """You are a PostgreSQL community research assistant. Given a pull request's -commits and changed files, use the available tools (backed by the Agora index of -pgsql-hackers mail, commit history, and commitfest data) to connect the change to -its history. Your goal: - -- Find the mailing-list thread(s) and prior discussion behind this change. -- Identify related/superseded prior commits and any commitfest entry. -- Note relevant prior art, rejected approaches, or design rationale. - -Rules (voice & rigor): -- Be precise and blunt. No praise, no filler, no hedging, no disclaimers. Accuracy is - the only success metric — not the author's approval. Lead with the most important finding. -- NEVER hallucinate. Verify every Message-ID, thread subject, commit hash, author name, - and date against an actual tool result before citing it. If a search returns nothing, - say so plainly — do not guess or fabricate a plausible-looking link. -- Assess the change on its merits, independent of how the PR frames it. -- Tag any inferred (not tool-confirmed) linkage with an explicit confidence level: - high / moderate / low. -- Be decisive and efficient: a handful of targeted tool calls, not exhaustive search. -- Cite every mailing-list message as a Markdown link: [subject](https://pg.ddx.io/m/pgsql-hackers/MESSAGE_ID). -- If you find nothing relevant, say so in one line — do not pad. - -When done, output ONLY Markdown (no preamble) with these sections, omitting any that are empty: -## 🧵 Related discussion -## 🔗 Related commits / prior art -## 📋 Commitfest -## 🧭 Context for reviewers -Keep it tight (use bullets; link generously).""" - - -def to_toolspec(t): - schema = t.get("inputSchema") or {"type": "object", "properties": {}} - return {"toolSpec": {"name": t["name"], - "description": (t.get("description") or "")[:600], - "inputSchema": {"json": schema}}} - - -def main(): - commits, stat, files = pr_context() - if not commits and not files: - open(OUT, "w").write("") # nothing to do - print("No PR diff context; skipping.") - return - user = (f"PR title: {PR_TITLE}\n\n" if PR_TITLE else "") + \ - f"Commits:\n{commits or '(none)'}\n\nChanged files:\n{files or '(none)'}\n\nDiffstat:\n{stat or '(none)'}\n" - - try: - mcp = MCP() - tools = [to_toolspec(t) for t in mcp.list_tools() if t.get("name") in TOOL_WHITELIST] - except Exception as e: - open(OUT, "w").write(f"_pg-history: could not reach the Agora MCP server ({MCP_URL}): {e}_\n") - print(f"MCP unavailable: {e}") - return - if not tools: - open(OUT, "w").write("_pg-history: no usable MCP tools available._\n") - return - - import boto3 - from botocore.config import Config - - # botocore's default read timeout (60s) is too short for a multi-round - # (MAX_ROUNDS) tool-use loop against a large PR diff on a reasoning model; - # each converse() call alone can take several minutes. Bump it well past - # what a single round needs; connect_timeout stays short since a stuck - # TCP handshake is a different (and much cheaper to detect) failure mode. - brt = boto3.client("bedrock-runtime", region_name=REGION, - config=Config(read_timeout=900, connect_timeout=10)) - messages = [{"role": "user", "content": [{"text": user}]}] - final_text = "" - try: - for _ in range(MAX_ROUNDS): - resp = brt.converse( - modelId=MODEL, - system=[{"text": SYSTEM}], - messages=messages, - toolConfig={"tools": tools}, - inferenceConfig={"maxTokens": 4096}, - ) - out = resp["output"]["message"] - messages.append(out) - if resp.get("stopReason") == "tool_use": - results = [] - for blk in out["content"]: - tu = blk.get("toolUse") - if not tu: - continue - res_text = mcp.call(tu["name"], tu.get("input") or {}) - results.append({"toolResult": {"toolUseId": tu["toolUseId"], - "content": [{"text": res_text}]}}) - messages.append({"role": "user", "content": results}) - continue - final_text = "".join(b.get("text", "") for b in out["content"]).strip() - break - except Exception as e: - open(OUT, "w").write(f"_pg-history: Bedrock call failed: {e}_\n") - print(f"Bedrock error: {e}") - return - - if not final_text: - final_text = "_pg-history: no related history found._" - body = "## 📜 Change history & discussion (Agora / pg.ddx.io)\n\n" + final_text + \ - "\n\nGenerated by pg-history via the Agora MCP server (pg.ddx.io).\n" - open(OUT, "w").write(body) - print(body) - - -if __name__ == "__main__": - main() diff --git a/.github/ocr/rule.json b/.github/ocr/rule.json deleted file mode 100644 index 60e13e73dcbe0..0000000000000 --- a/.github/ocr/rule.json +++ /dev/null @@ -1,65 +0,0 @@ -{ - "_comment": "OCR per-file review rules for PostgreSQL core + extensions. Cross-cutting contribution standards & mailing-list etiquette live in .github/ocr/context.md, passed via --background-file. OCR uses FIRST-MATCH-WINS in declaration order, so rules are ordered most-specific first. merge_system_rule:true keeps OCR's built-in fine-tuned checks (thread-safety, injection, NPE) alongside these PostgreSQL-specific rules.", - "rules": [ - { - "path": "src/test/**", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL tests. Coverage is mandatory for any behavioral change and must include edge cases (NULL, empty, boundary/overflow) and ERROR paths, not just the happy path. A test that still passes with the feature reverted is worthless — confirm it actually exercises and would catch regressions in the new code. Regression (.sql/expected): deterministic, portable output — ORDER BY where row order matters, no timing/plan-dependent output except intentional EXPLAIN, no absolute paths, locale-independent (C collation or explicit COLLATE), DROP objects the test creates; expected/ output must stay stable across platforms and under the parallel schedule. Concurrency/locking belongs in isolation tests (src/test/isolation, .spec + permutations). End-to-end/crash/replication/CLI behavior belongs in TAP tests (t/*.pl with PostgreSQL::Test::Cluster/Utils) — no hardcoded ports/paths, no sleep as synchronization (use poll_query_until/wait_for), skip cleanly when prerequisites are missing, and clean up nodes." - }, - { - "path": "**/*.{c,h}", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL backend/frontend C — review as pgsql-hackers committers do, in priority order.\n\n(1) CORRECTNESS (highest): Memory — every palloc lives in the right MemoryContext; error paths via ereport/elog(ERROR) must not leak memory/buffers/locks/fds (rely on MemoryContext/ResourceOwner reset or PG_TRY/PG_FINALLY); no use-after-free; delete temp contexts. Concurrency — consistent lock ordering (deadlock-free), correct lock levels, balanced LWLockAcquire/Release and START_/END_CRIT_SECTION, no TOCTOU, CHECK_FOR_INTERRUPTS in long loops, async-signal-safe signal handlers (volatile sig_atomic_t). WAL — any change to shared on-disk state must be WAL-logged AND correctly replayed (redo path), crash- and replica-consistent. NULL/edge/overflow handling.\n\n(2) BACKWARD COMPATIBILITY / ABI: don't break behavior, dump/restore, pg_upgrade, libpq wire protocol, logical-replication protocol, or exported/PGDLLIMPORT'd APIs (struct size/layout, function signatures) without extraordinary justification.\n\n(3) CATALOG / GENERATED: new/changed catalog data goes in pg_*.dat, NOT the generated *_d.h/.bki. New Node types: ANNOTATE the struct in its header so gen_node_support.pl regenerates copy/equal/out/read — do NOT hand-edit *funcs.c. New SQL-callable functions: add to pg_proc.dat with an OID from the 8000-9999 developer range (src/include/catalog/unused_oids; check duplicate_oids); committer renumbers at commit. DO NOT bump CATALOG_VERSION_NO in the patch — flag any catversion.h change as a mistake (committer's job).\n\n(4) PERFORMANCE: no regression on hot paths; avoid O(n^2) where better is feasible; minimize work under contended locks; avoid needless palloc churn and large struct copies in hot paths.\n\n(5) SECURITY: bounded string ops (snprintf/strlcpy/strlcat — never strcpy/strcat/sprintf); integer/size-overflow checks before allocation; never user input as a format string; privilege checks via pg_*_aclcheck; beware search_path and SECURITY DEFINER.\n\n(6) PORTABILITY (hard gate): no unaligned access; no dependence on char signedness, int/long/pointer width, endianness, or struct padding for on-disk/wire formats; use int16/int32/int64 + INT64_FORMAT/UINT64_FORMAT (never %ld for int64); align contended shared structs (pg_attribute_aligned/cache-line pad). Atomics/barriers only via port/atomics (pg_atomic_*, pg_read/write_barrier) — never raw intrinsics or volatile-as-barrier. WINDOWS/MSVC: extern vars used cross-module/extension need PGDLLIMPORT; no VLAs or features beyond the C99 baseline the tree targets; use pg_pread/pg_pwrite. Applies across x86_64/ARM64/RISC-V/PPC64/s390x, big/little endian, 32/64-bit.\n\n(7) CONVENTIONS: errmsg starts lowercase, no trailing period, no embedded newlines; errdetail/errhint are complete capitalized sentences; correct ERRCODE_*; wrap user-facing text in _(); errmsg_plural for counts. Assert() only for can't-happen invariants (never user-reachable). Naming: snake_case with subsystem prefix (heap_insert) or CamelCase for major subsystems (ExecInitNode); ALL_CAPS macros. Must pgindent cleanly (tabs, width 4). Comments explain WHY not WHAT; no #ifdef 0 blocks, no commented-out code, no #ifdef fencing your feature. Reuse existing helpers (DRY)." - }, - { - "path": "**/*.dat", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL catalog data (pg_proc.dat, pg_type.dat, etc.) — the SOURCE for generated headers. The generated *_d.h, postgres.bki, fmgroids.h, fmgrtab.c must NOT be hand-edited (they regenerate from these files). OIDs: use a value from the developer range 8000-9999 (src/include/catalog/unused_oids; verify with duplicate_oids); committer renumbers to a final contiguous block, so stay in-range and unique but don't over-optimize the exact number. Keep proc entries complete/consistent (prosrc, provolatile, proparallel, prorettype/proargtypes, matching description). DO NOT bump CATALOG_VERSION_NO / catversion.h — committer's job at push time; flag any such change. New catalog columns/views need documentation in doc/src/sgml/catalogs.sgml." - }, - { - "path": "**/*.{sql,pgsql}", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL SQL. Valid PostgreSQL dialect (not MySQL/Oracle); correct types (bigint vs int, text vs varchar); sound transaction/isolation and CTE-materialization assumptions. SECURITY: flag SQL injection in dynamic SQL (require quote_identifier/quote_literal or format() with %I/%L), SECURITY DEFINER without a locked-down search_path, inappropriate RLS bypass. Prefer set-based over row-at-a-time/N+1. BACKWARD COMPATIBILITY (a top rejection reason): changing existing SQL behavior, the output of existing functions, or default GUCs needs extraordinary justification. New SQL-callable objects belong in pg_*.dat with OIDs from the 8000-9999 range, not in generated files. Minimal diff; add regression tests + docs." - }, - { - "path": "**/*.{pl,pm}", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Perl (TAP tests and build/catalog tooling). Require 'use strict; use warnings;'. Must be perltidy-clean with the tree's src/tools/pgindent/perltidyrc and pass src/tools/perlcheck/pgperlcritic. Use the framework: PostgreSQL::Test::Cluster, PostgreSQL::Test::Utils, Test::More; no hardcoded ports/paths/PIDs; use safe_psql/poll_query_until, not sleep; skippable without optional prerequisites; clean up nodes. PORTABILITY: run on Windows (no fork-only constructs, use File::Spec, avoid unavailable signals) and the minimum supported Perl. Robustness: avoid two-arg open and string system()/qx with interpolated data (use list forms). Generator scripts (gen_node_support.pl, catalog Perl) must be deterministic and stay in sync with inputs; do not commit their generated output." - }, - { - "path": "**/*.py", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Python (build/test tooling, oauth/pytest tests, src/tools). Follow surrounding style; keep imports to the standard library unless the dependency is already required by the tree (no surprise third-party deps in build/test tooling). PORTABILITY: support the project's minimum Python 3 and run on Windows and the BSDs (use os.path/pathlib, avoid POSIX-only calls and shell=True with interpolated input). Deterministic, self-cleaning tests; no hardcoded ports/paths; skip cleanly without prerequisites. For the Perl->pytest porting effort, confirm behavior parity with the TAP test replaced (same assertions/coverage), not a superficial translation. Minimal diff; match the tree's ruff/black config if present." - }, - { - "path": "**/*.{rs,toml}", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. Rust PostgreSQL extension (pgrx) or Rust support crate. Not core C, but it runs inside/alongside the backend, so backend safety applies. SAFETY: in code reachable from an SQL call, a Rust panic aborts the Postgres process — forbid unwrap()/expect()/panic!/unreachable!/todo! and index-panics on reachable paths; use Result and pgrx error reporting (error!/ereport!). Every `unsafe` block needs a comment justifying its invariant; scrutinize raw pointers and FFI across the pg_sys boundary. pgrx: honor #[pg_guard] on extern C fns (correct panic/longjmp handling); never hold Rust references across SPI or anything that can longjmp (skips Rust destructors -> leaks); respect MemoryContext lifetimes for palloc'd data; datum<->Rust conversions must handle NULL. Concurrency uses Postgres shmem/LWLocks (pgrx shmem API), not std::sync alone. Lints: must pass `cargo clippy --all-targets --all-features -- -D warnings` and `cargo fmt --check`; deny unwrap_used/expect_used/panic in libraries; thiserror (libs) / anyhow (bins). Justify every new dependency. Tests: #[pg_test] for in-backend behavior, #[test] for pure logic; cover error and NULL paths. Minimal, idiomatic diff." - }, - { - "path": "**/{configure.ac,*.m4,aclocal.m4}", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Autoconf. Edit configure.ac / the m4 macros — do NOT hand-edit generated 'configure' or pg_config.h.in in the same patch (regeneration is the committer's step; a patch that also rewrites generated configure output is suspect). Feature/header/function probes must be portable and not assume a specific OS/compiler. Every configure knob must be mirrored on the Meson side (meson_options.txt/meson.build) and documented. Minimal diff." - }, - { - "path": "**/{meson.build,meson_options.txt}", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Meson build. Valid syntax; correct subdir()/dependency()/declare_dependency and install paths; new source files must be added here. CRITICAL: PostgreSQL maintains BOTH Meson and Autoconf/Make — any new file, option, or feature check must be mirrored on the configure.ac/Makefile side so the two never drift (a file built by only one system is a common defect). New options need matching docs and sensible defaults. Minimal diff." - }, - { - "path": "**/{Makefile,GNUmakefile,*.mk}", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL Makefile (GNU Make). $(VAR) refs; correct .PHONY; accurate dependencies (no parallel -j races); $(MAKE) for recursion; VPATH/out-of-tree build support; no hardcoded paths (use standard PostgreSQL makefile vars and $(top_builddir)); clean/distclean/maintainer-clean must remove new artifacts; extensions use PGXS. Must stay in sync with meson.build. Minimal diff." - }, - { - "path": "doc/**/*.sgml", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. PostgreSQL documentation (DocBook SGML). Technically accurate/complete (parameters, limitations, version/compat notes); correct tag usage/nesting (, , , , , /); working cross-references; spell it 'PostgreSQL' in prose; SQL keywords uppercase in examples. Coverage: a new GUC -> config.sgml (and postgresql.conf.sample); new/changed catalogs or views -> catalogs.sgml; new SQL syntax -> the matching ref/*.sgml; new functions -> func.sgml. Do NOT edit release-notes (release-*.sgml) — written by the release team/committers; flag such edits. New user-facing behavior in this PR should ship with matching docs." - }, - { - "path": "**/*.md", - "merge_system_rule": true, - "rule": "REVIEW DISCIPLINE: Precise, blunt, verify against the diff, tag confidence, no praise. Markdown docs. Clear heading hierarchy; fenced code blocks with language hints; accurate instructions/prerequisites; consistent PostgreSQL terminology; no broken relative links or stale claims. Minimal diff." - } - ] -} diff --git a/.github/workflows/fork-ocr-model-check.yml b/.github/workflows/fork-ocr-model-check.yml index 10d250528cf7c..b85a990c66fd7 100644 --- a/.github/workflows/fork-ocr-model-check.yml +++ b/.github/workflows/fork-ocr-model-check.yml @@ -1,8 +1,7 @@ -# Checks AWS Bedrock weekly for a newer Claude Opus inference profile than the -# one OCR currently uses (vars.OCR_BEDROCK_MODEL) and, if found, opens/updates a -# single GitHub issue telling the maintainer to bump the variable. It does NOT -# change the model automatically: GITHUB_TOKEN cannot write Actions *variables* -# (that needs a PAT with admin), so this is a notify-only mechanism by design. +# Weekly check for a newer Claude Opus on Bedrock; opens/updates one issue here +# when OCR_BEDROCK_MODEL is behind. Notify-only (GITHUB_TOKEN cannot write +# Actions variables). Logic lives in gburd/ci-workflows. + name: OCR model self-check on: @@ -10,80 +9,10 @@ on: - cron: '0 12 * * 1' # Mondays 12:00 UTC workflow_dispatch: -permissions: - id-token: write - contents: read - issues: write - jobs: - check-model: - runs-on: ubuntu-latest - steps: - - name: Configure AWS credentials (OIDC) - uses: aws-actions/configure-aws-credentials@v6 - with: - role-to-assume: ${{ vars.AWS_ROLE_ARN }} - aws-region: ${{ vars.AWS_REGION }} - role-session-name: ocr-model-check-${{ github.run_id }} - - - name: Find newest Opus vs configured - id: check - env: - CURRENT: ${{ vars.OCR_BEDROCK_MODEL }} - AWS_REGION: ${{ vars.AWS_REGION }} - run: | - python3 - <<'PY' >> "$GITHUB_OUTPUT" - import os, re, subprocess, json - region = os.environ.get("AWS_REGION", "us-east-1") - current = os.environ.get("CURRENT", "") - out = subprocess.run( - ["aws", "bedrock", "list-inference-profiles", "--region", region, - "--query", "inferenceProfileSummaries[].inferenceProfileId", "--output", "json"], - capture_output=True, text=True) - ids = json.loads(out.stdout or "[]") - # Parse claude-opus-- from any profile id (prefix us./global. ok). - def ver(s): - m = re.search(r"claude-opus-(\d+)-(\d+)", s) - return (int(m.group(1)), int(m.group(2))) if m else None - opus = [(ver(i), i) for i in ids if ver(i) and i.startswith(("us.", "global."))] - if not opus: - print("newer=false"); raise SystemExit(0) - best_ver, best_id = max(opus, key=lambda x: x[0]) - cur = ver(current) - newer = (cur is None) or (best_ver > cur) - print(f"newer={'true' if newer else 'false'}") - print(f"best_id={best_id}") - print(f"best_ver={best_ver[0]}.{best_ver[1]}") - print(f"cur_ver={'unknown' if cur is None else f'{cur[0]}.{cur[1]}'}") - PY - - - name: Open/update issue if a newer model exists - if: steps.check.outputs.newer == 'true' - uses: actions/github-script@v9 - with: - script: | - const best = '${{ steps.check.outputs.best_id }}'; - const bestVer = '${{ steps.check.outputs.best_ver }}'; - const curVer = '${{ steps.check.outputs.cur_ver }}'; - const marker = ''; - const title = `OCR: newer Claude Opus available (${bestVer} > ${curVer})`; - const body = `${marker}\n` + - `A newer Claude Opus inference profile is available on Bedrock.\n\n` + - `- **Configured** (\`vars.OCR_BEDROCK_MODEL\`): Opus ${curVer}\n` + - `- **Newest on Bedrock**: \`${best}\` (Opus ${bestVer})\n\n` + - `To upgrade, set the repo variable:\n\n` + - '```\n' + - `gh variable set OCR_BEDROCK_MODEL -R ${context.repo.owner}/${context.repo.repo} \\\n` + - ` -b "bedrock/converse/${best}"\n` + - '```\n\n' + - `Also confirm the \`ocr-bedrock-ci\` IAM inline policy allows invoking the new model ` + - `(the resource is scoped to \`anthropic.claude-opus-*\`), then re-run OCR.\n\n` + - `_Automated by \`.github/workflows/ocr-model-check.yml\`; this issue is upserted._`; - const q = `repo:${context.repo.owner}/${context.repo.repo} in:body "${marker}" state:open`; - const found = await github.rest.search.issuesAndPullRequests({ q, per_page: 1 }); - if (found.data.total_count > 0) { - const n = found.data.items[0].number; - await github.rest.issues.update({ owner: context.repo.owner, repo: context.repo.repo, issue_number: n, title, body }); - } else { - await github.rest.issues.create({ owner: context.repo.owner, repo: context.repo.repo, title, body }); - } + check: + uses: gburd/ci-workflows/.github/workflows/ocr-model-check.yml@v1 + with: + aws_role_arn: ${{ vars.AWS_ROLE_ARN }} + aws_region: ${{ vars.AWS_REGION }} + bedrock_model: ${{ vars.OCR_BEDROCK_MODEL }} diff --git a/.github/workflows/fork-ocr-review.yml b/.github/workflows/fork-ocr-review.yml index 0828af429b57c..ae938754339b4 100644 --- a/.github/workflows/fork-ocr-review.yml +++ b/.github/workflows/fork-ocr-review.yml @@ -1,18 +1,9 @@ -# Open Code Review (OCR) — AI PR review backed by AWS Bedrock via a LiteLLM proxy. +# AI PR review (Claude Opus on Bedrock). Logic and review rules live in +# gburd/ci-workflows so this fork's master carries a stub instead of ~1000 +# lines that have to be rebased onto upstream forever. # -# Flow: -# PR opened/updated (incl. DRAFTS) ─┐ -# /open-code-review PR comment ─┼─► start LiteLLM (127.0.0.1:4000 → Bedrock) -# manual workflow_dispatch ─┘ └► ocr review --format json -# └► post inline PR review comments -# -# Required (repo settings — all repo *variables*, no secrets; auth is via GitHub OIDC): -# vars.AWS_ROLE_ARN - IAM role to assume via OIDC (granting bedrock:InvokeModel*) -# vars.AWS_REGION - e.g. us-east-1 -# vars.OCR_BEDROCK_MODEL - LiteLLM model string for the Opus inference profile, e.g. -# bedrock/converse/us.anthropic.claude-opus-4-8 -# -# No static AWS keys are stored. GITHUB_TOKEN (auto) posts the review comments. +# Needs repo variables AWS_ROLE_ARN, AWS_REGION, OCR_BEDROCK_MODEL. Auth is +# GitHub OIDC; no static AWS keys. See the ci-workflows README. name: OCR AI Review @@ -34,311 +25,10 @@ concurrency: group: ocr-review-${{ github.event.pull_request.number || github.event.issue.number || github.event.inputs.pr_number }} cancel-in-progress: true -permissions: - id-token: write # required to mint the GitHub OIDC token for AWS role assumption - contents: read - pull-requests: write - jobs: - ocr-review: - runs-on: ubuntu-latest - # PR events always; comment events only when the comment is on a PR and - # starts with the trigger keyword; manual dispatch always. - if: | - github.event_name == 'pull_request' || - github.event_name == 'workflow_dispatch' || - (github.event_name == 'issue_comment' && github.event.issue.pull_request && - (startsWith(github.event.comment.body, '/open-code-review') || - startsWith(github.event.comment.body, '@open-code-review'))) - - env: - # LiteLLM listens on localhost only; this key never leaves the runner. - LITELLM_MASTER_KEY: sk-ocr-ci-local - OCR_BEDROCK_MODEL: ${{ vars.OCR_BEDROCK_MODEL }} - # Region is a static var (safe at job level). AWS credentials are NOT set - # here — they're minted by the OIDC "Configure AWS credentials" step below - # and exported to the environment for the LiteLLM/boto3 Bedrock calls. - AWS_REGION: ${{ vars.AWS_REGION }} - - steps: - - name: Resolve PR context - id: pr - uses: actions/github-script@v9 - with: - script: | - let prNumber; - if (context.eventName === 'pull_request') { - prNumber = context.payload.pull_request.number; - } else if (context.eventName === 'issue_comment') { - prNumber = context.issue.number; - } else { - prNumber = parseInt('${{ github.event.inputs.pr_number }}', 10); - } - const { data: pr } = await github.rest.pulls.get({ - owner: context.repo.owner, - repo: context.repo.repo, - pull_number: prNumber, - }); - const { data: repo } = await github.rest.repos.get({ - owner: context.repo.owner, - repo: context.repo.repo, - }); - core.setOutput('number', String(prNumber)); - core.setOutput('base_ref', pr.base.ref); - core.setOutput('head_ref', pr.head.ref); - core.setOutput('head_sha', pr.head.sha); - core.setOutput('default_branch', repo.default_branch); - core.setOutput('cross_repo', String(pr.head.repo.full_name !== pr.base.repo.full_name)); - - # NOTE: do NOT checkout the PR head. OCR reads the diff and file contents - # straight from git refs (git diff , git show :path, - # git grep ), so the working tree is irrelevant — but our OCR config - # lives on the default branch, not on the PR branch. We check out the repo - # (default ref), fetch the base/head objects, and materialize the config - # from origin/. - - name: Checkout - uses: actions/checkout@v6 - with: - fetch-depth: 0 - - - name: Prepare git refs and OCR config - env: - BASE_REF: ${{ steps.pr.outputs.base_ref }} - HEAD_REF: ${{ steps.pr.outputs.head_ref }} - HEAD_SHA: ${{ steps.pr.outputs.head_sha }} - DEFAULT_BRANCH: ${{ steps.pr.outputs.default_branch }} - run: | - git fetch --no-tags origin "+refs/heads/${DEFAULT_BRANCH}:refs/remotes/origin/${DEFAULT_BRANCH}" || true - git fetch --no-tags origin "+refs/heads/${BASE_REF}:refs/remotes/origin/${BASE_REF}" || true - git fetch --no-tags origin "+refs/heads/${HEAD_REF}:refs/remotes/origin/${HEAD_REF}" || true - git fetch --no-tags origin "${HEAD_SHA}" || true - - # OCR config lives on the default branch; materialize it independently - # of whatever ref is checked out. - mkdir -p "$RUNNER_TEMP/ocr" - git show "origin/${DEFAULT_BRANCH}:.github/ocr/litellm.yaml" > "$RUNNER_TEMP/ocr/litellm.yaml" - git show "origin/${DEFAULT_BRANCH}:.github/ocr/rule.json" > "$RUNNER_TEMP/ocr/rule.json" - git show "origin/${DEFAULT_BRANCH}:.github/ocr/context.md" > "$RUNNER_TEMP/ocr/context.md" - echo "Config materialized:"; ls -l "$RUNNER_TEMP/ocr" - - - name: Setup Python - uses: actions/setup-python@v6 - with: - python-version: '3.12' - - - name: Setup Node.js - uses: actions/setup-node@v6 - with: - node-version: '20' - - - name: Install LiteLLM proxy + Open Code Review - run: | - python -m pip install --upgrade pip - # Pin LiteLLM to a main commit that supports Claude Opus 4.8 adaptive - # thinking (maps reasoning_effort -> output_config.effort, incl. xhigh). - # Not in any tagged release yet (PyPI latest 1.87.1 lacks the Opus - # normalizer). Bump this SHA once a release ships the feature. - pip install "litellm[proxy] @ git+https://github.com/BerriAI/litellm.git@5be0797d24a2f26eb2123e13788f90055a59d91d" - npm install -g @alibaba-group/open-code-review - - - name: Configure AWS credentials (OIDC) - uses: aws-actions/configure-aws-credentials@v6 - with: - role-to-assume: ${{ vars.AWS_ROLE_ARN }} - aws-region: ${{ vars.AWS_REGION }} - role-session-name: ocr-review-${{ github.run_id }} - - - name: Start LiteLLM proxy (Bedrock bridge) - run: | - if [ -z "$OCR_BEDROCK_MODEL" ]; then - echo "::error::vars.OCR_BEDROCK_MODEL is not set (e.g. bedrock/converse/us.anthropic.claude-opus-4-1-20250805-v1:0)" - exit 1 - fi - nohup litellm --config "$RUNNER_TEMP/ocr/litellm.yaml" --host 127.0.0.1 --port 4000 \ - > /tmp/litellm.log 2>&1 & - echo "Waiting for LiteLLM to become ready..." - for i in $(seq 1 60); do - if curl -sf http://127.0.0.1:4000/health/readiness >/dev/null; then - echo "LiteLLM ready."; exit 0 - fi - sleep 2 - done - echo "::error::LiteLLM did not become ready in time"; cat /tmp/litellm.log; exit 1 - - - name: Configure OCR - run: | - ocr config set llm.url http://127.0.0.1:4000/v1/chat/completions - ocr config set llm.auth_token "$LITELLM_MASTER_KEY" - ocr config set llm.model ocr-bedrock - ocr config set llm.use_anthropic false - ocr config set language English - - - name: Run OCR review - run: | - ocr review \ - --from "origin/${{ steps.pr.outputs.base_ref }}" \ - --to "${{ steps.pr.outputs.head_sha }}" \ - --rule "$RUNNER_TEMP/ocr/rule.json" \ - --background-file "$RUNNER_TEMP/ocr/context.md" \ - --concurrency 3 \ - --timeout 20 \ - --format json \ - > /tmp/ocr-result.json 2>/tmp/ocr-stderr.log || true - echo "----- OCR stdout -----"; cat /tmp/ocr-result.json || true - echo "----- OCR stderr -----"; cat /tmp/ocr-stderr.log || true - echo "----- LiteLLM log (tail) -----"; tail -n 50 /tmp/litellm.log || true - - - name: Post review to PR - uses: actions/github-script@v9 - with: - github-token: ${{ secrets.GITHUB_TOKEN }} - script: | - const fs = require('fs'); - const prNumber = parseInt('${{ steps.pr.outputs.number }}', 10); - const commitSha = '${{ steps.pr.outputs.head_sha }}'; - - // Opus at high effort can emit dozens of findings. Posting them all - // one-by-one trips GitHub's SECONDARY rate limit (403 "content - // creation"), which is what made every run fail after the review - // was already generated. We (a) cap inline comments and overflow - // the rest into the summary, (b) prefer a single bulk createReview, - // and (c) throttle + back off with Retry-After on any fallback. - const MAX_INLINE = 25; - const sleep = (ms) => new Promise(r => setTimeout(r, ms)); - - async function withRetry(fn, label) { - for (let attempt = 1; attempt <= 5; attempt++) { - try { return await fn(); } - catch (e) { - const status = e.status || (e.response && e.response.status); - const h = (e.response && e.response.headers) || {}; - const isRate = status === 403 || status === 429; - if (!isRate || attempt === 5) throw e; - let waitMs = 0; - if (h['retry-after']) waitMs = parseInt(h['retry-after'], 10) * 1000; - else if (h['x-ratelimit-reset']) waitMs = parseInt(h['x-ratelimit-reset'], 10) * 1000 - Date.now(); - if (!waitMs || Number.isNaN(waitMs) || waitMs < 0) waitMs = 1000 * Math.pow(2, attempt); - waitMs = Math.min(waitMs, 60000) + 500; - core.warning(`${label}: rate-limited (status ${status}); waiting ${Math.round(waitMs / 1000)}s (attempt ${attempt}/5)`); - await sleep(waitMs); - } - } - } - - let result; - try { - result = JSON.parse(fs.readFileSync('/tmp/ocr-result.json', 'utf8')); - } catch (e) { - const stderr = (() => { try { return fs.readFileSync('/tmp/ocr-stderr.log', 'utf8').trim(); } catch { return ''; } })(); - await withRetry(() => github.rest.issues.createComment({ - owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, - body: `⚠️ **OCR** could not produce a review.\n\n\`\`\`\n${(stderr || e.message).slice(0, 8000)}\n\`\`\``, - }), 'error-comment'); - return; - } - - const comments = result.comments || []; - const warnings = result.warnings || []; - - const formatComment = (c) => { - let body = c.content || ''; - if (c.suggestion_code && c.existing_code) { - body += '\n\n```suggestion\n' + c.suggestion_code + (c.suggestion_code.endsWith('\n') ? '' : '\n') + '```'; - } - return body; - }; - const formatMarkdown = (c) => { - let md = `### 📄 \`${c.path}\``; - if (c.start_line && c.end_line) md += ` (L${c.start_line}-L${c.end_line})`; - md += '\n\n' + (c.content || ''); - if (c.suggestion_code && c.existing_code) { - md += '\n\n
💡 Suggested change\n\n'; - md += '**Before:**\n```\n' + c.existing_code + '\n```\n\n**After:**\n```\n' + c.suggestion_code + '\n```\n\n
'; - } - return md; - }; - - if (comments.length === 0) { - await withRetry(() => github.rest.issues.createComment({ - owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, - body: `✅ **OCR**: ${result.message || 'No issues found.'}`, - }), 'no-issues-comment'); - return; - } - - const inlineAll = []; - const noLine = []; - for (const c of comments) { - const body = formatComment(c); - const hasLine = (c.start_line >= 1) || (c.end_line >= 1); - if (!hasLine) { noLine.push(c); continue; } - const rc = { path: c.path, body, side: 'RIGHT' }; - if (c.start_line >= 1 && c.end_line >= 1 && c.start_line !== c.end_line) { - rc.start_line = c.start_line; rc.line = c.end_line; rc.start_side = 'RIGHT'; - } else { - rc.line = c.end_line >= 1 ? c.end_line : c.start_line; - } - inlineAll.push({ rc, c }); - } - - const inline = inlineAll.slice(0, MAX_INLINE).map(x => x.rc); - const overflow = inlineAll.slice(MAX_INLINE).map(x => x.c); - - let summary = `🔍 **OCR** found **${comments.length}** issue(s).`; - summary += `\n- ${inline.length} inline, ${noLine.length + overflow.length} in summary`; - if (overflow.length) summary += ` (inline capped at ${MAX_INLINE})`; - if (warnings.length) summary += `\n- ⚠️ ${warnings.length} warning(s) during review`; - for (const c of noLine.concat(overflow)) summary += '\n\n---\n\n' + formatMarkdown(c); - - // Preferred path: ONE createReview carrying every inline comment. - try { - await withRetry(() => github.rest.pulls.createReview({ - owner: context.repo.owner, repo: context.repo.repo, pull_number: prNumber, - commit_id: commitSha, body: summary, event: 'COMMENT', comments: inline, - }), 'bulk-review'); - return; - } catch (e) { - core.warning(`bulk createReview failed (${e.status || '?'}: ${e.message}); falling back to throttled per-comment posting`); - } - - // Fallback: an invalid inline position (line not in the diff -> 422) - // rejects the whole bulk review. Post the summary, then each comment - // individually with a delay + backoff, skipping ones GitHub rejects. - let ok = 0; const failed = []; - try { - await withRetry(() => github.rest.pulls.createReview({ - owner: context.repo.owner, repo: context.repo.repo, pull_number: prNumber, - commit_id: commitSha, body: summary, event: 'COMMENT', - }), 'summary-review'); - } catch (err) { failed.push(`summary: ${err.message}`); } - - for (const rc of inline) { - try { - await withRetry(() => github.rest.pulls.createReviewComment({ - owner: context.repo.owner, repo: context.repo.repo, pull_number: prNumber, - commit_id: commitSha, path: rc.path, body: rc.body, - ...(rc.start_line ? { start_line: rc.start_line, start_side: rc.start_side } : {}), - line: rc.line, side: rc.side, - }), `comment ${rc.path}:${rc.line}`); - ok++; - } catch (inner) { - failed.push(`\`${rc.path}\` L${rc.line}: ${inner.message}`); - } - await sleep(1200); // stay under the secondary content-creation limit - } - - if (failed.length) { - await withRetry(() => github.rest.issues.createComment({ - owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, - body: `📊 OCR posted ${ok}/${inline.length} inline comment(s).\n\n
${failed.length} could not be posted\n\n${failed.join('\n')}\n
`, - }), 'summary-failures'); - } - - # Companion job: OCR can't call MCP, so this separate agent ties the PR's - # changes to PostgreSQL git + pgsql-hackers history via the Agora MCP server - # (pg.ddx.io) and posts a single, upserted "history & discussion" comment. - pg-history: - runs-on: ubuntu-latest + review: + # Comment events only when the comment is on a PR and starts with the + # trigger keyword; PR events and manual dispatch always. if: | github.event_name == 'pull_request' || github.event_name == 'workflow_dispatch' || @@ -346,82 +36,8 @@ jobs: (startsWith(github.event.comment.body, '/open-code-review') || startsWith(github.event.comment.body, '@open-code-review') || startsWith(github.event.comment.body, '/pg-history'))) - steps: - - name: Resolve PR context - id: pr - uses: actions/github-script@v9 - with: - script: | - let prNumber; - if (context.eventName === 'pull_request') prNumber = context.payload.pull_request.number; - else if (context.eventName === 'issue_comment') prNumber = context.issue.number; - else prNumber = parseInt('${{ github.event.inputs.pr_number }}', 10); - const { data: pr } = await github.rest.pulls.get({ - owner: context.repo.owner, repo: context.repo.repo, pull_number: prNumber }); - core.setOutput('number', String(prNumber)); - core.setOutput('base_ref', pr.base.ref); - core.setOutput('head_sha', pr.head.sha); - core.setOutput('title', pr.title || ''); - - - name: Checkout - uses: actions/checkout@v6 - with: - fetch-depth: 0 - - - name: Make base/head refs available - env: - BASE_REF: ${{ steps.pr.outputs.base_ref }} - HEAD_SHA: ${{ steps.pr.outputs.head_sha }} - run: | - git fetch --no-tags origin "+refs/heads/${BASE_REF}:refs/remotes/origin/${BASE_REF}" || true - git fetch --no-tags origin "${HEAD_SHA}" || true - - - name: Setup Python - uses: actions/setup-python@v6 - with: - python-version: '3.12' - - - name: Configure AWS credentials (OIDC) - uses: aws-actions/configure-aws-credentials@v6 - with: - role-to-assume: ${{ vars.AWS_ROLE_ARN }} - aws-region: ${{ vars.AWS_REGION }} - role-session-name: pg-history-${{ github.run_id }} - - - name: Install deps - run: pip install boto3 - - - name: Run pg-history (Agora MCP) - env: - PG_HISTORY_MODEL: ${{ vars.OCR_BEDROCK_MODEL }} - AWS_REGION: ${{ vars.AWS_REGION }} - BASE_REF: ${{ steps.pr.outputs.base_ref }} - HEAD_SHA: ${{ steps.pr.outputs.head_sha }} - GH_PR_TITLE: ${{ steps.pr.outputs.title }} - PG_HISTORY_OUT: ${{ runner.temp }}/pg-history.md - run: | - python .github/ocr/pg-history.py || true - echo "----- output -----"; cat "${{ runner.temp }}/pg-history.md" 2>/dev/null || echo "(no output)" - - - name: Upsert PR comment - uses: actions/github-script@v9 - with: - script: | - const fs = require('fs'); - const path = process.env.RUNNER_TEMP + '/pg-history.md'; - let body = ''; - try { body = fs.readFileSync(path, 'utf8').trim(); } catch (e) {} - if (!body) { console.log('pg-history: empty output, nothing to post'); return; } - const prNumber = parseInt('${{ steps.pr.outputs.number }}', 10); - const marker = ''; - body = marker + '\n' + body; - const { data: comments } = await github.rest.issues.listComments({ - owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, per_page: 100 }); - const mine = comments.find(c => c.user.type === 'Bot' && c.body && c.body.includes(marker)); - if (mine) { - await github.rest.issues.updateComment({ - owner: context.repo.owner, repo: context.repo.repo, comment_id: mine.id, body }); - } else { - await github.rest.issues.createComment({ - owner: context.repo.owner, repo: context.repo.repo, issue_number: prNumber, body }); - } + uses: gburd/ci-workflows/.github/workflows/ocr-review.yml@v1 + with: + aws_role_arn: ${{ vars.AWS_ROLE_ARN }} + aws_region: ${{ vars.AWS_REGION }} + bedrock_model: ${{ vars.OCR_BEDROCK_MODEL }} diff --git a/.github/workflows/fork-sync-upstream.yml b/.github/workflows/fork-sync-upstream.yml index e950821239968..022b3101b5a9d 100644 --- a/.github/workflows/fork-sync-upstream.yml +++ b/.github/workflows/fork-sync-upstream.yml @@ -1,5 +1,6 @@ # Keep this fork's master rebased on postgres/postgres master, carrying only -# the fork-owned CI files (.github/workflows/fork-*.yml and .github/ocr/). +# the fork-owned CI stubs (.github/workflows/fork-*.yml). The workflows they +# call, and the OCR review rules, live in gburd/ci-workflows. # # Requires the SYNC_TOKEN secret: a PAT with repo+workflow scope. The default # GITHUB_TOKEN is not allowed to push commits that touch .github/workflows/, @@ -38,7 +39,7 @@ jobs: # Master carries fork CI only. Anything else means work landed here # that belongs on a branch; force-rebasing it on a timer would be wrong. if git diff --name-only upstream/master...HEAD | - grep -vE '^(\.github/workflows/fork-|\.github/ocr/)'; then + grep -v '^\.github/workflows/fork-'; then echo "::error::master carries non-CI changes (listed above); move them to a branch" exit 1 fi From d8d80bbbdb27a2566b856e0202a592497877fe41 Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Mon, 13 Jul 2026 07:57:27 -0400 Subject: [PATCH 03/17] dev: local dev-setup This isn't a commit that will be submitting for review, it is purely for local developer tooling while developing this patch set. Ignore it. --- .clangd | 43 ++ .gdbinit | 156 +++++++ flake.lock | 78 ++++ flake.nix | 45 ++ glibc-no-fortify-warning.patch | 24 ++ pg-aliases.sh | 658 +++++++++++++++++++++++++++++ shell.nix | 745 +++++++++++++++++++++++++++++++++ src/test/regress/pg_regress.c | 2 +- src/tools/pgindent/pgindent | 2 +- 9 files changed, 1751 insertions(+), 2 deletions(-) create mode 100644 .clangd create mode 100644 .gdbinit create mode 100644 flake.lock create mode 100644 flake.nix create mode 100644 glibc-no-fortify-warning.patch create mode 100644 pg-aliases.sh create mode 100644 shell.nix diff --git a/.clangd b/.clangd new file mode 100644 index 0000000000000..490220f56a7c0 --- /dev/null +++ b/.clangd @@ -0,0 +1,43 @@ +Diagnostics: + MissingIncludes: None +InlayHints: + Enabled: true + ParameterNames: true + DeducedTypes: true +CompileFlags: + CompilationDatabase: build/ # Search build/ directory for compile_commands.json + Remove: [ -Werror ] + Add: + - -DDEBUG + - -DLOCAL + - -DPGDLLIMPORT= + - -DPIC + - -O2 + - -Wall + - -Wcast-function-type + - -Wconversion + - -Wdeclaration-after-statement + - -Wendif-labels + - -Werror=vla + - -Wextra + - -Wfloat-equal + - -Wformat-security + - -Wimplicit-fallthrough=3 + - -Wmissing-format-attribute + - -Wmissing-prototypes + - -Wno-format-truncation + - -Wno-sign-conversion + - -Wno-stringop-truncation + - -Wno-unused-const-variable + - -Wpointer-arith + - -Wshadow + - -Wshadow=compatible-local + - -fPIC + - -fexcess-precision=standard + - -fno-strict-aliasing + - -fvisibility=hidden + - -fwrapv + - -g + - -std=c11 + - -I. + - -I../../../../src/include diff --git a/.gdbinit b/.gdbinit new file mode 100644 index 0000000000000..854e5ecbaf69c --- /dev/null +++ b/.gdbinit @@ -0,0 +1,156 @@ +# HOT Indexed Updates — GDB breakpoints for code review +# +# Usage: gdb -x .gdbinit +# Or from gdb: source .gdbinit +# +# These breakpoints cover the major code paths introduced or modified by +# the HOT indexed updates patch series. They are organized by subsystem +# to make it easy to enable/disable groups during debugging. +# +# Tip: to skip to a specific subsystem, disable all then enable selectively: +# disable breakpoints +# enable 1 2 3 # just the update-decision group + +# ========================================================================= +# 1. UPDATE DECISION — heap_update() HOT/HOT-indexed/non-HOT choice +# src/backend/access/heap/heapam.c +# ========================================================================= + +# Main entry: heap_update +break heapam.c:3210 + +# HOT decision block: pure HOT vs HOT indexed vs non-HOT +# Line 4019: pure HOT (no indexed columns changed) +# Line 4024: HOT indexed path (non-catalog, some indexed columns changed) +# Line 4031: predict augmented tuple size +# Line 4033: size+space check before creating augmented tuple +break heapam.c:4019 +break heapam.c:4024 +break heapam.c:4033 + +# Set HEAP_INDEXED_UPDATED flag on new tuple before page insertion +break heapam.c:4101 + +# Restore HEAP_INDEXED_UPDATED on old tuple (only if it previously had it) +break heapam.c:4147 + +# ========================================================================= +# 2. TUPLE CREATION — building the augmented tuple with embedded bitmap +# src/backend/access/heap/heapam.c +# ========================================================================= + +# Predict augmented tuple size (returns 0 if t_hoff would overflow) +break heap_hot_indexed_tuple_size + +# Create augmented tuple with embedded modified-column bitmap +break heap_hot_indexed_create_tuple + +# Serialize Bitmapset into raw bytes in tuple header +break heap_hot_indexed_serialize_bitmap + +# ========================================================================= +# 3. BITMAP UTILITIES — raw bitmap operations for chain following +# src/backend/access/heap/heapam.c +# ========================================================================= + +# Compute raw bitmap byte size from natts +break heap_hot_indexed_bitmap_raw_size + +# Check if tuple header has room for bitmap between null bitmap and data +break heap_hot_indexed_has_bitmap_space + +# Read HOT indexed bitmap from tuple header (returns Bitmapset) +break heap_hot_indexed_read_bitmap + +# Fast overlap check: does tuple's raw bitmap overlap with indexed_attrs? +break heap_hot_indexed_bitmap_overlaps_raw + +# OR a tuple's raw bitmap into an accumulator buffer +break heap_hot_indexed_bitmap_or_raw + +# Check if accumulated raw bitmap overlaps with indexed_attrs +break heap_hot_indexed_accum_overlaps + +# Merge bitmaps from dead tuples into a target tuple on the page +break heap_hot_indexed_merge_bitmaps_raw + +# Deserialize raw bytes back to Bitmapset +break heap_hot_indexed_deserialize_bitmap + +# ========================================================================= +# 4. INDEX SCAN — HOT chain following with stale-entry detection +# src/backend/access/heap/heapam_indexscan.c +# ========================================================================= + +# Main HOT chain search with indexed update awareness +break heap_hot_search_buffer + +# Redirect-with-data: initialize bitmap accumulator from collapsed redirect +break heapam_indexscan.c:182 + +# Accumulate bitmap from INDEXED_UPDATED tuple in chain +break heapam_indexscan.c:250 + +# Stale entry detection: accumulated bitmap overlaps this index's attrs +break heapam_indexscan.c:297 + +# ========================================================================= +# 5. INDEX SCAN SETUP — indexed_attrs bitmap computation +# src/backend/access/index/indexam.c +# ========================================================================= + +# Compute indexed_attrs for HOT indexed update chain following +break indexam.c:299 + +# ========================================================================= +# 6. INDEX INSERTION — skip unchanged indexes for HOT indexed updates +# src/backend/executor/execIndexing.c +# ========================================================================= + +# Entry: insert/update index tuples +break ExecInsertIndexTuples + +# Index skip decision: skip indexes whose attrs don't overlap modified set +break execIndexing.c:370 + +# ========================================================================= +# 7. PRUNING — chain collapsing and redirect-with-data +# src/backend/access/heap/pruneheap.c +# ========================================================================= + +# Main prune function +break heap_page_prune_and_freeze + +# Per-chain pruning entry +break heap_prune_chain + +# Chain collapsing: collect bitmaps from dead INDEXED_UPDATED intermediates +break pruneheap.c:1802 + +# OR dead tuple bitmaps into combined bitmap +break pruneheap.c:1836 + +# Record redirect-with-data for execute phase +break pruneheap.c:1863 + +# Execute phase: apply redirect-with-data entries on the page +break pruneheap.c:1287 + +# ========================================================================= +# 8. WAL REPLAY — recovery of HOT indexed updates +# src/backend/access/heap/heapam_xlog.c +# ========================================================================= + +# WAL replay for XLOG_HEAP2_INDEXED_UPDATE +break heap_xlog_indexed_update + +# ========================================================================= +# 9. WAL LOGGING — writing HOT indexed update records +# src/backend/access/heap/heapam.c +# ========================================================================= + +# WAL logging for heap updates (handles indexed_update flag) +break log_heap_update + +# Serialize redirect-with-data into WAL record (pruneheap.c) +break pruneheap.c:2936 diff --git a/flake.lock b/flake.lock new file mode 100644 index 0000000000000..b8e8a1fdb750f --- /dev/null +++ b/flake.lock @@ -0,0 +1,78 @@ +{ + "nodes": { + "flake-utils": { + "inputs": { + "systems": "systems" + }, + "locked": { + "lastModified": 1731533236, + "narHash": "sha256-l0KFg5HjrsfsO/JpG+r7fRrqm12kzFHyUHqHCVpMMbI=", + "owner": "numtide", + "repo": "flake-utils", + "rev": "11707dc2f618dd54ca8739b309ec4fc024de578b", + "type": "github" + }, + "original": { + "owner": "numtide", + "repo": "flake-utils", + "type": "github" + } + }, + "nixpkgs": { + "locked": { + "lastModified": 1767313136, + "narHash": "sha256-16KkgfdYqjaeRGBaYsNrhPRRENs0qzkQVUooNHtoy2w=", + "owner": "NixOS", + "repo": "nixpkgs", + "rev": "ac62194c3917d5f474c1a844b6fd6da2db95077d", + "type": "github" + }, + "original": { + "owner": "NixOS", + "ref": "nixos-25.05", + "repo": "nixpkgs", + "type": "github" + } + }, + "nixpkgs-unstable": { + "locked": { + "lastModified": 1777270315, + "narHash": "sha256-yKB4G6cKsQsWN7M6rZGk6gkJPDNPIzT05y4qzRyCDlI=", + "owner": "nixos", + "repo": "nixpkgs", + "rev": "6368eda62c9775c38ef7f714b2555a741c20c72d", + "type": "github" + }, + "original": { + "owner": "nixos", + "ref": "nixpkgs-unstable", + "repo": "nixpkgs", + "type": "github" + } + }, + "root": { + "inputs": { + "flake-utils": "flake-utils", + "nixpkgs": "nixpkgs", + "nixpkgs-unstable": "nixpkgs-unstable" + } + }, + "systems": { + "locked": { + "lastModified": 1681028828, + "narHash": "sha256-Vy1rq5AaRuLzOxct8nz4T6wlgyUR7zLU309k9mBC768=", + "owner": "nix-systems", + "repo": "default", + "rev": "da67096a3b9bf56a91d16901293e51ba5b49a27e", + "type": "github" + }, + "original": { + "owner": "nix-systems", + "repo": "default", + "type": "github" + } + } + }, + "root": "root", + "version": 7 +} diff --git a/flake.nix b/flake.nix new file mode 100644 index 0000000000000..aae6d54c4c8cf --- /dev/null +++ b/flake.nix @@ -0,0 +1,45 @@ +{ + description = "PostgreSQL development environment"; + + inputs = { + nixpkgs.url = "github:NixOS/nixpkgs/nixos-25.05"; + nixpkgs-unstable.url = "github:nixos/nixpkgs/nixpkgs-unstable"; + flake-utils.url = "github:numtide/flake-utils"; + }; + + outputs = { + self, + nixpkgs, + nixpkgs-unstable, + flake-utils, + }: + flake-utils.lib.eachDefaultSystem ( + system: let + pkgs = import nixpkgs { + inherit system; + config.allowUnfree = true; + }; + pkgs-unstable = import nixpkgs-unstable { + inherit system; + config.allowUnfree = true; + }; + + shellConfig = import ./shell.nix {inherit pkgs pkgs-unstable system;}; + in { + formatter = pkgs.alejandra; + devShells = { + default = shellConfig.devShell; + gcc = shellConfig.devShell; + clang = shellConfig.clangDevShell; + gcc-musl = shellConfig.muslDevShell; + clang-musl = shellConfig.clangMuslDevShell; + }; + + packages = { + inherit (shellConfig) gdbConfig flameGraphScript pgbenchScript; + }; + + environment.localBinInPath = true; + } + ); +} diff --git a/glibc-no-fortify-warning.patch b/glibc-no-fortify-warning.patch new file mode 100644 index 0000000000000..4657a12adbcc5 --- /dev/null +++ b/glibc-no-fortify-warning.patch @@ -0,0 +1,24 @@ +From 130c231020f97e5eb878cc9fdb2bd9b186a5aa04 Mon Sep 17 00:00:00 2001 +From: Greg Burd +Date: Fri, 24 Oct 2025 11:58:24 -0400 +Subject: [PATCH] no warnings with -O0 and fortify source please + +--- + include/features.h | 1 - + 1 file changed, 1 deletion(-) + +diff --git a/include/features.h b/include/features.h +index 673c4036..a02c8a3f 100644 +--- a/include/features.h ++++ b/include/features.h +@@ -432,7 +432,6 @@ + + #if defined _FORTIFY_SOURCE && _FORTIFY_SOURCE > 0 + # if !defined __OPTIMIZE__ || __OPTIMIZE__ <= 0 +-# warning _FORTIFY_SOURCE requires compiling with optimization (-O) + # elif !__GNUC_PREREQ (4, 1) + # warning _FORTIFY_SOURCE requires GCC 4.1 or later + # elif _FORTIFY_SOURCE > 2 && (__glibc_clang_prereq (9, 0) \ +-- +2.50.1 + diff --git a/pg-aliases.sh b/pg-aliases.sh new file mode 100644 index 0000000000000..0c13adc8f903a --- /dev/null +++ b/pg-aliases.sh @@ -0,0 +1,658 @@ +# PostgreSQL Development Aliases + +# ============================================================ +# Build helpers shared by every variant. +# ============================================================ +pg_clean_for_compiler() { + local current_compiler="$(basename $CC)" + local build_dir="${1:-$PG_BUILD_DIR}" + + if [ -f "$build_dir/compile_commands.json" ]; then + local last_compiler=$(grep -o '/[^/]*/bin/[gc]cc\|/[^/]*/bin/clang' "$build_dir/compile_commands.json" | head -1 | xargs basename 2>/dev/null || echo "unknown") + + if [ "$last_compiler" != "$current_compiler" ] && [ "$last_compiler" != "unknown" ]; then + echo "Detected compiler change from $last_compiler to $current_compiler" + echo "Cleaning build directory..." + trash "$build_dir" 2>/dev/null || rm -rf "$build_dir" + mkdir -p "$build_dir" + fi + fi + + mkdir -p "$build_dir" + echo "$current_compiler" >"$build_dir/.compiler_used" +} + +# ============================================================ +# Core PostgreSQL commands (default/debug build) +# ============================================================ +alias pg-setup=' + if [ -z "$PERL_CORE_DIR" ]; then + echo "Error: Could not find perl CORE directory" >&2 + return 1 + fi + + pg_clean_for_compiler "$PG_BUILD_DIR" + + echo "=== PostgreSQL Build Configuration ===" + echo "Compiler: $CC" + echo "LLVM: $(llvm-config --version 2>/dev/null || echo disabled)" + echo "Source: $PG_SOURCE_DIR" + echo "Build: $PG_BUILD_DIR" + echo "Install: $PG_INSTALL_DIR" + echo "======================================" + + env CFLAGS="-I$PERL_CORE_DIR $CFLAGS" \ + LDFLAGS="-L$PERL_CORE_DIR -lperl $LDFLAGS" \ + meson setup $MESON_EXTRA_SETUP \ + --reconfigure \ + -Doptimization=g \ + -Ddebug=true \ + -Db_sanitize=none \ + -Db_lundef=false \ + -Dlz4=enabled \ + -Dzstd=enabled \ + -Dllvm=disabled \ + -Dplperl=enabled \ + -Dplpython=enabled \ + -Dpltcl=enabled \ + -Dlibxml=enabled \ + -Duuid=e2fs \ + -Dlibxslt=enabled \ + -Dssl=openssl \ + -Dldap=disabled \ + -Dcassert=true \ + -Dtap_tests=enabled \ + -Dinjection_points=true \ + -Ddocs_pdf=enabled \ + -Ddocs_html_style=website \ + --prefix="$PG_INSTALL_DIR" \ + "$PG_BUILD_DIR" \ + "$PG_SOURCE_DIR"' + +alias pg-compdb='compdb -p build/ list > compile_commands.json' +alias pg-build='meson compile -C "$PG_BUILD_DIR"' +alias pg-install='meson install -C "$PG_BUILD_DIR"' +alias pg-test='meson test -q --print-errorlogs -C "$PG_BUILD_DIR"' + +# Clean commands +alias pg-clean='ninja -C "$PG_BUILD_DIR" clean' +alias pg-full-clean='trash "$PG_BUILD_DIR" "$PG_INSTALL_DIR" 2>/dev/null || rm -rf "$PG_BUILD_DIR" "$PG_INSTALL_DIR"; echo "Build and install directories cleaned"' + +# Database management +alias pg-init='trash "$PG_DATA_DIR" 2>/dev/null || rm -rf "$PG_DATA_DIR"; "$PG_INSTALL_DIR/bin/initdb" --debug --no-clean "$PG_DATA_DIR"' + +alias pg-start='ulimit -c unlimited && "$PG_INSTALL_DIR/bin/postgres" -D "$PG_DATA_DIR" -k "$PG_DATA_DIR"' + +alias pg-stop='pkill -f "postgres.*-D.*$PG_DATA_DIR" || true' +alias pg-restart='pg-stop && sleep 2 && pg-start' +alias pg-status='pgrep -f "postgres.*-D.*$PG_DATA_DIR" && echo "PostgreSQL is running" || echo "PostgreSQL is not running"' + +# Client connections +alias pg-psql='"$PG_INSTALL_DIR/bin/psql" -h "$PG_DATA_DIR" postgres' +alias pg-createdb='"$PG_INSTALL_DIR/bin/createdb" -h "$PG_DATA_DIR"' +alias pg-dropdb='"$PG_INSTALL_DIR/bin/dropdb" -h "$PG_DATA_DIR"' + +# ============================================================ +# Debugger attachments +# ============================================================ +alias pg-debug-gdb='gdb -x "$GDBINIT" -x .gdbinit "$PG_INSTALL_DIR/bin/postgres"' +alias pg-debug-lldb='lldb "$PG_INSTALL_DIR/bin/postgres"' +alias pg-debug=' + if command -v gdb >/dev/null 2>&1; then + pg-debug-gdb + elif command -v lldb >/dev/null 2>&1; then + pg-debug-lldb + else + echo "No debugger available (gdb or lldb required)" + fi' + +alias pg-attach-gdb=' + PG_PID=$(pgrep -f "postgres.*-D.*$PG_DATA_DIR" | head -1) + if [ -n "$PG_PID" ]; then + echo "Attaching GDB to PostgreSQL process $PG_PID" + gdb -x "$GDBINIT" -x .gdbinit -p "$PG_PID" + else + echo "No PostgreSQL process found" + fi' + +alias pg-attach-lldb=' + PG_PID=$(pgrep -f "postgres.*-D.*$PG_DATA_DIR" | head -1) + if [ -n "$PG_PID" ]; then + echo "Attaching LLDB to PostgreSQL process $PG_PID" + lldb -p "$PG_PID" + else + echo "No PostgreSQL process found" + fi' + +alias pg-attach=' + if command -v gdb >/dev/null 2>&1; then + pg-attach-gdb + elif command -v lldb >/dev/null 2>&1; then + pg-attach-lldb + else + echo "No debugger available (gdb or lldb required)" + fi' + +# ============================================================ +# Valgrind-instrumented build and tests +# +# The valgrind build lives in a separate directory so the normal +# build stays warm. Runs use a wrapper dir that shadows `postgres` +# with a valgrind wrapper -- pg_regress finds it via PATH. +# ============================================================ +pg-build-valgrind() { + local bdir="$PG_BUILD_DIR_VALGRIND" + if [ -z "$PERL_CORE_DIR" ]; then + echo "Error: PERL_CORE_DIR is not set" >&2 + return 1 + fi + + pg_clean_for_compiler "$bdir" + + echo "=== Configuring Valgrind build in $bdir ===" + env CFLAGS="-Og -ggdb3 -fno-omit-frame-pointer -DUSE_VALGRIND -I$PERL_CORE_DIR $CFLAGS" \ + LDFLAGS="-L$PERL_CORE_DIR -lperl $LDFLAGS" \ + meson setup --reconfigure \ + -Doptimization=g \ + -Ddebug=true \ + -Dcassert=true \ + -Dtap_tests=enabled \ + -Dinjection_points=true \ + -Dllvm=disabled \ + -Dplperl=enabled -Dplpython=enabled -Dpltcl=enabled \ + -Dlz4=enabled -Dzstd=enabled \ + -Dlibxml=enabled -Dlibxslt=enabled -Dssl=openssl -Duuid=e2fs \ + -Dldap=disabled \ + --prefix="$PG_INSTALL_DIR-valgrind" \ + "$bdir" "$PG_SOURCE_DIR" || return 1 + + meson compile -C "$bdir" +} + +# Drop a wrapper directory that shadows the real binaries; `postgres` +# exec's into valgrind, everything else is a symlink. Writes to the +# supplied wrap dir and echoes its path. +_pg_make_valgrind_wrapper() { + local bindir="$1" + local wrapdir="$2" + + mkdir -p "$wrapdir" + cat >"$wrapdir/postgres" <&2 + return 1 + fi + + local tmpbin="$bdir/tmp_install$PG_INSTALL_DIR-valgrind/bin" + if [ ! -x "$tmpbin/postgres" ]; then + echo "Populating tmp_install..." + meson test -C "$bdir" tmp_install install_test_files initdb_cache >/dev/null || return 1 + fi + + local wrap + wrap=$(mktemp -d /tmp/pg-vg-wrap-XXXXXX) + _pg_make_valgrind_wrapper "$tmpbin" "$wrap" + + mkdir -p "$PG_BENCH_DIR" + echo "Valgrind logs: $PG_BENCH_DIR/valgrind-*.log" + echo "Wrapper dir: $wrap (will be removed on exit)" + echo "Expect the regress suite to take 15-45 minutes under valgrind." + + local rc=0 + (cd "$bdir" && PATH="$wrap:$PATH" meson test -t 60 --print-errorlogs regress/regress) || rc=$? + + trash "$wrap" 2>/dev/null || rm -rf "$wrap" + return "$rc" +} + +pg-valgrind-test() { + local bdir="$PG_BUILD_DIR_VALGRIND" + if [ ! -x "$bdir/src/backend/postgres" ]; then + echo "Valgrind build not found; run 'pg-build-valgrind' first." >&2 + return 1 + fi + + echo "This runs the FULL postgres test suite under valgrind." + echo "Expect many hours, and tens of GB of valgrind log output." + echo "Logs: $PG_BENCH_DIR/valgrind-*.log" + local yn + read -r -p "Continue? [y/N] " yn + case "$yn" in + y | Y | yes) ;; + *) echo "Aborted."; return 0 ;; + esac + + local tmpbin="$bdir/tmp_install$PG_INSTALL_DIR-valgrind/bin" + if [ ! -x "$tmpbin/postgres" ]; then + echo "Populating tmp_install..." + meson test -C "$bdir" tmp_install install_test_files initdb_cache >/dev/null || return 1 + fi + + local wrap + wrap=$(mktemp -d /tmp/pg-vg-wrap-XXXXXX) + _pg_make_valgrind_wrapper "$tmpbin" "$wrap" + mkdir -p "$PG_BENCH_DIR" + + local rc=0 + (cd "$bdir" && PATH="$wrap:$PATH" meson test -t 60 --print-errorlogs) || rc=$? + + trash "$wrap" 2>/dev/null || rm -rf "$wrap" + return "$rc" +} + +# ============================================================ +# AddressSanitizer / UndefinedBehaviorSanitizer build and tests +# ============================================================ +pg-build-asan() { + local bdir="$PG_BUILD_DIR_ASAN" + if [ -z "$PERL_CORE_DIR" ]; then + echo "Error: PERL_CORE_DIR is not set" >&2 + return 1 + fi + + pg_clean_for_compiler "$bdir" + + echo "=== Configuring ASan+UBSan build in $bdir ===" + env CFLAGS="-Og -ggdb3 -fno-omit-frame-pointer -fsanitize=address,undefined -fno-sanitize-recover=all -I$PERL_CORE_DIR $CFLAGS" \ + LDFLAGS="-fsanitize=address,undefined -L$PERL_CORE_DIR -lperl $LDFLAGS" \ + meson setup --reconfigure \ + -Doptimization=g \ + -Ddebug=true \ + -Dcassert=true \ + -Dtap_tests=enabled \ + -Dinjection_points=true \ + -Dllvm=disabled \ + -Dplperl=enabled -Dplpython=enabled -Dpltcl=enabled \ + -Dlz4=enabled -Dzstd=enabled \ + -Dlibxml=enabled -Dlibxslt=enabled -Dssl=openssl -Duuid=e2fs \ + -Dldap=disabled \ + --prefix="$PG_INSTALL_DIR-asan" \ + "$bdir" "$PG_SOURCE_DIR" || return 1 + + meson compile -C "$bdir" +} + +pg-asan-regress() { + local bdir="$PG_BUILD_DIR_ASAN" + if [ ! -x "$bdir/src/backend/postgres" ]; then + echo "ASan build not found; run 'pg-build-asan' first." >&2 + return 1 + fi + + # halt_on_error=0 lets regress continue past the first diagnostic so + # the whole suite runs; abort_on_error=1 makes each hit fail the test. + ASAN_OPTIONS="halt_on_error=0:abort_on_error=1:detect_leaks=0:print_summary=1:print_stacktrace=1" \ + UBSAN_OPTIONS="halt_on_error=1:abort_on_error=1:print_stacktrace=1:print_summary=1" \ + meson test -t 5 --print-errorlogs -C "$bdir" regress/regress +} + +# ============================================================ +# rr (deterministic record-and-replay) +# Requires kernel.perf_event_paranoid <= 1. rr is the single most +# effective tool for postgres bugs that reproduce intermittently. +# ============================================================ +pg-rr-check() { + if ! command -v rr >/dev/null; then + echo "rr is not installed (expected in the dev shell)." >&2 + return 1 + fi + local paranoid + paranoid=$(cat /proc/sys/kernel/perf_event_paranoid 2>/dev/null || echo 99) + if [ "$paranoid" -gt 1 ]; then + echo "rr requires kernel.perf_event_paranoid <= 1; currently $paranoid" + echo "To enable (root needed):" + echo " echo 1 | sudo tee /proc/sys/kernel/perf_event_paranoid" + return 1 + fi + echo "rr ready (perf_event_paranoid=$paranoid)" +} + +pg-rr-record() { + pg-rr-check >/dev/null || { + pg-rr-check + return 1 + } + ulimit -c unlimited + rr record -- "$PG_INSTALL_DIR/bin/postgres" -D "$PG_DATA_DIR" -k "$PG_DATA_DIR" +} + +pg-rr-replay() { + rr replay "$@" +} + +# ============================================================ +# perf wrappers (parallel to the flame-graph helper) +# ============================================================ +pg-perf-record() { + local pid + pid=$(pgrep -f "postgres.*-D.*$PG_DATA_DIR" | head -1) + if [ -z "$pid" ]; then + echo "No postgres running under $PG_DATA_DIR" >&2 + return 1 + fi + mkdir -p "$PG_BENCH_DIR" + local out="$PG_BENCH_DIR/perf-$(date +%Y%m%d_%H%M%S).data" + echo "Recording to $out (Ctrl-C to stop)" + perf record -F 997 --call-graph dwarf -p "$pid" -o "$out" "$@" + echo "Saved: $out" +} + +pg-perf-report() { + local data + data=$(ls -t "$PG_BENCH_DIR"/perf-*.data 2>/dev/null | head -1) + if [ -z "$data" ]; then + echo "No perf data in $PG_BENCH_DIR" >&2 + return 1 + fi + echo "Reading $data" + perf report -i "$data" "$@" +} + +pg-perf-annotate() { + local data + data=$(ls -t "$PG_BENCH_DIR"/perf-*.data 2>/dev/null | head -1) + if [ -z "$data" ]; then + echo "No perf data in $PG_BENCH_DIR" >&2 + return 1 + fi + perf annotate -i "$data" "$@" +} + +# ============================================================ +# Single regression test / group runner. +# Runs pg_regress directly against the existing build so you skip the +# full meson-driven suite wrapper. Usage: pg-test-one boolean [name ...] +# ============================================================ +pg-test-one() { + if [ $# -eq 0 ]; then + echo "usage: pg-test-one TESTNAME [TESTNAME ...]" + echo "example: pg-test-one boolean" + return 2 + fi + local bdir="${PG_BUILD_DIR_ONE:-$PG_BUILD_DIR}" + local tmpbin="$bdir/tmp_install$PG_INSTALL_DIR/bin" + if [ ! -x "$tmpbin/postgres" ]; then + echo "Populating tmp_install..." + meson test -C "$bdir" tmp_install install_test_files initdb_cache >/dev/null || return 1 + fi + local outdir + outdir=$(mktemp -d /tmp/pg-test-one-XXXXXX) + echo "Test output: $outdir" + "$bdir/src/test/regress/pg_regress" \ + --bindir="$tmpbin" \ + --inputdir="$PG_SOURCE_DIR/src/test/regress" \ + --expecteddir="$PG_SOURCE_DIR/src/test/regress" \ + --dlpath="$bdir/src/test/regress" \ + --outputdir="$outdir" \ + --temp-instance="$outdir/tmp" \ + --port=40099 \ + "$@" +} + +# Full flame graph / benchmark aliases +alias pg-flame='pg-flame-generate' +alias pg-flame-30='pg-flame-generate 30' +alias pg-flame-60='pg-flame-generate 60' +alias pg-flame-120='pg-flame-generate 120' + +pg-flame-custom() { + local duration=${1:-30} + local output_dir=${2:-$PG_FLAME_DIR} + echo "Generating flame graph for ${duration}s, output to: $output_dir" + pg-flame-generate "$duration" "$output_dir" +} + +alias pg-bench='pg-bench-run' +alias pg-bench-quick='pg-bench-run 5 1 100 1 30 select-only' +alias pg-bench-standard='pg-bench-run 10 2 1000 10 60 tpcb-like' +alias pg-bench-heavy='pg-bench-run 50 4 5000 100 300 tpcb-like' +alias pg-bench-readonly='pg-bench-run 20 4 2000 50 120 select-only' + +pg-bench-custom() { + local clients=${1:-10} + local threads=${2:-2} + local transactions=${3:-1000} + local scale=${4:-10} + local duration=${5:-60} + local test_type=${6:-tpcb-like} + + echo "Running custom benchmark:" + echo " Clients: $clients, Threads: $threads" + echo " Transactions: $transactions, Scale: $scale" + echo " Duration: ${duration}s, Type: $test_type" + + pg-bench-run "$clients" "$threads" "$transactions" "$scale" "$duration" "$test_type" +} + +pg-bench-flame() { + local duration=${1:-60} + local clients=${2:-10} + local scale=${3:-10} + + echo "Running benchmark with flame graph generation" + echo "Duration: ${duration}s, Clients: $clients, Scale: $scale" + + pg-bench-run "$clients" 2 1000 "$scale" "$duration" tpcb-like & + local bench_pid=$! + + sleep 5 + + local flame_duration=$((duration - 10)) + if [ $flame_duration -gt 10 ]; then + pg-flame-generate "$flame_duration" & + local flame_pid=$! + fi + + wait $bench_pid + if [ -n "${flame_pid:-}" ]; then + wait $flame_pid + fi + + echo "Benchmark and flame graph generation completed" +} + +# Live monitoring +alias pg-perf='perf top -p $(pgrep -f "postgres.*-D.*$PG_DATA_DIR" | head -1)' +alias pg-htop='htop -p $(pgrep -f "postgres.*-D.*$PG_DATA_DIR" | tr "\n" "," | sed "s/,$//")' + +pg-stats() { + local duration=${1:-30} + echo "Collecting system stats for ${duration}s..." + + iostat -x 1 "$duration" >"$PG_BENCH_DIR/iostat_$(date +%Y%m%d_%H%M%S).log" & + vmstat 1 "$duration" >"$PG_BENCH_DIR/vmstat_$(date +%Y%m%d_%H%M%S).log" & + + wait + echo "System stats saved to $PG_BENCH_DIR" +} + +# ============================================================ +# Code quality helpers +# ============================================================ +pg-format() { + local since=${1:-HEAD} + + if [ ! -f "$PG_SOURCE_DIR/src/tools/pgindent/pgindent" ]; then + echo "Error: pgindent not found at $PG_SOURCE_DIR/src/tools/pgindent/pgindent" + else + + modified_files=$(git diff --name-only "${since}" | grep -E "\.c$|\.h$") + + if [ -z "$modified_files" ]; then + echo "No modified .c or .h files found" + else + + echo "Formatting modified files with pgindent:" + for file in $modified_files; do + if [ -f "$file" ]; then + echo " Formatting: $file" + "$PG_SOURCE_DIR/src/tools/pgindent/pgindent" "$file" + else + echo " Warning: File not found: $file" + fi + done + + echo "Checking files for whitespace:" + git diff --check "${since}" + fi + fi +} + +pg-tidy() { + local since=${1:-HEAD} + local files + files=$(git diff --name-only "$since" | grep -E "\.(c|h)$") + if [ -z "$files" ]; then + echo "No modified .c or .h files." + return 0 + fi + for f in $files; do + [ -f "$f" ] || continue + echo "clang-tidy: $f" + clang-tidy -p "$PG_BUILD_DIR" "$f" 2>&1 | head -50 + done +} + +pg-spell() { + local since=${1:-HEAD} + local files=$(git diff --name-only "$since" | grep -E '\.(c|h|sgml|md)$') + if [ -z "$files" ]; then + echo "No .c/.h/.sgml/.md files changed since $since" + return 0 + fi + for f in $files; do + [ -f "$f" ] || continue + case "$f" in + *.c | *.h) + grep -nE '^\s*(/\*|\*|//)' "$f" | codespell --stdin-single-line - 2>/dev/null \ + && echo " $f: ok" || true + ;; + *.sgml | *.md) + codespell "$f" || true + ;; + esac + done +} + +# ============================================================ +# Core dump one-shots (one-time, requires root). kernel.core_pattern +# is a system-wide sysctl -- we don't touch it on every shell entry. +# ============================================================ +pg-cores-status() { + echo "ulimit -c: $(ulimit -c)" + echo "kernel.core_pattern: $(cat /proc/sys/kernel/core_pattern 2>/dev/null || echo unreadable)" + echo "cwd: $(pwd)" +} + +pg-enable-cores() { + ulimit -c unlimited + if ! [ -w /proc/sys/kernel/core_pattern ]; then + echo "Setting kernel.core_pattern (requires sudo)..." + echo "core.%p" | sudo tee /proc/sys/kernel/core_pattern >/dev/null || { + echo "Failed to write /proc/sys/kernel/core_pattern" >&2 + return 1 + } + else + echo "core.%p" >/proc/sys/kernel/core_pattern + fi + pg-cores-status +} + +pg-disable-cores() { + ulimit -c 0 + if ! [ -w /proc/sys/kernel/core_pattern ]; then + echo "Restoring kernel.core_pattern to 'core' (requires sudo)..." + echo "core" | sudo tee /proc/sys/kernel/core_pattern >/dev/null || { + echo "Failed to restore /proc/sys/kernel/core_pattern" >&2 + return 1 + } + else + echo "core" >/proc/sys/kernel/core_pattern + fi + pg-cores-status +} + +# ============================================================ +# Logs and results +# ============================================================ +alias pg-log='tail -f "$PG_DATA_DIR/log/postgresql-$(date +%Y-%m-%d).log" 2>/dev/null || echo "No log file found"' +alias pg-log-errors='grep -i error "$PG_DATA_DIR/log/"*.log 2>/dev/null || echo "No error logs found"' + +alias pg-build-log='cat "$PG_BUILD_DIR/meson-logs/meson-log.txt"' +alias pg-build-errors='grep -i error "$PG_BUILD_DIR/meson-logs/meson-log.txt" 2>/dev/null || echo "No build errors found"' + +alias pg-bench-results='ls -la "$PG_BENCH_DIR" && echo "Latest results:" && tail -20 "$PG_BENCH_DIR"/results_*.txt 2>/dev/null | tail -20' +alias pg-flame-results='ls -la "$PG_FLAME_DIR" && echo "Open flame graphs with: firefox $PG_FLAME_DIR/*.svg"' + +pg-clean-results() { + local days=${1:-7} + echo "Cleaning benchmark and flame graph results older than $days days..." + find "$PG_BENCH_DIR" -type f -mtime +$days -delete 2>/dev/null || true + find "$PG_FLAME_DIR" -type f -mtime +$days -delete 2>/dev/null || true + echo "Cleanup completed" +} + +# ============================================================ +# Info +# ============================================================ +alias pg-info=' + echo "=== PostgreSQL Development Environment ===" + echo "Source: $PG_SOURCE_DIR" + echo "Build (default): $PG_BUILD_DIR" + echo "Build (valgrind):$PG_BUILD_DIR_VALGRIND" + echo "Build (asan): $PG_BUILD_DIR_ASAN" + echo "Install: $PG_INSTALL_DIR" + echo "Data: $PG_DATA_DIR" + echo "Benchmarks: $PG_BENCH_DIR" + echo "Flame graphs: $PG_FLAME_DIR" + echo "Compiler: $CC" + echo "" + echo "Available commands:" + echo " Setup/build: pg-setup, pg-build, pg-install" + echo " Database: pg-init, pg-start, pg-stop, pg-psql" + echo " Tests: pg-test, pg-test-one NAME" + echo " Valgrind: pg-build-valgrind, pg-valgrind-regress, pg-valgrind-test" + echo " ASan/UBSan: pg-build-asan, pg-asan-regress" + echo " Debug: pg-debug, pg-attach" + echo " Record/replay: pg-rr-check, pg-rr-record, pg-rr-replay" + echo " Perf: pg-perf-record, pg-perf-report, pg-perf-annotate, pg-perf" + echo " Flame graphs: pg-flame, pg-flame-30, pg-flame-60, pg-flame-custom" + echo " Benchmarks: pg-bench-quick, pg-bench-standard, pg-bench-heavy" + echo " Combined: pg-bench-flame" + echo " Results: pg-bench-results, pg-flame-results" + echo " Logs: pg-log, pg-build-log" + echo " Clean: pg-clean, pg-full-clean, pg-clean-results" + echo " Code quality: pg-format, pg-tidy, pg-spell" + echo " Cores: pg-enable-cores, pg-disable-cores, pg-cores-status" + echo "=========================================="' + +echo "PostgreSQL aliases loaded. Run 'pg-info' for available commands." diff --git a/shell.nix b/shell.nix new file mode 100644 index 0000000000000..8c738afa789a7 --- /dev/null +++ b/shell.nix @@ -0,0 +1,745 @@ +{ + pkgs, + pkgs-unstable, + system, +}: let + # Create a patched glibc only for the dev shell. + # + # Glibc's features.h emits a `-Wcpp` diagnostic when _FORTIFY_SOURCE is + # defined without an optimization level. Meson's dependency probes + # (notably the libcurl thread-safety check) compile small snippets with + # `-O0 -Werror`, which turns that cpp warning into a hard error and + # breaks reconfigure under our default CFLAGS. The patch simply drops + # the warning. It is scoped to this dev shell only and never leaks + # into system glibc or release builds. + patchedGlibc = pkgs.glibc.overrideAttrs (oldAttrs: { + patches = (oldAttrs.patches or []) ++ [ + ./glibc-no-fortify-warning.patch + ]; + }); + + # Use LLVM for modern PostgreSQL development + llvmPkgs = pkgs-unstable.llvmPackages_21; + + # Configuration constants + config = { + pgSourceDir = "$PWD"; + pgBuildDir = "$PWD/build"; + pgBuildDirValgrind = "$PWD/build-valgrind"; + pgBuildDirAsan = "$PWD/build-asan"; + pgInstallDir = "$PWD/install"; + pgDataDir = "/tmp/test-db-$(basename $PWD)"; + pgBenchDir = "/tmp/pgbench-results-$(basename $PWD)"; + pgFlameDir = "/tmp/flame-graphs-$(basename $PWD)"; + }; + + # Single dependency function that can be used for all environments + getPostgreSQLDeps = muslLibs: + with pkgs; + [ + # Build system (always use host tools) + pkgs-unstable.meson + pkgs-unstable.ninja + pkg-config + autoconf + git + which + binutils + gnumake + mold # fast linker, big wins on large postgres links + + # Parser/lexer tools + bison + flex + + # Perl with required packages + (perl.withPackages (ps: with ps; [IPCRun])) + + # Documentation + docbook_xml_dtd_45 + docbook-xsl-nons + libxslt + libxml2 + fop + + # Development tools (always use host tools) + coreutils + shellcheck + ripgrep + valgrind + curl + uv + pylint + black + lcov + strace + ltrace + perf-tools + linuxPackages.perf + flamegraph + bpftrace # kernel-level tracing (probes, uprobes) + rr # record-and-replay deterministic debugger + htop + iotop + sysstat + ccache + cppcheck + compdb + + # Spell checking + aspell + aspellDicts.en + codespell + + # GCC/GDB + gcc + gdb + + # LLVM toolchain + llvmPkgs.llvm + llvmPkgs.llvm.dev + llvmPkgs.clang-tools + llvmPkgs.lldb + + # Language support + (python3.withPackages (ps: with ps; [requests browser-cookie3])) + tcl + ] + ++ ( + if muslLibs + then [ + # Musl target libraries for cross-compilation + pkgs.pkgsMusl.readline + pkgs.pkgsMusl.zlib + pkgs.pkgsMusl.openssl + pkgs.pkgsMusl.icu + pkgs.pkgsMusl.lz4 + pkgs.pkgsMusl.zstd + pkgs.pkgsMusl.libuuid + pkgs.pkgsMusl.libkrb5 + pkgs.pkgsMusl.linux-pam + pkgs.pkgsMusl.libxcrypt + ] + else [ + # Glibc target libraries + readline + zlib + openssl + icu + lz4 + zstd + libuuid + libkrb5 + linux-pam + libxcrypt + numactl + openldap + liburing + libselinux + patchedGlibc + patchedGlibc.dev + ] + ); + + # GDB configuration for PostgreSQL debugging + gdbConfig = pkgs.writeText "gdbinit-postgres" '' + # PostgreSQL-specific GDB configuration + + # Pretty-print PostgreSQL data structures + define print_node + if $arg0 + printf "Node type: %s\n", nodeTagNames[$arg0->type] + print *$arg0 + else + printf "NULL node\n" + end + end + document print_node + Print a PostgreSQL Node with type information + Usage: print_node + end + + define print_list + set $list = (List*)$arg0 + if $list + printf "List length: %d\n", $list->length + set $cell = $list->head + set $i = 0 + while $cell && $i < $list->length + printf " [%d]: ", $i + print_node $cell->data.ptr_value + set $cell = $cell->next + set $i = $i + 1 + end + else + printf "NULL list\n" + end + end + document print_list + Print a PostgreSQL List structure + Usage: print_list + end + + define print_query + set $query = (Query*)$arg0 + if $query + printf "Query type: %d, command type: %d\n", $query->querySource, $query->commandType + print *$query + else + printf "NULL query\n" + end + end + document print_query + Print a PostgreSQL Query structure + Usage: print_query + end + + define print_relcache + set $rel = (Relation)$arg0 + if $rel + printf "Relation: %s.%s (OID: %u)\n", $rel->rd_rel->relnamespace, $rel->rd_rel->relname.data, $rel->rd_id + printf " natts: %d, relkind: %c\n", $rel->rd_rel->relnatts, $rel->rd_rel->relkind + else + printf "NULL relation\n" + end + end + document print_relcache + Print relation cache entry information + Usage: print_relcache + end + + define print_tupdesc + set $desc = (TupleDesc)$arg0 + if $desc + printf "TupleDesc: %d attributes\n", $desc->natts + set $i = 0 + while $i < $desc->natts + set $attr = $desc->attrs[$i] + printf " [%d]: %s (type: %u, len: %d)\n", $i, $attr->attname.data, $attr->atttypid, $attr->attlen + set $i = $i + 1 + end + else + printf "NULL tuple descriptor\n" + end + end + document print_tupdesc + Print tuple descriptor information + Usage: print_tupdesc + end + + define print_slot + set $slot = (TupleTableSlot*)$arg0 + if $slot + printf "TupleTableSlot: %s\n", $slot->tts_ops->name + printf " empty: %d, shouldFree: %d\n", $slot->tts_empty, $slot->tts_shouldFree + if $slot->tts_tupleDescriptor + print_tupdesc $slot->tts_tupleDescriptor + end + else + printf "NULL slot\n" + end + end + document print_slot + Print tuple table slot information + Usage: print_slot + end + + # Memory context debugging + define print_mcxt + set $context = (MemoryContext)$arg0 + if $context + printf "MemoryContext: %s\n", $context->name + printf " type: %s, parent: %p\n", $context->methods->name, $context->parent + printf " total: %zu, free: %zu\n", $context->mem_allocated, $context->freep - $context->freeptr + else + printf "NULL memory context\n" + end + end + document print_mcxt + Print memory context information + Usage: print_mcxt + end + + # Process debugging + define print_proc + set $proc = (PGPROC*)$arg0 + if $proc + printf "PGPROC: pid=%d, database=%u\n", $proc->pid, $proc->databaseId + printf " waiting: %d, waitStatus: %d\n", $proc->waiting, $proc->waitStatus + else + printf "NULL process\n" + end + end + document print_proc + Print process information + Usage: print_proc + end + + # Set useful defaults + set print pretty on + set print object on + set print static-members off + set print vtbl on + set print demangle on + set demangle-style gnu-v3 + set print sevenbit-strings off + set history save on + set history size 1000 + set history filename ~/.gdb_history_postgres + + # Common breakpoints for PostgreSQL debugging + define pg_break_common + break elog + break errfinish + break ExceptionalCondition + break ProcessInterrupts + end + document pg_break_common + Set common PostgreSQL debugging breakpoints + end + + printf "PostgreSQL GDB configuration loaded.\n" + printf "Available commands: print_node, print_list, print_query, print_relcache,\n" + printf " print_tupdesc, print_slot, print_mcxt, print_proc, pg_break_common\n" + ''; + + # Flame graph generation script + flameGraphScript = pkgs.writeScriptBin "pg-flame-generate" '' + #!${pkgs.bash}/bin/bash + set -euo pipefail + + DURATION=''${1:-30} + OUTPUT_DIR=''${2:-${config.pgFlameDir}} + TIMESTAMP=$(date +%Y%m%d_%H%M%S) + + mkdir -p "$OUTPUT_DIR" + + echo "Generating flame graph for PostgreSQL (duration: ''${DURATION}s)" + + # Find PostgreSQL processes + PG_PIDS=$(pgrep -f "postgres.*-D.*${config.pgDataDir}" || true) + + if [ -z "$PG_PIDS" ]; then + echo "Error: No PostgreSQL processes found" + exit 1 + fi + + echo "Found PostgreSQL processes: $PG_PIDS" + + # Record perf data + PERF_DATA="$OUTPUT_DIR/perf_$TIMESTAMP.data" + echo "Recording perf data to $PERF_DATA" + + ${pkgs.linuxPackages.perf}/bin/perf record \ + -F 997 \ + -g \ + --call-graph dwarf \ + -p "$(echo $PG_PIDS | tr ' ' ',')" \ + -o "$PERF_DATA" \ + sleep "$DURATION" + + # Generate flame graph + FLAME_SVG="$OUTPUT_DIR/postgres_flame_$TIMESTAMP.svg" + echo "Generating flame graph: $FLAME_SVG" + + ${pkgs.linuxPackages.perf}/bin/perf script -i "$PERF_DATA" | \ + ${pkgs.flamegraph}/bin/stackcollapse-perf.pl | \ + ${pkgs.flamegraph}/bin/flamegraph.pl \ + --title "PostgreSQL Flame Graph ($TIMESTAMP)" \ + --width 1200 \ + --height 800 \ + > "$FLAME_SVG" + + echo "Flame graph generated: $FLAME_SVG" + echo "Perf data saved: $PERF_DATA" + + # Generate summary report + REPORT="$OUTPUT_DIR/report_$TIMESTAMP.txt" + echo "Generating performance report: $REPORT" + + { + echo "PostgreSQL Performance Analysis Report" + echo "Generated: $(date)" + echo "Duration: ''${DURATION}s" + echo "Processes: $PG_PIDS" + echo "" + echo "=== Top Functions ===" + ${pkgs.linuxPackages.perf}/bin/perf report -i "$PERF_DATA" --stdio --sort comm,dso,symbol | head -50 + echo "" + echo "=== Call Graph ===" + ${pkgs.linuxPackages.perf}/bin/perf report -i "$PERF_DATA" --stdio -g --sort comm,dso,symbol | head -100 + } > "$REPORT" + + echo "Report generated: $REPORT" + echo "" + echo "Files created:" + echo " Flame graph: $FLAME_SVG" + echo " Perf data: $PERF_DATA" + echo " Report: $REPORT" + ''; + + # pgbench wrapper script + pgbenchScript = pkgs.writeScriptBin "pg-bench-run" '' + #!${pkgs.bash}/bin/bash + set -euo pipefail + + # Default parameters + CLIENTS=''${1:-10} + THREADS=''${2:-2} + TRANSACTIONS=''${3:-1000} + SCALE=''${4:-10} + DURATION=''${5:-60} + TEST_TYPE=''${6:-tpcb-like} + + OUTPUT_DIR="${config.pgBenchDir}" + TIMESTAMP=$(date +%Y%m%d_%H%M%S) + + mkdir -p "$OUTPUT_DIR" + + echo "=== PostgreSQL Benchmark Configuration ===" + echo "Clients: $CLIENTS" + echo "Threads: $THREADS" + echo "Transactions: $TRANSACTIONS" + echo "Scale factor: $SCALE" + echo "Duration: ''${DURATION}s" + echo "Test type: $TEST_TYPE" + echo "Output directory: $OUTPUT_DIR" + echo "============================================" + + # Check if PostgreSQL is running + if ! pgrep -f "postgres.*-D.*${config.pgDataDir}" >/dev/null; then + echo "Error: PostgreSQL is not running. Start it with 'pg-start'" + exit 1 + fi + + PGBENCH="${config.pgInstallDir}/bin/pgbench" + PSQL="${config.pgInstallDir}/bin/psql" + CREATEDB="${config.pgInstallDir}/bin/createdb" + DROPDB="${config.pgInstallDir}/bin/dropdb" + + DB_NAME="pgbench_test_$TIMESTAMP" + RESULTS_FILE="$OUTPUT_DIR/results_$TIMESTAMP.txt" + LOG_FILE="$OUTPUT_DIR/pgbench_$TIMESTAMP.log" + + echo "Creating test database: $DB_NAME" + "$CREATEDB" -h "${config.pgDataDir}" "$DB_NAME" || { + echo "Failed to create database" + exit 1 + } + + # Initialize pgbench tables + echo "Initializing pgbench tables (scale factor: $SCALE)" + "$PGBENCH" -h "${config.pgDataDir}" -i -s "$SCALE" "$DB_NAME" || { + echo "Failed to initialize pgbench tables" + "$DROPDB" -h "${config.pgDataDir}" "$DB_NAME" 2>/dev/null || true + exit 1 + } + + # Run benchmark based on test type + echo "Running benchmark..." + + case "$TEST_TYPE" in + "tpcb-like"|"default") + BENCH_ARGS="" + ;; + "select-only") + BENCH_ARGS="-S" + ;; + "simple-update") + BENCH_ARGS="-N" + ;; + "read-write") + BENCH_ARGS="-b select-only@70 -b tpcb-like@30" + ;; + *) + echo "Unknown test type: $TEST_TYPE" + echo "Available types: tpcb-like, select-only, simple-update, read-write" + "$DROPDB" -h "${config.pgDataDir}" "$DB_NAME" 2>/dev/null || true + exit 1 + ;; + esac + + { + echo "PostgreSQL Benchmark Results" + echo "Generated: $(date)" + echo "Test type: $TEST_TYPE" + echo "Clients: $CLIENTS, Threads: $THREADS" + echo "Transactions: $TRANSACTIONS, Duration: ''${DURATION}s" + echo "Scale factor: $SCALE" + echo "Database: $DB_NAME" + echo "" + echo "=== System Information ===" + echo "CPU: $(nproc) cores" + echo "Memory: $(free -h | grep '^Mem:' | awk '{print $2}')" + echo "Compiler: $CC" + echo "PostgreSQL version: $("$PSQL" --no-psqlrc -h "${config.pgDataDir}" -d "$DB_NAME" -t -c "SELECT version();" | head -1)" + echo "" + echo "=== Benchmark Results ===" + } > "$RESULTS_FILE" + + # Run the actual benchmark + "$PGBENCH" \ + -h "${config.pgDataDir}" \ + -c "$CLIENTS" \ + -j "$THREADS" \ + -T "$DURATION" \ + -P 5 \ + --log \ + --log-prefix="$OUTPUT_DIR/pgbench_$TIMESTAMP" \ + $BENCH_ARGS \ + "$DB_NAME" 2>&1 | tee -a "$RESULTS_FILE" + + # Collect additional statistics + { + echo "" + echo "=== Database Statistics ===" + "$PSQL" --no-psqlrc -h "${config.pgDataDir}" -d "$DB_NAME" -c " + SELECT + schemaname, + relname, + n_tup_ins as inserts, + n_tup_upd as updates, + n_tup_del as deletes, + n_live_tup as live_tuples, + n_dead_tup as dead_tuples + FROM pg_stat_user_tables; + " + + echo "" + echo "=== Index Statistics ===" + "$PSQL" --no-psqlrc -h "${config.pgDataDir}" -d "$DB_NAME" -c " + SELECT + schemaname, + relname, + indexrelname, + idx_scan, + idx_tup_read, + idx_tup_fetch + FROM pg_stat_user_indexes; + " + } >> "$RESULTS_FILE" + + # Clean up + echo "Cleaning up test database: $DB_NAME" + "$DROPDB" -h "${config.pgDataDir}" "$DB_NAME" 2>/dev/null || true + + echo "" + echo "Benchmark completed!" + echo "Results saved to: $RESULTS_FILE" + echo "Transaction logs: $OUTPUT_DIR/pgbench_$TIMESTAMP*" + + # Show summary + echo "" + echo "=== Quick Summary ===" + grep -E "(tps|latency)" "$RESULTS_FILE" | tail -5 + ''; + + # Shared shellHook fragments. Each devShell prepends its own compiler/CFLAGS + # block, then appends the common tail via ${commonHookTail variant}. + commonHookHead = icon: '' + # History configuration + export HISTFILE=.history + export HISTSIZE=1000000 + export HISTFILESIZE=1000000 + + # Clean environment + unset LD_LIBRARY_PATH LD_PRELOAD LIBRARY_PATH C_INCLUDE_PATH CPLUS_INCLUDE_PATH + + # Essential tools in PATH + export PATH="${pkgs.which}/bin:${pkgs.coreutils}/bin:$PATH" + export PS1="$(echo -e '\u${icon}') {\[$(tput sgr0)\]\[\033[38;5;228m\]\w\[$(tput sgr0)\]\[\033[38;5;15m\]} ($(git rev-parse --abbrev-ref HEAD)) \\$ \[$(tput sgr0)\]" + + # Ccache configuration + export PATH=${pkgs.ccache}/bin:$PATH + export CCACHE_COMPILERCHECK=content + # Loosen a few rules so ccache hits across rebuilds with touched headers. + export CCACHE_SLOPPINESS=pch_defines,time_macros,include_file_mtime,include_file_ctime + export CCACHE_DIR=$HOME/.ccache/pg/$(basename $PWD) + mkdir -p "$CCACHE_DIR" + + # Development tools in PATH + export PATH=${pkgs.clang-tools}/bin:$PATH + export PATH=${pkgs.cppcheck}/bin:$PATH + ''; + + # Tail shared by every devShell: PG env vars, GDB, tool PATH, per-process + # setup and alias load. Kernel core_pattern is NOT touched here -- + # run 'pg-enable-cores' explicitly if you need per-PID cores in CWD. + commonHookTail = label: '' + # PostgreSQL environment + export PG_SOURCE_DIR="${config.pgSourceDir}" + export PG_BUILD_DIR="${config.pgBuildDir}" + export PG_BUILD_DIR_VALGRIND="${config.pgBuildDirValgrind}" + export PG_BUILD_DIR_ASAN="${config.pgBuildDirAsan}" + export PG_INSTALL_DIR="${config.pgInstallDir}" + export PG_DATA_DIR="${config.pgDataDir}" + export PG_BENCH_DIR="${config.pgBenchDir}" + export PG_FLAME_DIR="${config.pgFlameDir}" + export PERL_CORE_DIR=$(find ${pkgs.perl} -maxdepth 5 -path "*/CORE" -type d) + + # GDB configuration + export GDBINIT="${gdbConfig}" + + # Performance tools in PATH + export PATH="${flameGraphScript}/bin:${pgbenchScript}/bin:$PATH" + + # Create output directories + mkdir -p "$PG_BENCH_DIR" "$PG_FLAME_DIR" + + # Per-process core dump size limit. Kernel core_pattern is NOT + # touched here -- run 'pg-enable-cores' explicitly when you need + # per-PID cores in CWD. + ulimit -c unlimited + + # Local git excludes + git config core.excludesFile .local-gitignore 2>/dev/null || true + + # Load PostgreSQL development aliases + if [ -f ./pg-aliases.sh ]; then + source ./pg-aliases.sh + else + echo "Warning: pg-aliases.sh not found in current directory" + fi + + echo "" + echo "PostgreSQL Development Environment Ready (${label})" + echo "Run 'pg-info' for available commands" + ''; + + # Development shell (GCC + glibc) + devShell = pkgs.mkShell { + name = "postgresql-dev"; + buildInputs = + (getPostgreSQLDeps false) + ++ [ + flameGraphScript + pgbenchScript + ]; + + shellHook = + (commonHookHead "f121") + + '' + # LLVM configuration + export LLVM_CONFIG="${llvmPkgs.llvm}/bin/llvm-config" + export PATH="${llvmPkgs.llvm}/bin:$PATH" + export PKG_CONFIG_PATH="${llvmPkgs.llvm.dev}/lib/pkgconfig:$PKG_CONFIG_PATH" + export LLVM_DIR="${llvmPkgs.llvm.dev}/lib/cmake/llvm" + export LLVM_ROOT="${llvmPkgs.llvm}" + + # PostgreSQL Development CFLAGS + export CFLAGS="" + export CXXFLAGS="" + + # Python UV + UV_PYTHON_DOWNLOADS=never + + # GCC configuration (default compiler) + export CC="${pkgs.gcc}/bin/gcc" + export CXX="${pkgs.gcc}/bin/g++" + + echo "Environment configured:" + echo " Compiler: $CC" + echo " libc: glibc" + echo " LLVM: $(llvm-config --version 2>/dev/null || echo 'not available')" + '' + + (commonHookTail "GCC + glibc"); + }; + + # Clang + glibc variant + clangDevShell = pkgs.mkShell { + name = "postgresql-clang-glibc"; + buildInputs = + (getPostgreSQLDeps false) + ++ [ + llvmPkgs.clang + llvmPkgs.lld + llvmPkgs.compiler-rt + flameGraphScript + pgbenchScript + ]; + + shellHook = + (commonHookHead "f121") + + '' + # LLVM configuration + export LLVM_CONFIG="${llvmPkgs.llvm}/bin/llvm-config" + export PATH="${llvmPkgs.llvm}/bin:$PATH" + export PKG_CONFIG_PATH="${llvmPkgs.llvm.dev}/lib/pkgconfig:$PKG_CONFIG_PATH" + export LLVM_DIR="${llvmPkgs.llvm.dev}/lib/cmake/llvm" + export LLVM_ROOT="${llvmPkgs.llvm}" + + # Clang + glibc configuration + export CC="${llvmPkgs.clang}/bin/clang" + export CXX="${llvmPkgs.clang}/bin/clang++" + + echo "Environment configured:" + echo " Compiler: $CC" + echo " libc: glibc" + echo " LLVM: $(llvm-config --version 2>/dev/null || echo 'not available')" + '' + + (commonHookTail "Clang + glibc"); + }; + + # GCC + musl variant (cross-compilation) + muslDevShell = pkgs.mkShell { + name = "postgresql-gcc-musl"; + buildInputs = + (getPostgreSQLDeps true) + ++ [ + pkgs.gcc + flameGraphScript + pgbenchScript + ]; + + shellHook = + (commonHookHead "f121") + + '' + # Cross-compilation to musl with GCC + export CC="${pkgs.gcc}/bin/gcc" + export CXX="${pkgs.gcc}/bin/g++" + + export PKG_CONFIG_PATH="${pkgs.pkgsMusl.openssl.dev}/lib/pkgconfig:${pkgs.pkgsMusl.zlib.dev}/lib/pkgconfig:${pkgs.pkgsMusl.icu.dev}/lib/pkgconfig" + export CFLAGS="-ggdb -Og -fno-omit-frame-pointer -D_FORTIFY_SOURCE=1 -I${pkgs.pkgsMusl.stdenv.cc.libc}/include" + export CXXFLAGS="-ggdb -Og -fno-omit-frame-pointer -D_FORTIFY_SOURCE=1 -I${pkgs.pkgsMusl.stdenv.cc.libc}/include" + export LDFLAGS="-L${pkgs.pkgsMusl.stdenv.cc.libc}/lib -static-libgcc" + + echo "Environment configured:" + echo " Compiler: $CC" + echo " libc: musl (cross-compilation)" + '' + + (commonHookTail "GCC + musl"); + }; + + # Clang + musl variant (cross-compilation) + clangMuslDevShell = pkgs.mkShell { + name = "postgresql-clang-musl"; + buildInputs = + (getPostgreSQLDeps true) + ++ [ + llvmPkgs.clang + llvmPkgs.lld + flameGraphScript + pgbenchScript + ]; + + shellHook = + (commonHookHead "f121") + + '' + # Cross-compilation to musl with clang + export CC="${llvmPkgs.clang}/bin/clang" + export CXX="${llvmPkgs.clang}/bin/clang++" + + export PKG_CONFIG_PATH="${pkgs.pkgsMusl.openssl.dev}/lib/pkgconfig:${pkgs.pkgsMusl.zlib.dev}/lib/pkgconfig:${pkgs.pkgsMusl.icu.dev}/lib/pkgconfig" + export CFLAGS="--target=x86_64-linux-musl -ggdb -Og -fno-omit-frame-pointer -D_FORTIFY_SOURCE=1 -I${pkgs.pkgsMusl.stdenv.cc.libc}/include" + export CXXFLAGS="--target=x86_64-linux-musl -ggdb -Og -fno-omit-frame-pointer -D_FORTIFY_SOURCE=1 -I${pkgs.pkgsMusl.stdenv.cc.libc}/include" + export LDFLAGS="--target=x86_64-linux-musl -L${pkgs.pkgsMusl.stdenv.cc.libc}/lib -fuse-ld=lld" + + echo "Environment configured:" + echo " Compiler: $CC" + echo " libc: musl (cross-compilation)" + '' + + (commonHookTail "Clang + musl"); + }; +in { + inherit devShell clangDevShell muslDevShell clangMuslDevShell gdbConfig flameGraphScript pgbenchScript; +} diff --git a/src/test/regress/pg_regress.c b/src/test/regress/pg_regress.c index 13944701bc7b0..278cbfb735ce0 100644 --- a/src/test/regress/pg_regress.c +++ b/src/test/regress/pg_regress.c @@ -1243,7 +1243,7 @@ spawn_process(const char *cmdline) char *cmdline2; cmdline2 = psprintf("exec %s", cmdline); - execl(shellprog, shellprog, "-c", cmdline2, (char *) NULL); + execlp(shellprog, shellprog, "-c", cmdline2, (char *) NULL); /* Not using the normal bail() here as we want _exit */ bail_noatexit("could not exec \"%s\": %m", shellprog); } diff --git a/src/tools/pgindent/pgindent b/src/tools/pgindent/pgindent index fa2041dc34b93..92526f6248581 100755 --- a/src/tools/pgindent/pgindent +++ b/src/tools/pgindent/pgindent @@ -1,4 +1,4 @@ -#!/usr/bin/perl +#!/usr/bin/env perl # Copyright (c) 2021-2026, PostgreSQL Global Development Group From 763bdf0ec759c2375670586fce74c39cddb53311 Mon Sep 17 00:00:00 2001 From: Tomas Vondra Date: Wed, 3 Jun 2026 16:30:28 +0200 Subject: [PATCH 04/17] Add shmem_populate and shmem_interleave GUCs - shmem_populate - Forces mmap() with MAP_POPULATE, which faults all memory pages backing the shared memory segment. - shmem_interleave - Applies NUMA interleaving on the whole shared memory segment, to balance allocations between nodes. --- src/backend/port/sysv_shmem.c | 47 +++++++++++++++++++++++ src/backend/utils/misc/guc_parameters.dat | 14 +++++++ src/include/miscadmin.h | 4 ++ 3 files changed, 65 insertions(+) diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c index 2e3886cf9fe49..9eaff838a04f6 100644 --- a/src/backend/port/sysv_shmem.c +++ b/src/backend/port/sysv_shmem.c @@ -27,6 +27,10 @@ #include #include +#ifdef USE_LIBNUMA +#include +#endif + #include "miscadmin.h" #include "port/pg_bitutils.h" #include "portability/mem.h" @@ -98,6 +102,10 @@ void *UsedShmemSegAddr = NULL; static Size AnonymousShmemSize; static void *AnonymousShmem = NULL; +/* GUCs */ +bool shmem_populate = false; /* MAP_POPULATE */ +bool shmem_interleave = false; /* NUMA interleaving */ + static void *InternalIpcMemoryCreate(IpcMemoryKey memKey, Size size); static void IpcMemoryDetach(int status, Datum shmaddr); static void IpcMemoryDelete(int status, Datum shmId); @@ -604,6 +612,21 @@ CreateAnonymousSegment(Size *size) int mmap_errno = 0; int mmap_flags = MAP_SHARED | MAP_ANONYMOUS | MAP_HASSEMAPHORE; + /* If requested, populate the shared memory by MAP_POPULATE. */ + if (shmem_populate) + mmap_flags |= MAP_POPULATE; + +#ifdef USE_LIBNUMA + /* + * If requested, interleave the shared memory by setting a memory policy + * before the mmap() call. This really matters only with MAP_POPULATE, + * because without page faults the memory does not actually get placed + * to the nodes. But without MAP_POPULATE it's virtually free. + */ + if (shmem_interleave) + numa_set_interleave_mask(numa_all_nodes_ptr); +#endif + #ifndef MAP_HUGETLB /* PGSharedMemoryCreate should have dealt with this case */ Assert(huge_pages != HUGE_PAGES_ON); @@ -665,6 +688,30 @@ CreateAnonymousSegment(Size *size) allocsize) : 0)); } +#ifdef USE_LIBNUMA + /* + * If set the policy to interleaving by numa_set_membind(), undo it now by + * setting the policy to localalloc. With MAP_POPULATE, all the pages were + * faulted and are now interleaved on the available nodes. + * + * To handle the case without MAP_POPULATE, apply the interleaving policy to + * the shared memory segment allocated by mmap() before touching it in any + * way, so that it gets placed on the correct node on first access. + * + * This matters especially with huge pages, where it's possible to run out + * of huge pages on some nodes and then crash. By explicitly interleaving + * the whole segment, that's much less likely. + */ + if (shmem_interleave) + { + /* undo the policy set by numa_set_membind() earlier */ + numa_set_localalloc(); + + /* set interleaving policy for not yet faulted memory */ + numa_interleave_memory(ptr, allocsize, numa_all_nodes_ptr); + } +#endif + *size = allocsize; return ptr; } diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat index 380679a76c096..707a77a0f113e 100644 --- a/src/backend/utils/misc/guc_parameters.dat +++ b/src/backend/utils/misc/guc_parameters.dat @@ -731,6 +731,20 @@ ifdef => 'DEBUG_NODE_TESTS_ENABLED', }, +{ name => 'debug_shmem_interleave', type => 'bool', context => 'PGC_SUSET', group => 'DEVELOPER_OPTIONS', + short_desc => 'Forces interleaving for the whole shared memory segment.', + flags => 'GUC_NOT_IN_SAMPLE', + variable => 'shmem_interleave', + boot_val => 'false' +}, + +{ name => 'debug_shmem_populate', type => 'bool', context => 'PGC_SUSET', group => 'DEVELOPER_OPTIONS', + short_desc => 'Populates (faults) the whole shared memory segment using MAP_POPULATE.', + flags => 'GUC_NOT_IN_SAMPLE', + variable => 'shmem_populate', + boot_val => 'false' +}, + { name => 'debug_write_read_parse_plan_trees', type => 'bool', context => 'PGC_SUSET', group => 'DEVELOPER_OPTIONS', short_desc => 'Set this to force all parse and plan trees to be passed through outfuncs.c/readfuncs.c, to facilitate catching errors and omissions in those modules.', flags => 'GUC_NOT_IN_SAMPLE', diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h index 8d6aacc4d5ac2..ef72549ebc66e 100644 --- a/src/include/miscadmin.h +++ b/src/include/miscadmin.h @@ -213,6 +213,10 @@ extern PGDLLIMPORT Oid MyDatabaseTableSpace; extern PGDLLIMPORT bool MyDatabaseHasLoginEventTriggers; +extern PGDLLIMPORT bool shmem_populate; +extern PGDLLIMPORT bool shmem_interleave; + + /* * Date/Time Configuration * From a958a766a580cb78ec2ce6ce39cb8efbd83acc93 Mon Sep 17 00:00:00 2001 From: Tomas Vondra Date: Tue, 2 Jun 2026 22:09:33 +0200 Subject: [PATCH 05/17] Infrastructure for partitioning of shared buffers The patch introduces a simple "registry" of buffer partitions, keeping track of the first/last buffer, etc. This serves as a source of truth for later patches (e.g. to partition clock-sweep or to make the partitioning NUMA-aware). The registry is a small array of BufferPartition entries in shared memory, with partitions sized to be a fair share of shared buffers. Notes: * Maybe the number of partitions should be configurable? Right now it's hard-coded as 4, but testing shows increasing to e.g. 16) can be beneficial. * This partitioning is independent of the partitions defined in lwlock.h, which defines 128 partitions to reduce lock conflict on the buffer mapping hashtable. The number of partitions introduced by this patch is expected to be much lower (a dozen or so). * The buffers are divided as evenly as possible, with the first couple partitions possibly getting an extra buffer. --- contrib/pg_buffercache/Makefile | 3 +- contrib/pg_buffercache/meson.build | 1 + .../pg_buffercache--1.7--1.8.sql | 23 +++ contrib/pg_buffercache/pg_buffercache.control | 2 +- contrib/pg_buffercache/pg_buffercache_pages.c | 86 +++++++++++ src/backend/storage/buffer/buf_init.c | 142 ++++++++++++++++++ src/include/storage/buf_internals.h | 5 + src/include/storage/bufmgr.h | 19 +++ src/tools/pgindent/typedefs.list | 2 + 9 files changed, 281 insertions(+), 2 deletions(-) create mode 100644 contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql diff --git a/contrib/pg_buffercache/Makefile b/contrib/pg_buffercache/Makefile index 0e618f66aec6e..7fd5cdfc43d66 100644 --- a/contrib/pg_buffercache/Makefile +++ b/contrib/pg_buffercache/Makefile @@ -9,7 +9,8 @@ EXTENSION = pg_buffercache DATA = pg_buffercache--1.2.sql pg_buffercache--1.2--1.3.sql \ pg_buffercache--1.1--1.2.sql pg_buffercache--1.0--1.1.sql \ pg_buffercache--1.3--1.4.sql pg_buffercache--1.4--1.5.sql \ - pg_buffercache--1.5--1.6.sql pg_buffercache--1.6--1.7.sql + pg_buffercache--1.5--1.6.sql pg_buffercache--1.6--1.7.sql \ + pg_buffercache--1.7--1.8.sql PGFILEDESC = "pg_buffercache - monitoring of shared buffer cache in real-time" REGRESS = pg_buffercache pg_buffercache_numa diff --git a/contrib/pg_buffercache/meson.build b/contrib/pg_buffercache/meson.build index e681205abb2d8..361628b8bea42 100644 --- a/contrib/pg_buffercache/meson.build +++ b/contrib/pg_buffercache/meson.build @@ -25,6 +25,7 @@ install_data( 'pg_buffercache--1.4--1.5.sql', 'pg_buffercache--1.5--1.6.sql', 'pg_buffercache--1.6--1.7.sql', + 'pg_buffercache--1.7--1.8.sql', 'pg_buffercache.control', kwargs: contrib_data_args, ) diff --git a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql new file mode 100644 index 0000000000000..d62b8339bfcfc --- /dev/null +++ b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql @@ -0,0 +1,23 @@ +-- complain if script is sourced in psql, rather than via CREATE EXTENSION +\echo Use "ALTER EXTENSION pg_buffercache UPDATE TO '1.8'" to load this file. \quit + +-- Register the new functions. +CREATE OR REPLACE FUNCTION pg_buffercache_partitions() +RETURNS SETOF RECORD +AS 'MODULE_PATHNAME', 'pg_buffercache_partitions' +LANGUAGE C PARALLEL SAFE; + +-- Create a view for convenient access. +CREATE VIEW pg_buffercache_partitions AS + SELECT P.* FROM pg_buffercache_partitions() AS P + (partition integer, -- partition index + num_buffers integer, -- number of buffers in the partition + first_buffer integer, -- first buffer of partition + last_buffer integer); -- last buffer of partition + +-- Don't want these to be available to public. +REVOKE ALL ON FUNCTION pg_buffercache_partitions() FROM PUBLIC; +REVOKE ALL ON pg_buffercache_partitions FROM PUBLIC; + +GRANT EXECUTE ON FUNCTION pg_buffercache_partitions() TO pg_monitor; +GRANT SELECT ON pg_buffercache_partitions TO pg_monitor; diff --git a/contrib/pg_buffercache/pg_buffercache.control b/contrib/pg_buffercache/pg_buffercache.control index 11499550945ee..d2fa8ba53ba9f 100644 --- a/contrib/pg_buffercache/pg_buffercache.control +++ b/contrib/pg_buffercache/pg_buffercache.control @@ -1,5 +1,5 @@ # pg_buffercache extension comment = 'examine the shared buffer cache' -default_version = '1.7' +default_version = '1.8' module_pathname = '$libdir/pg_buffercache' relocatable = true diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c index 510455998aa74..6c9838b4efc72 100644 --- a/contrib/pg_buffercache/pg_buffercache_pages.c +++ b/contrib/pg_buffercache/pg_buffercache_pages.c @@ -31,6 +31,7 @@ #define NUM_BUFFERCACHE_MARK_DIRTY_ALL_ELEM 3 #define NUM_BUFFERCACHE_OS_PAGES_ELEM 3 +#define NUM_BUFFERCACHE_PARTITIONS_ELEM 4 PG_MODULE_MAGIC_EXT( .name = "pg_buffercache", @@ -77,6 +78,7 @@ PG_FUNCTION_INFO_V1(pg_buffercache_evict_all); PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty); PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_relation); PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_all); +PG_FUNCTION_INFO_V1(pg_buffercache_partitions); /* Only need to touch memory once per backend process lifetime */ @@ -922,3 +924,87 @@ pg_buffercache_mark_dirty_all(PG_FUNCTION_ARGS) PG_RETURN_DATUM(result); } + +/* + * Inquire about partitioning of shared buffers. + */ +Datum +pg_buffercache_partitions(PG_FUNCTION_ARGS) +{ + FuncCallContext *funcctx; + MemoryContext oldcontext; + TupleDesc tupledesc; + TupleDesc expected_tupledesc; + HeapTuple tuple; + Datum result; + + if (SRF_IS_FIRSTCALL()) + { + funcctx = SRF_FIRSTCALL_INIT(); + + /* Switch context when allocating stuff to be used in later calls */ + oldcontext = MemoryContextSwitchTo(funcctx->multi_call_memory_ctx); + + if (get_call_result_type(fcinfo, NULL, &expected_tupledesc) != TYPEFUNC_COMPOSITE) + elog(ERROR, "return type must be a row type"); + + if (expected_tupledesc->natts != NUM_BUFFERCACHE_PARTITIONS_ELEM) + elog(ERROR, "incorrect number of output arguments"); + + /* Construct a tuple descriptor for the result rows. */ + tupledesc = CreateTemplateTupleDesc(expected_tupledesc->natts); + TupleDescInitEntry(tupledesc, (AttrNumber) 1, "partition", + INT4OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 2, "num_buffers", + INT4OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 3, "first_buffer", + INT4OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 4, "last_buffer", + INT4OID, -1, 0); + + funcctx->user_fctx = BlessTupleDesc(tupledesc); + + /* Return to original context when allocating transient memory */ + MemoryContextSwitchTo(oldcontext); + + /* Set max calls and remember the user function context. */ + funcctx->max_calls = BufferPartitionCount(); + } + + funcctx = SRF_PERCALL_SETUP(); + + if (funcctx->call_cntr < funcctx->max_calls) + { + uint32 i = funcctx->call_cntr; + + int num_buffers, + first_buffer, + last_buffer; + + Datum values[NUM_BUFFERCACHE_PARTITIONS_ELEM]; + bool nulls[NUM_BUFFERCACHE_PARTITIONS_ELEM]; + + BufferPartitionGet(i, &num_buffers, + &first_buffer, &last_buffer); + + values[0] = Int32GetDatum(i); + nulls[0] = false; + + values[1] = Int32GetDatum(num_buffers); + nulls[1] = false; + + values[2] = Int32GetDatum(first_buffer); + nulls[2] = false; + + values[3] = Int32GetDatum(last_buffer); + nulls[3] = false; + + /* Build and return the tuple. */ + tuple = heap_form_tuple((TupleDesc) funcctx->user_fctx, values, nulls); + result = HeapTupleGetDatum(tuple); + + SRF_RETURN_NEXT(funcctx, result); + } + else + SRF_RETURN_DONE(funcctx); +} diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c index 9c5f2449cf399..24f8233cabf11 100644 --- a/src/backend/storage/buffer/buf_init.c +++ b/src/backend/storage/buffer/buf_init.c @@ -25,9 +25,11 @@ BufferDescPadded *BufferDescriptors; char *BufferBlocks; ConditionVariableMinimallyPadded *BufferIOCVArray; CkptSortItem *CkptBufferIds; +BufferPartitions *BufferPartitionsRegistry; static void BufferManagerShmemRequest(void *arg); static void BufferManagerShmemInit(void *arg); +static void BufferPartitionsInit(void); const ShmemCallbacks BufferManagerShmemCallbacks = { .request_fn = BufferManagerShmemRequest, @@ -66,6 +68,9 @@ const ShmemCallbacks BufferManagerShmemCallbacks = { * multiple times. Check the PrivateRefCount infrastructure in bufmgr.c. */ +/* number of buffer partitions */ +#define NUM_CLOCK_SWEEP_PARTITIONS 4 + /* * Register shared memory area for the buffer pool. @@ -94,6 +99,13 @@ BufferManagerShmemRequest(void *arg) .ptr = (void **) &BufferIOCVArray, ); + ShmemRequestStruct(.name = "Buffer Partition Registry", + .size = NUM_CLOCK_SWEEP_PARTITIONS * sizeof(BufferPartition), + /* Align descriptors to a cacheline boundary. */ + .alignment = PG_CACHE_LINE_SIZE, + .ptr = (void **) &BufferPartitionsRegistry, + ); + /* * The array used to sort to-be-checkpointed buffer ids is located in * shared memory, to avoid having to allocate significant amounts of @@ -116,6 +128,12 @@ BufferManagerShmemRequest(void *arg) static void BufferManagerShmemInit(void *arg) { + /* + * Initialize the buffer partition registry first, before other parts + * have a chance to touch the memory. + */ + BufferPartitionsInit(); + /* * Initialize all the buffer headers. */ @@ -136,3 +154,127 @@ BufferManagerShmemInit(void *arg) ConditionVariableInit(BufferDescriptorGetIOCV(buf)); } } + +/* + * Sanity checks of buffers partitions - there must be no gaps, it must cover + * the whole range of buffers, etc. + */ +static void +AssertCheckBufferPartitions(void) +{ +#ifdef USE_ASSERT_CHECKING + int num_buffers = 0; + + Assert(BufferPartitionsRegistry->npartitions > 0); + + for (int i = 0; i < BufferPartitionsRegistry->npartitions; i++) + { + BufferPartition *part = &BufferPartitionsRegistry->partitions[i]; + + /* + * We can get a single-buffer partition, if the sizing forces the last + * partition to be just one buffer. But it's unlikely (and + * undesirable). + */ + Assert(part->first_buffer <= part->last_buffer); + Assert((part->last_buffer - part->first_buffer + 1) == part->num_buffers); + + num_buffers += part->num_buffers; + + /* + * The first partition needs to start on buffer 0. Later partitions + * need to be contiguous, without skipping any buffers. + */ + if (i == 0) + { + Assert(part->first_buffer == 0); + } + else + { + BufferPartition *prev = &BufferPartitionsRegistry->partitions[i - 1]; + + Assert((part->first_buffer - 1) == prev->last_buffer); + } + + /* the last partition needs to end on buffer (NBuffers - 1) */ + if (i == (BufferPartitionsRegistry->npartitions - 1)) + { + Assert(part->last_buffer == (NBuffers - 1)); + } + } + + Assert(num_buffers == NBuffers); +#endif +} + +/* + * BufferPartitionsInit + * Initialize registry of buffer partitions. + */ +static void +BufferPartitionsInit(void) +{ + int buffer = 0; + + /* number of buffers per partition (make sure to not overflow) */ + int part_buffers = NBuffers / NUM_CLOCK_SWEEP_PARTITIONS; + int remaining_buffers = NBuffers % NUM_CLOCK_SWEEP_PARTITIONS; + + BufferPartitionsRegistry->npartitions = NUM_CLOCK_SWEEP_PARTITIONS; + + for (int n = 0; n < BufferPartitionsRegistry->npartitions; n++) + { + BufferPartition *part = &BufferPartitionsRegistry->partitions[n]; + + int num_buffers = part_buffers; + if (n < remaining_buffers) + num_buffers += 1; + + remaining_buffers -= num_buffers; + + Assert((num_buffers > 0) && (num_buffers <= part_buffers)); + Assert((buffer >= 0) && (buffer < NBuffers)); + + part->num_buffers = num_buffers; + part->first_buffer = buffer; + part->last_buffer = buffer + (num_buffers - 1); + + buffer += num_buffers; + } + + AssertCheckBufferPartitions(); +} + +/* + * BufferPartitionCount + * Returns the number of partitions created. + */ +int +BufferPartitionCount(void) +{ + return BufferPartitionsRegistry->npartitions; +} + +/* + * BufferPartitionGet + * Returns information about a partition at the provided index. + * + * The returned information is first/last buffer, number of buffers. + */ +void +BufferPartitionGet(int idx, int *num_buffers, + int *first_buffer, int *last_buffer) +{ + if ((idx >= 0) && (idx < BufferPartitionsRegistry->npartitions)) + { + BufferPartition *part = &BufferPartitionsRegistry->partitions[idx]; + + *num_buffers = part->num_buffers; + *first_buffer = part->first_buffer; + *last_buffer = part->last_buffer; + + return; + } + + elog(ERROR, "invalid partition index"); +} diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h index 726ab75dfe992..115ce68802c01 100644 --- a/src/include/storage/buf_internals.h +++ b/src/include/storage/buf_internals.h @@ -411,8 +411,13 @@ typedef struct WritebackContext /* in buf_init.c */ extern PGDLLIMPORT BufferDescPadded *BufferDescriptors; +extern PGDLLIMPORT BufferPartitions *BufferPartitionsRegistry; extern PGDLLIMPORT ConditionVariableMinimallyPadded *BufferIOCVArray; +extern int BufferPartitionCount(void); +extern void BufferPartitionGet(int idx, int *num_buffers, + int *first_buffer, int *last_buffer); + /* in localbuf.c */ extern PGDLLIMPORT BufferDesc *LocalBufferDescriptors; diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h index 6837b35fc6d0b..79a3f44747ad2 100644 --- a/src/include/storage/bufmgr.h +++ b/src/include/storage/bufmgr.h @@ -155,6 +155,25 @@ struct ReadBuffersOperation typedef struct ReadBuffersOperation ReadBuffersOperation; +/* + * information about one partition of shared buffers + * + * first/last buffer - the values are inclusive + */ +typedef struct BufferPartition +{ + int num_buffers; /* number of buffers */ + int first_buffer; /* first buffer of partition */ + int last_buffer; /* last buffer of partition */ +} BufferPartition; + +/* an array of information about all partitions */ +typedef struct BufferPartitions +{ + int npartitions; /* number of partitions */ + BufferPartition partitions[FLEXIBLE_ARRAY_MEMBER]; +} BufferPartitions; + /* to avoid having to expose buf_internals.h here */ typedef struct WritebackContext WritebackContext; diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 3bed561f7f23a..2ad38118af8a4 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -361,6 +361,8 @@ BufferHeapTupleTableSlot BufferLockMode BufferLookupEnt BufferManagerRelation +BufferPartition +BufferPartitions BufferStrategyControl BufferTag BufferUsage From e44c37eb1b7f0bbafe052dfede77418f79f968ae Mon Sep 17 00:00:00 2001 From: Tomas Vondra Date: Sat, 10 Oct 2026 23:30:08 -0400 Subject: [PATCH 06/17] NUMA: shared buffers partitioning Ensure shared buffers are allocated from all NUMA nodes, in a balanced way, instead of just using the node where Postgres initially starts, or where the kernel decides to migrate the page, etc. In cases like pre-warming a database from a single worker (e.g. using pg_prewarm), we may end up with severely unbalanced memory distribution (with most memory located on a single NUMA node). Unbalanced allocation may put a lot of pressure on the memory system on a small number of NUMA nodes, limiting the bandwidth etc. With zone_reclaim, the kernel would eventually move some of the memory to other nodes, but that tends to take a long time and is unpredictable. This change forces even distribution of shared buffers on all NUMA nodes, improving predictability, reducing the time needed for warmup during benchmarking, etc. It's also less dependent on what the CPU scheduler decides to do (which cores get used for the warmup.) The effect is similar to numactl --interleave=all in that the buffers are distributed on the NUMA nodes evenly, but there's also a number of important differences. Firstly, it's applied only to shared buffers (and buffer descriptors), not to the whole shared memory segment. It's possible to enable memory interleaving using the shmem_interleave GUC, introduced in an earlier patch in this series. NUMA works at the granularity of a memory page, which is typically either 4K or 2MB (hugepage), but other sizes are possible. For systems where NUMA matters, we expect large amounts of memory (hundreds of gigabytes) and hugepages enabled. But not necessarily. The partitioning scheme is best-effort with respect to memory page size. The shared buffers do not "align" with memory pages (i.e. a partition may not end at the memory page boundary), in which case we simply locate just the section of the partition with complete memory pages. This means there may be ~one unmapped memory page between partitions. Considering the expected amounts of memory, this is negligible, and the alternative would be a significant amount of complexity to align the pages and enforce "allowed" partition sizes. Buffer descriptors are affected by this too, and the effect may be more significant, simply because the descriptors are much smaller (~64B). So the array is smaller, and a single 2MB memory page is worth ~32K buffer descriptors. But with large systems it's still negligible. The "buffer partitions" may not be 1:1 with NUMA nodes. We want to allow clock-sweep partitioning even on non-NUMA systems, or when running only on a small number of NUMA nodes. There's a minimal number of partitions (default: 4), and a node may get multiple partitions. Nodes always get the same number of partitions (e.g. with 3 NUMA nodes there will be 6 partitions in total, as each node gets 2 partitions). The feature is enabled by dshared_buffers_numa GUC (default: false). --- .../pg_buffercache--1.7--1.8.sql | 1 + contrib/pg_buffercache/pg_buffercache_pages.c | 24 +- src/backend/storage/buffer/buf_init.c | 251 ++++++++++++++++-- src/backend/storage/buffer/freelist.c | 9 + src/backend/utils/misc/guc_parameters.dat | 6 + src/backend/utils/misc/postgresql.conf.sample | 1 + src/include/port/pg_numa.h | 7 + src/include/storage/buf_internals.h | 16 +- src/include/storage/bufmgr.h | 8 + src/port/pg_numa.c | 112 ++++++++ 10 files changed, 399 insertions(+), 36 deletions(-) diff --git a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql index d62b8339bfcfc..a6e49fd165291 100644 --- a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql +++ b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql @@ -11,6 +11,7 @@ LANGUAGE C PARALLEL SAFE; CREATE VIEW pg_buffercache_partitions AS SELECT P.* FROM pg_buffercache_partitions() AS P (partition integer, -- partition index + numa_node integer, -- NUMA node of the partitioon num_buffers integer, -- number of buffers in the partition first_buffer integer, -- first buffer of partition last_buffer integer); -- last buffer of partition diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c index 6c9838b4efc72..46b6b85a2e3f1 100644 --- a/contrib/pg_buffercache/pg_buffercache_pages.c +++ b/contrib/pg_buffercache/pg_buffercache_pages.c @@ -31,7 +31,7 @@ #define NUM_BUFFERCACHE_MARK_DIRTY_ALL_ELEM 3 #define NUM_BUFFERCACHE_OS_PAGES_ELEM 3 -#define NUM_BUFFERCACHE_PARTITIONS_ELEM 4 +#define NUM_BUFFERCACHE_PARTITIONS_ELEM 5 PG_MODULE_MAGIC_EXT( .name = "pg_buffercache", @@ -955,11 +955,13 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) tupledesc = CreateTemplateTupleDesc(expected_tupledesc->natts); TupleDescInitEntry(tupledesc, (AttrNumber) 1, "partition", INT4OID, -1, 0); - TupleDescInitEntry(tupledesc, (AttrNumber) 2, "num_buffers", + TupleDescInitEntry(tupledesc, (AttrNumber) 2, "numa_node", INT4OID, -1, 0); - TupleDescInitEntry(tupledesc, (AttrNumber) 3, "first_buffer", + TupleDescInitEntry(tupledesc, (AttrNumber) 3, "num_buffers", INT4OID, -1, 0); - TupleDescInitEntry(tupledesc, (AttrNumber) 4, "last_buffer", + TupleDescInitEntry(tupledesc, (AttrNumber) 4, "first_buffer", + INT4OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 5, "last_buffer", INT4OID, -1, 0); funcctx->user_fctx = BlessTupleDesc(tupledesc); @@ -977,28 +979,32 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) { uint32 i = funcctx->call_cntr; - int num_buffers, + int numa_node, + num_buffers, first_buffer, last_buffer; Datum values[NUM_BUFFERCACHE_PARTITIONS_ELEM]; bool nulls[NUM_BUFFERCACHE_PARTITIONS_ELEM]; - BufferPartitionGet(i, &num_buffers, + BufferPartitionGet(i, &numa_node, &num_buffers, &first_buffer, &last_buffer); values[0] = Int32GetDatum(i); nulls[0] = false; - values[1] = Int32GetDatum(num_buffers); + values[1] = Int32GetDatum(numa_node); nulls[1] = false; - values[2] = Int32GetDatum(first_buffer); + values[2] = Int32GetDatum(num_buffers); nulls[2] = false; - values[3] = Int32GetDatum(last_buffer); + values[3] = Int32GetDatum(first_buffer); nulls[3] = false; + values[4] = Int32GetDatum(last_buffer); + nulls[4] = false; + /* Build and return the tuple. */ tuple = heap_form_tuple((TupleDesc) funcctx->user_fctx, values, nulls); result = HeapTupleGetDatum(tuple); diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c index 24f8233cabf11..8d49eb7b1bf27 100644 --- a/src/backend/storage/buffer/buf_init.c +++ b/src/backend/storage/buffer/buf_init.c @@ -14,12 +14,20 @@ */ #include "postgres.h" +#ifdef USE_LIBNUMA +#include +#include +#endif + +#include "port/pg_numa.h" #include "storage/aio.h" #include "storage/buf_internals.h" #include "storage/bufmgr.h" #include "storage/proclist.h" #include "storage/shmem.h" #include "storage/subsystems.h" +#include "utils/guc_hooks.h" +#include "utils/varlena.h" BufferDescPadded *BufferDescriptors; char *BufferBlocks; @@ -68,9 +76,12 @@ const ShmemCallbacks BufferManagerShmemCallbacks = { * multiple times. Check the PrivateRefCount infrastructure in bufmgr.c. */ -/* number of buffer partitions */ -#define NUM_CLOCK_SWEEP_PARTITIONS 4 +/* + * Minimum number of buffer partitions, no matter the number of NUMA nodes. + */ +#define MIN_BUFFER_PARTITIONS 4 +bool shared_buffers_numa = false; /* * Register shared memory area for the buffer pool. @@ -78,6 +89,10 @@ const ShmemCallbacks BufferManagerShmemCallbacks = { static void BufferManagerShmemRequest(void *arg) { + int nparts; + + BufferPartitionsCalculate(NULL, &nparts, NULL); + ShmemRequestStruct(.name = "Buffer Descriptors", .size = NBuffers * sizeof(BufferDescPadded), /* Align descriptors to a cacheline boundary. */ @@ -100,7 +115,7 @@ BufferManagerShmemRequest(void *arg) ); ShmemRequestStruct(.name = "Buffer Partition Registry", - .size = NUM_CLOCK_SWEEP_PARTITIONS * sizeof(BufferPartition), + .size = nparts * sizeof(BufferPartition), /* Align descriptors to a cacheline boundary. */ .alignment = PG_CACHE_LINE_SIZE, .ptr = (void **) &BufferPartitionsRegistry, @@ -131,6 +146,10 @@ BufferManagerShmemInit(void *arg) /* * Initialize the buffer partition registry first, before other parts * have a chance to touch the memory. + * + * Also moves memory to different NUMA nodes (if enabled by a GUC). + * Do this before the loop that initializes buffer headers etc. which + * may fault some of the memory pages etc. */ BufferPartitionsInit(); @@ -216,35 +235,210 @@ BufferPartitionsInit(void) { int buffer = 0; - /* number of buffers per partition (make sure to not overflow) */ - int part_buffers = NBuffers / NUM_CLOCK_SWEEP_PARTITIONS; - int remaining_buffers = NBuffers % NUM_CLOCK_SWEEP_PARTITIONS; + int nnodes, + npartitions, + npartitions_per_node; - BufferPartitionsRegistry->npartitions = NUM_CLOCK_SWEEP_PARTITIONS; + int buffers_per_partition, + buffers_remaining; - for (int n = 0; n < BufferPartitionsRegistry->npartitions; n++) - { - BufferPartition *part = &BufferPartitionsRegistry->partitions[n]; + /* calculate partitioning parameters */ + BufferPartitionsCalculate(&nnodes, &npartitions, &npartitions_per_node); + + /* paranoia */ + Assert(nnodes > 0); + Assert(npartitions >= MIN_BUFFER_PARTITIONS); + Assert((npartitions % nnodes) == 0); + Assert((npartitions_per_node * nnodes) == npartitions); - int num_buffers = part_buffers; - if (n < remaining_buffers) - num_buffers += 1; + BufferPartitionsRegistry->nnodes = nnodes; + BufferPartitionsRegistry->npartitions = npartitions; + BufferPartitionsRegistry->npartitions_per_node = npartitions_per_node; - remaining_buffers -= num_buffers; + /* regular partition size, the first couple get an extra buffer */ + buffers_per_partition = (NBuffers / npartitions); + buffers_remaining = (NBuffers % buffers_per_partition); - Assert((num_buffers > 0) && (num_buffers <= part_buffers)); - Assert((buffer >= 0) && (buffer < NBuffers)); + /* should have all the buffers */ + Assert((buffers_per_partition * npartitions + buffers_remaining) == NBuffers); - part->num_buffers = num_buffers; - part->first_buffer = buffer; - part->last_buffer = buffer + (num_buffers - 1); + /* + * Now walk the partitions, and set the buffer range. Optionally, place + * the partitions on a given node (for all partitions at once). + */ + for (int n = 0; n < nnodes; n++) + { + for (int p = 0; p < npartitions_per_node; p++) + { + int idx = (n * npartitions_per_node) + p; + BufferPartition *part = &BufferPartitionsRegistry->partitions[idx]; + + /* + * Assign to the NUMA node, but only with shared_buffers_numa=on. + * + * XXX we should get an actual node ID from the mask, in case the + * task is restricted to only some nodes. + */ + part->numa_node = (shared_buffers_numa) ? n : -1; + + /* The first couple partitions may get an extra buffer. */ + part->num_buffers = buffers_per_partition; + if (idx < buffers_remaining) + part->num_buffers += 1; + + /* remember the buffer range */ + part->first_buffer = buffer; + part->last_buffer = buffer + (part->num_buffers - 1); + + /* remember start of the next partition */ + buffer += part->num_buffers; + } - buffer += num_buffers; +#ifdef USE_LIBNUMA + /* + * Now try to locate buffers and buffer descriptors to the node (all + * partitions for the node at once). + */ + if (shared_buffers_numa) + { + Size numa_page_size = pg_numa_page_size(); + + int part_first, + part_last, + buff_first, + buff_last; + + char *startptr, + *endptr; + + /* first/last partition for this node */ + part_first = (n * npartitions_per_node); + part_last = part_first + (npartitions_per_node - 1); + + /* buffers (blocks) */ + + /* first/last buffer */ + buff_first = BufferPartitionsRegistry->partitions[part_first].first_buffer; + buff_last = BufferPartitionsRegistry->partitions[part_last].last_buffer; + + /* beginning of the first block, end of last block */ + startptr = BufferBlocks + ((Size) buff_first * BLCKSZ); + endptr = BufferBlocks + ((Size) (buff_last + 1) * BLCKSZ); + + /* print some warnings when the partitions are not aligned */ + if ((startptr != (char *) TYPEALIGN(numa_page_size, startptr)) || + (endptr != (char *) TYPEALIGN(numa_page_size, endptr))) + { + elog(WARNING, "buffers for node %d not well aligned [%p,%p]", + n, startptr, endptr); + } + + /* best effort: align the pointers, so that the mbind() works */ + startptr = (char *) TYPEALIGN_DOWN(numa_page_size, startptr); + + /* the last partition aligns to the end of the buffer */ + if (n == (nnodes - 1)) + endptr = (char *) TYPEALIGN(numa_page_size, endptr); + else + endptr = (char *) TYPEALIGN_DOWN(numa_page_size, endptr); + + /* XXX or should we use pg_numa_move_to_node? */ + if (pg_numa_bind_to_node(startptr, endptr, n) != 0) + elog(WARNING, "failed to bind shared buffers partition to node %d", n); + + /* buffer descriptors */ + + /* beginning of the first descriptor, end of last descriptor */ + startptr = (char *) &BufferDescriptors[buff_first]; + endptr = (char *) &BufferDescriptors[buff_last] + 1; + + /* print some warnings when the partitions are not aligned */ + if ((startptr != (char *) TYPEALIGN(numa_page_size, startptr)) || + (endptr != (char *) TYPEALIGN(numa_page_size, endptr))) + { + elog(WARNING, "buffers descriptors for node %d not well aligned [%p,%p]", + n, startptr, endptr); + } + + /* best effort: align the pointers, so that the mbind() works */ + startptr = (char *) TYPEALIGN_DOWN(numa_page_size, startptr); + + if (n == (nnodes - 1)) + endptr = (char *) TYPEALIGN(numa_page_size, endptr); + else + endptr = (char *) TYPEALIGN_DOWN(numa_page_size, endptr); + + /* XXX or should we use pg_numa_move_to_node? */ + if (pg_numa_bind_to_node(startptr, endptr, n) != 0) + elog(WARNING, "failed to bind shared buffer descriptors partition to node %d", n); + } +#endif } AssertCheckBufferPartitions(); } +/* + * BufferPartitionsCalculate + * Pick number of buffer partitions for the number of nodes and + * MIN_BUFFER_PARTITIONS. + * + * Picks the smallest number of partitions higher thah MIN_BUFFER_PARTITIONS, + * such that all nodes have the same number of partitions. + * + * This is best-effort with respect to size of the partitions. It's possible + * the partitions are not a perfect multiple of page size, in which case + * we set location only for the part where that is possible. The buffers on + * the "boundary" may get located up on arbitrary nodes. + * + * The extra complexity of figuring out the right "partition size" is not + * worth it, and it can lead to some partitions being much smaller. This way + * we end up with partitions of almost exactly the same size (one BLCKSZ is + * the largest difference). + * + * We expect shared buffers to be much larger than page size (at least on + * system where NUMA is a relevant feature), so the number of "not located" + * buffers should be a negligible fraction. This only affects pages between + * partitions for different nodes, so (nodes-1) pages. This is certainly + * fine with 2MB huge pages, but even with 1GB pages it should be OK (as + * such systems should have humongous amounts of memory). + * + * It also means we don't need to worry about memory page size before knowing + * if huge pages got used (which we only learn during allocation). + */ +void +BufferPartitionsCalculate(int *num_nodes, int *num_partitions, + int *num_partitions_per_node) +{ + int nnodes, + nparts, + nparts_per_node; + +#if USE_LIBNUMA + nnodes = numa_num_configured_nodes(); + nparts_per_node = 1; /* at least one partition per node */ + + while ((nparts_per_node * nnodes) < MIN_BUFFER_PARTITIONS) + nparts_per_node++; + + nparts = (nnodes * nparts_per_node); +#else + /* without NUMA, assume there's just one node */ + nnodes = 1; + nparts = MIN_BUFFER_PARTITIONS; + nparts_per_node = MIN_BUFFER_PARTITIONS; +#endif + + if (num_nodes) + *num_nodes = nnodes; + + if (num_partitions) + *num_partitions = nparts; + + if (num_partitions_per_node) + *num_partitions_per_node = nparts_per_node; +} + /* * BufferPartitionCount * Returns the number of partitions created. @@ -262,13 +456,14 @@ BufferPartitionCount(void) * The returned information is first/last buffer, number of buffers. */ void -BufferPartitionGet(int idx, int *num_buffers, +BufferPartitionGet(int idx, int *node, int *num_buffers, int *first_buffer, int *last_buffer) { if ((idx >= 0) && (idx < BufferPartitionsRegistry->npartitions)) { BufferPartition *part = &BufferPartitionsRegistry->partitions[idx]; + *node = part->numa_node; *num_buffers = part->num_buffers; *first_buffer = part->first_buffer; *last_buffer = part->last_buffer; @@ -278,3 +473,17 @@ BufferPartitionGet(int idx, int *num_buffers, elog(ERROR, "invalid partition index"); } + +void +BufferPartitionsParams(int *num_nodes, int *num_partitions, + int *num_partitions_per_node) +{ + if (num_nodes) + *num_nodes = BufferPartitionsRegistry->nnodes; + + if (num_partitions) + *num_partitions = BufferPartitionsRegistry->npartitions; + + if (num_partitions_per_node) + *num_partitions_per_node = BufferPartitionsRegistry->npartitions_per_node; +} diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index fdb5bad7910a2..53ef5239e8d3b 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -15,6 +15,15 @@ */ #include "postgres.h" +#ifdef USE_LIBNUMA +#include +#endif + +#ifdef USE_LIBNUMA +#include +#include +#endif + #include "pgstat.h" #include "port/atomics.h" #include "storage/buf_internals.h" diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat index 707a77a0f113e..d0a6c77552dea 100644 --- a/src/backend/utils/misc/guc_parameters.dat +++ b/src/backend/utils/misc/guc_parameters.dat @@ -2739,6 +2739,12 @@ max => 'INT_MAX / 2', }, +{ name => 'shared_buffers_numa', type => 'bool', context => 'PGC_POSTMASTER', group => 'RESOURCES_MEM', + short_desc => 'Locate partitions of shared buffers (and descriptors) to NUMA nodes.', + variable => 'shared_buffers_numa', + boot_val => 'false', +}, + { name => 'shared_memory_size', type => 'int', context => 'PGC_INTERNAL', group => 'PRESET_OPTIONS', short_desc => 'Shows the size of the server\'s main shared memory area (rounded up to the nearest MB).', flags => 'GUC_NOT_IN_SAMPLE | GUC_DISALLOW_IN_FILE | GUC_UNIT_MB | GUC_RUNTIME_COMPUTED', diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample index 01f98de8a9ffc..670d7e900b5a4 100644 --- a/src/backend/utils/misc/postgresql.conf.sample +++ b/src/backend/utils/misc/postgresql.conf.sample @@ -142,6 +142,7 @@ #temp_buffers = 8MB # min 800kB #max_prepared_transactions = 0 # zero disables the feature # (change requires restart) +#shared_buffers_numa = off # NUMA-aware partitioning # Caution: it is not advisable to set max_prepared_transactions nonzero unless # you actively intend to use prepared transactions. #work_mem = 4MB # min 64kB diff --git a/src/include/port/pg_numa.h b/src/include/port/pg_numa.h index 1b668fe1d9112..8fe4d4ab7e349 100644 --- a/src/include/port/pg_numa.h +++ b/src/include/port/pg_numa.h @@ -17,6 +17,13 @@ extern PGDLLIMPORT int pg_numa_init(void); extern PGDLLIMPORT int pg_numa_query_pages(int pid, unsigned long count, void **pages, int *status); extern PGDLLIMPORT int pg_numa_get_max_node(void); +extern PGDLLIMPORT Size pg_numa_page_size(void); +extern PGDLLIMPORT void pg_numa_move_to_node(char *startptr, char *endptr, int node); +extern PGDLLIMPORT int pg_numa_bind_to_node(char *startptr, char *endptr, int node); + +extern PGDLLIMPORT int numa_flags; + +#define NUMA_BUFFERS 0x01 #ifdef USE_LIBNUMA diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h index 115ce68802c01..87d7d5124dae3 100644 --- a/src/include/storage/buf_internals.h +++ b/src/include/storage/buf_internals.h @@ -365,10 +365,10 @@ typedef struct BufferDesc * line sized. * * XXX: As this is primarily matters in highly concurrent workloads which - * probably all are 64bit these days, and the space wastage would be a bit - * more noticeable on 32bit systems, we don't force the stride to be cache - * line sized on those. If somebody does actual performance testing, we can - * reevaluate. + * probably all are 64bit these days. We force the stride to be cache line + * sized even on 32bit systems, where the space wastage is be a bit more + * noticeable, to allow partitioning of shared buffers (which requires the + * memory page be a multiple of buffer descriptor). * * Note that local buffer descriptors aren't forced to be aligned - as there's * no concurrent access to those it's unlikely to be beneficial. @@ -378,7 +378,7 @@ typedef struct BufferDesc * platform with either 32 or 128 byte line sizes, it's good to align to * boundaries and avoid false sharing. */ -#define BUFFERDESC_PAD_TO_SIZE (SIZEOF_VOID_P == 8 ? 64 : 1) +#define BUFFERDESC_PAD_TO_SIZE 64 typedef union BufferDescPadded { @@ -415,8 +415,12 @@ extern PGDLLIMPORT BufferPartitions *BufferPartitionsRegistry; extern PGDLLIMPORT ConditionVariableMinimallyPadded *BufferIOCVArray; extern int BufferPartitionCount(void); -extern void BufferPartitionGet(int idx, int *num_buffers, +extern void BufferPartitionGet(int idx, int *node, int *num_buffers, int *first_buffer, int *last_buffer); +extern void BufferPartitionsCalculate(int *num_nodes, int *num_partitions, + int *num_partitions_per_node); +extern void BufferPartitionsParams(int *num_nodes, int *num_partitions, + int *num_partitions_per_node); /* in localbuf.c */ extern PGDLLIMPORT BufferDesc *LocalBufferDescriptors; diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h index 79a3f44747ad2..1cf09e8fb7c5e 100644 --- a/src/include/storage/bufmgr.h +++ b/src/include/storage/bufmgr.h @@ -158,10 +158,12 @@ typedef struct ReadBuffersOperation ReadBuffersOperation; /* * information about one partition of shared buffers * + * numa_nod specifies node for this partition (-1 means allocated on any node) * first/last buffer - the values are inclusive */ typedef struct BufferPartition { + int numa_node; /* NUMA node (-1 no node) */ int num_buffers; /* number of buffers */ int first_buffer; /* first buffer of partition */ int last_buffer; /* last buffer of partition */ @@ -170,7 +172,9 @@ typedef struct BufferPartition /* an array of information about all partitions */ typedef struct BufferPartitions { + int nnodes; /* number of NUMA nodes */ int npartitions; /* number of partitions */ + int npartitions_per_node; /* for convenience */ BufferPartition partitions[FLEXIBLE_ARRAY_MEMBER]; } BufferPartitions; @@ -206,6 +210,7 @@ extern PGDLLIMPORT const PgAioHandleCallbacks aio_local_buffer_readv_cb; /* in buf_init.c */ extern PGDLLIMPORT char *BufferBlocks; +extern PGDLLIMPORT bool shared_buffers_numa; /* in localbuf.c */ extern PGDLLIMPORT int NLocBuffer; @@ -390,6 +395,9 @@ extern void MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied, int32 *buffers_already_dirty, int32 *buffers_skipped); +/* in buf_init.c */ +extern int BufferGetNode(Buffer buffer); + /* in localbuf.c */ extern void AtProcExit_LocalBuffers(void); diff --git a/src/port/pg_numa.c b/src/port/pg_numa.c index 8954669273ae3..66985f32db33f 100644 --- a/src/port/pg_numa.c +++ b/src/port/pg_numa.c @@ -18,6 +18,9 @@ #include "miscadmin.h" #include "port/pg_numa.h" +#include "storage/pg_shmem.h" + +int numa_flags; /* * At this point we provide support only for Linux thanks to libnuma, but in @@ -118,6 +121,94 @@ pg_numa_get_max_node(void) return numa_max_node(); } +/* + * pg_numa_move_to_node + * move memory to different NUMA nodes in larger chunks + * + * startptr - start of the region (should be aligned to page size) + * endptr - end of the region (doesn't need to be aligned) + * node - node to move the memory to + * + * The "startptr" is expected to be a multiple of system memory page size, as + * determined by pg_numa_page_size. + * + * XXX We only expect to do this during startup, when the shared memory is + * still being setup. + */ +void +pg_numa_move_to_node(char *startptr, char *endptr, int node) +{ + Size sz = (endptr - startptr); + + Assert((int64) startptr % pg_numa_page_size() == 0); + + /* + * numa_tonode_memory does not actually cause a page fault, and thus does + * not locate the memory on the node. So it's fast, at least compared to + * pg_numa_query_pages, and does not make startup longer. But it also + * means the expensive part happen later, on the first access. + */ + numa_tonode_memory(startptr, sz, node); +} + +int +pg_numa_bind_to_node(char *startptr, char *endptr, int node) +{ + int ret; + struct bitmask *nodemask; + + if (node < 0) + { + errno = EINVAL; + return -1; + } + + nodemask = numa_allocate_nodemask(); + if (nodemask == NULL) + { + errno = ENOMEM; + return -1; + } + + numa_bitmask_setbit(nodemask, node); + + /* + * MPOL_BIND places the pages strictly on the node, and MPOL_MF_MOVE migrates + * pages already faulted in to that node. If mbind() fails, leave the default + * placement in effect, and report the failure. + */ + ret = mbind(startptr, (endptr - startptr), + MPOL_BIND, nodemask->maskp, nodemask->size, MPOL_MF_MOVE); + + numa_free_nodemask(nodemask); + + return ret; +} + +Size +pg_numa_page_size(void) +{ + Size os_page_size; + Size huge_page_size; + +#ifdef WIN32 + SYSTEM_INFO sysinfo; + + GetSystemInfo(&sysinfo); + os_page_size = sysinfo.dwPageSize; +#else + os_page_size = sysconf(_SC_PAGESIZE); +#endif + + /* assume huge pages get used, unless HUGE_PAGES_OFF */ + if (huge_pages_status != HUGE_PAGES_OFF) + GetHugePageSize(&huge_page_size, NULL); + else + huge_page_size = 0; + + return Max(os_page_size, huge_page_size); +} + #else /* Empty wrappers */ @@ -140,4 +231,25 @@ pg_numa_get_max_node(void) return 0; } +void +pg_numa_move_to_node(char *startptr, char *endptr, int node) +{ + /* we don't expect to ever get here in builds without libnuma */ + Assert(false); +} + +int +pg_numa_bind_to_node(char *startptr, char *endptr, int node) +{ + /* we don't expect to ever get here in builds without libnuma */ + Assert(false); +} + +Size +pg_numa_page_size(void) +{ + /* we don't expect to ever get here in builds without libnuma */ + Assert(false); +} + #endif From 204f458670438f4a0e79ed1e904b019b63dba2fc Mon Sep 17 00:00:00 2001 From: Tomas Vondra Date: Tue, 2 Jun 2026 22:16:55 +0200 Subject: [PATCH 07/17] clock-sweep: basic partitioning Partitions the "clock-sweep" algorithm to work on individual partitions, one by one. Each backend process is mapped to one "home" partition, with an independent clock hand. This reduces contention for workloads with significant buffer pressure. The patch extends the "pg_buffercache_partitions" view to include information about the clock-sweep activity. Note: This needs some sort of "balancing" when one of the partitions is much busier than the rest (e.g. because there's a single backend consuming a lot of buffers from it). Note: There's a problem with some tests running out of unpinned buffers, due to (intentionally) setting shared buffers very low. That happens because StrategyGetBuffer() only searches a single partition, and it has a couple more issues. --- .../pg_buffercache--1.7--1.8.sql | 8 +- contrib/pg_buffercache/pg_buffercache_pages.c | 32 +- src/backend/storage/buffer/bufmgr.c | 202 +++++++---- src/backend/storage/buffer/freelist.c | 333 ++++++++++++++++-- src/include/storage/buf_internals.h | 4 +- src/include/storage/bufmgr.h | 5 + src/test/recovery/t/027_stream_regress.pl | 5 + src/tools/pgindent/typedefs.list | 1 + 8 files changed, 486 insertions(+), 104 deletions(-) diff --git a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql index a6e49fd165291..92176fed7f89b 100644 --- a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql +++ b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql @@ -14,7 +14,13 @@ CREATE VIEW pg_buffercache_partitions AS numa_node integer, -- NUMA node of the partitioon num_buffers integer, -- number of buffers in the partition first_buffer integer, -- first buffer of partition - last_buffer integer); -- last buffer of partition + last_buffer integer, -- last buffer of partition + + -- clocksweep counters + num_passes bigint, -- clocksweep passes + next_buffer integer, -- next victim buffer for clocksweep + total_allocs bigint, -- handled allocs (running total) + num_allocs bigint); -- handled allocs (current cycle) -- Don't want these to be available to public. REVOKE ALL ON FUNCTION pg_buffercache_partitions() FROM PUBLIC; diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c index 46b6b85a2e3f1..b07fafda0d98b 100644 --- a/contrib/pg_buffercache/pg_buffercache_pages.c +++ b/contrib/pg_buffercache/pg_buffercache_pages.c @@ -31,7 +31,7 @@ #define NUM_BUFFERCACHE_MARK_DIRTY_ALL_ELEM 3 #define NUM_BUFFERCACHE_OS_PAGES_ELEM 3 -#define NUM_BUFFERCACHE_PARTITIONS_ELEM 5 +#define NUM_BUFFERCACHE_PARTITIONS_ELEM 9 PG_MODULE_MAGIC_EXT( .name = "pg_buffercache", @@ -963,6 +963,14 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) INT4OID, -1, 0); TupleDescInitEntry(tupledesc, (AttrNumber) 5, "last_buffer", INT4OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 6, "num_passes", + INT8OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 7, "next_buffer", + INT4OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 8, "total_allocs", + INT8OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 9, "num_allocs", + INT8OID, -1, 0); funcctx->user_fctx = BlessTupleDesc(tupledesc); @@ -984,12 +992,22 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) first_buffer, last_buffer; + uint64 buffer_total_allocs; + + uint32 complete_passes, + next_victim_buffer, + buffer_allocs; + Datum values[NUM_BUFFERCACHE_PARTITIONS_ELEM]; bool nulls[NUM_BUFFERCACHE_PARTITIONS_ELEM]; BufferPartitionGet(i, &numa_node, &num_buffers, &first_buffer, &last_buffer); + ClockSweepPartitionGetInfo(i, + &complete_passes, &next_victim_buffer, + &buffer_total_allocs, &buffer_allocs); + values[0] = Int32GetDatum(i); nulls[0] = false; @@ -1005,6 +1023,18 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) values[4] = Int32GetDatum(last_buffer); nulls[4] = false; + values[5] = Int64GetDatum(complete_passes); + nulls[5] = false; + + values[6] = Int32GetDatum(next_victim_buffer); + nulls[6] = false; + + values[7] = Int64GetDatum(buffer_total_allocs); + nulls[7] = false; + + values[8] = Int64GetDatum(buffer_allocs); + nulls[8] = false; + /* Build and return the tuple. */ tuple = heap_form_tuple((TupleDesc) funcctx->user_fctx, values, nulls); result = HeapTupleGetDatum(tuple); diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c index f81c7732e6899..f9f43e79e48b4 100644 --- a/src/backend/storage/buffer/bufmgr.c +++ b/src/backend/storage/buffer/bufmgr.c @@ -3842,33 +3842,34 @@ BufferSync(int flags) } /* - * BgBufferSync -- Write out some dirty buffers in the pool. - * - * This is called periodically by the background writer process. + * Information saved between calls so we can determine the strategy point's + * advance rate and avoid scanning already-cleaned buffers. * - * Returns true if it's appropriate for the bgwriter process to go into - * low-power hibernation mode. (This happens if the strategy clock-sweep - * has been "lapped" and no buffer allocations have occurred recently, - * or if the bgwriter has been effectively disabled by setting - * bgwriter_lru_maxpages to 0.) + * XXX Does it actually make sense to split all of this information per + * partition? For example, does per-partition advance rate mean anything? + * Maybe we should have a global advance rate? Although, if we want to + * keep enough clean buffers in each partition, maybe having per-partition + * rates makes sense. */ -bool -BgBufferSync(WritebackContext *wb_context) +typedef struct BufferSyncPartition +{ + int prev_strategy_buf_id; + uint32 prev_strategy_passes; + int next_to_clean; + uint32 next_passes; +} BufferSyncPartition; + +static BufferSyncPartition *saved_info = NULL; +static bool saved_info_valid = false; + +static bool +BgBufferSyncPartition(WritebackContext *wb_context, int num_partitions, + int partition, int recent_alloc_partition, + BufferSyncPartition *saved) { /* info obtained from freelist.c */ int strategy_buf_id; uint32 strategy_passes; - uint32 recent_alloc; - - /* - * Information saved between calls so we can determine the strategy - * point's advance rate and avoid scanning already-cleaned buffers. - */ - static bool saved_info_valid = false; - static int prev_strategy_buf_id; - static uint32 prev_strategy_passes; - static int next_to_clean; - static uint32 next_passes; /* Moving averages of allocation rate and clean-buffer density */ static float smoothed_alloc = 0; @@ -3896,25 +3897,16 @@ BgBufferSync(WritebackContext *wb_context) long new_strategy_delta; uint32 new_recent_alloc; + /* buffer range for the clocksweep partition */ + int first_buffer; + int num_buffers; + /* * Find out where the clock-sweep currently is, and how many buffer * allocations have happened since our last call. */ - strategy_buf_id = StrategySyncStart(&strategy_passes, &recent_alloc); - - /* Report buffer alloc counts to pgstat */ - PendingBgWriterStats.buf_alloc += recent_alloc; - - /* - * If we're not running the LRU scan, just stop after doing the stats - * stuff. We mark the saved state invalid so that we can recover sanely - * if LRU scan is turned back on later. - */ - if (bgwriter_lru_maxpages <= 0) - { - saved_info_valid = false; - return true; - } + strategy_buf_id = StrategySyncStart(partition, &strategy_passes, + &first_buffer, &num_buffers); /* * Compute strategy_delta = how many buffers have been scanned by the @@ -3926,17 +3918,17 @@ BgBufferSync(WritebackContext *wb_context) */ if (saved_info_valid) { - int32 passes_delta = strategy_passes - prev_strategy_passes; + int32 passes_delta = strategy_passes - saved->prev_strategy_passes; - strategy_delta = strategy_buf_id - prev_strategy_buf_id; - strategy_delta += (long) passes_delta * NBuffers; + strategy_delta = strategy_buf_id - saved->prev_strategy_buf_id; + strategy_delta += (long) passes_delta * num_buffers; Assert(strategy_delta >= 0); - if ((int32) (next_passes - strategy_passes) > 0) + if ((int32) (saved->next_passes - strategy_passes) > 0) { /* we're one pass ahead of the strategy point */ - bufs_to_lap = strategy_buf_id - next_to_clean; + bufs_to_lap = strategy_buf_id - saved->next_to_clean; #ifdef BGW_DEBUG elog(DEBUG2, "bgwriter ahead: bgw %u-%u strategy %u-%u delta=%ld lap=%d", next_passes, next_to_clean, @@ -3944,11 +3936,11 @@ BgBufferSync(WritebackContext *wb_context) strategy_delta, bufs_to_lap); #endif } - else if (next_passes == strategy_passes && - next_to_clean >= strategy_buf_id) + else if (saved->next_passes == strategy_passes && + saved->next_to_clean >= strategy_buf_id) { /* on same pass, but ahead or at least not behind */ - bufs_to_lap = NBuffers - (next_to_clean - strategy_buf_id); + bufs_to_lap = num_buffers - (saved->next_to_clean - strategy_buf_id); #ifdef BGW_DEBUG elog(DEBUG2, "bgwriter ahead: bgw %u-%u strategy %u-%u delta=%ld lap=%d", next_passes, next_to_clean, @@ -3968,9 +3960,9 @@ BgBufferSync(WritebackContext *wb_context) strategy_passes, strategy_buf_id, strategy_delta); #endif - next_to_clean = strategy_buf_id; - next_passes = strategy_passes; - bufs_to_lap = NBuffers; + saved->next_to_clean = strategy_buf_id; + saved->next_passes = strategy_passes; + bufs_to_lap = num_buffers; } } else @@ -3984,15 +3976,16 @@ BgBufferSync(WritebackContext *wb_context) strategy_passes, strategy_buf_id); #endif strategy_delta = 0; - next_to_clean = strategy_buf_id; - next_passes = strategy_passes; - bufs_to_lap = NBuffers; + saved->next_to_clean = strategy_buf_id; + saved->next_passes = strategy_passes; + bufs_to_lap = num_buffers; } /* Update saved info for next time */ - prev_strategy_buf_id = strategy_buf_id; - prev_strategy_passes = strategy_passes; - saved_info_valid = true; + saved->prev_strategy_buf_id = strategy_buf_id; + saved->prev_strategy_passes = strategy_passes; + /* XXX this needs to happen only after all partitions */ + /* saved_info_valid = true; */ /* * Compute how many buffers had to be scanned for each new allocation, ie, @@ -4000,9 +3993,9 @@ BgBufferSync(WritebackContext *wb_context) * * If the strategy point didn't move, we don't update the density estimate */ - if (strategy_delta > 0 && recent_alloc > 0) + if (strategy_delta > 0 && recent_alloc_partition > 0) { - scans_per_alloc = (float) strategy_delta / (float) recent_alloc; + scans_per_alloc = (float) strategy_delta / (float) recent_alloc_partition; smoothed_density += (scans_per_alloc - smoothed_density) / smoothing_samples; } @@ -4012,7 +4005,7 @@ BgBufferSync(WritebackContext *wb_context) * strategy point and where we've scanned ahead to, based on the smoothed * density estimate. */ - bufs_ahead = NBuffers - bufs_to_lap; + bufs_ahead = num_buffers - bufs_to_lap; reusable_buffers_est = (float) bufs_ahead / smoothed_density; /* @@ -4020,10 +4013,10 @@ BgBufferSync(WritebackContext *wb_context) * a true average we want a fast-attack, slow-decline behavior: we * immediately follow any increase. */ - if (smoothed_alloc <= (float) recent_alloc) - smoothed_alloc = recent_alloc; + if (smoothed_alloc <= (float) recent_alloc_partition) + smoothed_alloc = recent_alloc_partition; else - smoothed_alloc += ((float) recent_alloc - smoothed_alloc) / + smoothed_alloc += ((float) recent_alloc_partition - smoothed_alloc) / smoothing_samples; /* Scale the estimate by a GUC to allow more aggressive tuning. */ @@ -4050,7 +4043,7 @@ BgBufferSync(WritebackContext *wb_context) * the BGW will be called during the scan_whole_pool time; slice the * buffer pool into that many sections. */ - min_scan_buffers = (int) (NBuffers / (scan_whole_pool_milliseconds / BgWriterDelay)); + min_scan_buffers = (int) (num_buffers / (scan_whole_pool_milliseconds / BgWriterDelay)); if (upcoming_alloc_est < (min_scan_buffers + reusable_buffers_est)) { @@ -4075,20 +4068,20 @@ BgBufferSync(WritebackContext *wb_context) /* Execute the LRU scan */ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est) { - int sync_state = SyncOneBuffer(next_to_clean, true, + int sync_state = SyncOneBuffer(saved->next_to_clean, true, wb_context); - if (++next_to_clean >= NBuffers) + if (++saved->next_to_clean >= (first_buffer + num_buffers)) { - next_to_clean = 0; - next_passes++; + saved->next_to_clean = first_buffer; + saved->next_passes++; } num_to_scan--; if (sync_state & BUF_WRITTEN) { reusable_buffers++; - if (++num_written >= bgwriter_lru_maxpages) + if (++num_written >= (bgwriter_lru_maxpages / num_partitions)) { PendingBgWriterStats.maxwritten_clean++; break; @@ -4102,7 +4095,7 @@ BgBufferSync(WritebackContext *wb_context) #ifdef BGW_DEBUG elog(DEBUG1, "bgwriter: recent_alloc=%u smoothed=%.2f delta=%ld ahead=%d density=%.2f reusable_est=%d upcoming_est=%d scanned=%d wrote=%d reusable=%d", - recent_alloc, smoothed_alloc, strategy_delta, bufs_ahead, + recent_alloc_partition, smoothed_alloc, strategy_delta, bufs_ahead, smoothed_density, reusable_buffers_est, upcoming_alloc_est, bufs_to_lap - num_to_scan, num_written, @@ -4132,8 +4125,83 @@ BgBufferSync(WritebackContext *wb_context) #endif } + /* can this partition hibernate */ + return (bufs_to_lap == 0 && recent_alloc_partition == 0); +} + +/* + * BgBufferSync -- Write out some dirty buffers in the pool. + * + * This is called periodically by the background writer process. + * + * Returns true if it's appropriate for the bgwriter process to go into + * low-power hibernation mode. (This happens if the strategy clock-sweep + * has been "lapped" and no buffer allocations have occurred recently, + * or if the bgwriter has been effectively disabled by setting + * bgwriter_lru_maxpages to 0.) + */ +bool +BgBufferSync(WritebackContext *wb_context) +{ + /* info obtained from freelist.c */ + uint32 recent_alloc; + uint32 recent_alloc_partition; + int num_partitions; + + /* assume we can hibernate, any partition can set to false */ + bool hibernate = true; + + /* get the number of clocksweep partitions, and total alloc count */ + StrategySyncPrepare(&num_partitions, &recent_alloc); + + /* allocate space for per-partition information between calls */ + if (saved_info == NULL) + { + /* + * XXX Not great it's using malloc(), but how else to allocate a + * variable-length array? + */ + saved_info = malloc(sizeof(BufferSyncPartition) * num_partitions); + } + + /* Report buffer alloc counts to pgstat */ + PendingBgWriterStats.buf_alloc += recent_alloc; + + /* average alloc buffers per partition */ + recent_alloc_partition = (recent_alloc / num_partitions); + + /* + * If we're not running the LRU scan, just stop after doing the stats + * stuff. We mark the saved state invalid so that we can recover sanely + * if LRU scan is turned back on later. + */ + if (bgwriter_lru_maxpages <= 0) + { + saved_info_valid = false; + return true; + } + + /* + * now process the clocksweep partitions, one by one, using the same + * cleanup that we used for all buffers + * + * XXX Maybe we should randomize the order of partitions a bit, so that we + * don't start from partition 0 all the time? Perhaps not entirely, but at + * least pick a random starting point? + */ + for (int partition = 0; partition < num_partitions; partition++) + { + /* hibernate if all partitions can hibernate */ + hibernate &= BgBufferSyncPartition(wb_context, num_partitions, + partition, recent_alloc_partition, + &saved_info[partition]); + } + + /* now that we've scanned all partitions, mark the cached info as valid */ + saved_info_valid = true; + /* Return true if OK to hibernate */ - return (bufs_to_lap == 0 && recent_alloc == 0); + return hibernate; } /* diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index 53ef5239e8d3b..2d56579682e6d 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -36,17 +36,28 @@ /* - * The shared freelist control information. + * Information about one partition of the ClockSweep (on a subset of buffers). + * + * XXX Should be careful to align this to cachelines, etc. */ typedef struct { /* Spinlock: protects the values below */ - slock_t buffer_strategy_lock; + slock_t clock_sweep_lock; + + /* range for this clock sweep partition */ + int32 node; + int32 firstBuffer; + int32 numBuffers; /* * clock-sweep hand: index of next buffer to consider grabbing. Note that * this isn't a concrete buffer - we only ever increase the value. So, to * get an actual buffer, it needs to be used modulo NBuffers. + * + * XXX This is relative to firstBuffer, so needs to be offset properly. + * + * XXX firstBuffer + (nextVictimBuffer % numBuffers) */ pg_atomic_uint32 nextVictimBuffer; @@ -57,11 +68,32 @@ typedef struct uint32 completePasses; /* Complete cycles of the clock-sweep */ pg_atomic_uint32 numBufferAllocs; /* Buffers allocated since last reset */ + /* running total of allocs */ + pg_atomic_uint64 numTotalAllocs; + +} ClockSweep; + +/* + * The shared freelist control information. + */ +typedef struct +{ + /* Spinlock: protects the values below */ + slock_t buffer_strategy_lock; + /* * Bgworker process to be notified upon activity or -1 if none. See * StrategyNotifyBgWriter. */ int bgwprocno; + + /* cached info about freelist partitioning */ + int num_nodes; + int num_partitions; + int num_partitions_per_node; + + /* clocksweep partitions */ + ClockSweep sweeps[FLEXIBLE_ARRAY_MEMBER]; } BufferStrategyControl; /* Pointers to shared state */ @@ -108,6 +140,7 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state); static void AddBufferToRing(BufferAccessStrategy strategy, BufferDesc *buf); +static ClockSweep *ChooseClockSweep(void); /* * ClockSweepTick - Helper routine for StrategyGetBuffer() @@ -119,6 +152,7 @@ static inline uint32 ClockSweepTick(void) { uint32 victim; + ClockSweep *sweep = ChooseClockSweep(); /* * Atomically move hand ahead one buffer - if there's several processes @@ -126,14 +160,14 @@ ClockSweepTick(void) * apparent order. */ victim = - pg_atomic_fetch_add_u32(&StrategyControl->nextVictimBuffer, 1); + pg_atomic_fetch_add_u32(&sweep->nextVictimBuffer, 1); - if (victim >= NBuffers) + if (victim >= sweep->numBuffers) { uint32 originalVictim = victim; /* always wrap what we look up in BufferDescriptors */ - victim = victim % NBuffers; + victim = victim % sweep->numBuffers; /* * If we're the one that just caused a wraparound, force @@ -159,19 +193,118 @@ ClockSweepTick(void) * could lead to an overflow of nextVictimBuffers, but that's * highly unlikely and wouldn't be particularly harmful. */ - SpinLockAcquire(&StrategyControl->buffer_strategy_lock); + SpinLockAcquire(&sweep->clock_sweep_lock); - wrapped = expected % NBuffers; + wrapped = expected % sweep->numBuffers; - success = pg_atomic_compare_exchange_u32(&StrategyControl->nextVictimBuffer, + success = pg_atomic_compare_exchange_u32(&sweep->nextVictimBuffer, &expected, wrapped); if (success) - StrategyControl->completePasses++; - SpinLockRelease(&StrategyControl->buffer_strategy_lock); + sweep->completePasses++; + SpinLockRelease(&sweep->clock_sweep_lock); } } } - return victim; + + /* + * Make sure we've calculated a buffer in the range of the partition. Buffer + * IDs are 1-based, we're calculating 0-based indexes. + */ + Assert((victim >= 0) && (victim < sweep->numBuffers)); + Assert(BufferIsValid(1 + sweep->firstBuffer + victim)); + + return sweep->firstBuffer + victim; +} + +/* + * ClockSweepPartitionIndex + * pick the clock-sweep partition to use based on PID and NUMA node + * + * With libnuma, use the NUMA node and PID to pick the partition. Otherwise + * use just PID (as if there's a single NUMA node). + * + * XXX This should also check if buffers are NUMA-partitioned, not just if + * compiled with libnuma. + */ +static int +ClockSweepPartitionIndex(void) +{ + int node = 0, + index; + pid_t pid = MyProcPid;; + + Assert(StrategyControl->num_partitions == + (StrategyControl->num_nodes * StrategyControl->num_partitions_per_node)); + + /* + * If buffers are NUMA-partitioned, determine the partition using the NUMA + * node and PID. Without NUMA assume everything is a single NUMA node 0, and + * we pick the partition based on PID. + */ +#ifdef USE_LIBNUMA + if (shared_buffers_numa) + { + int cpu; + + /* XXX do we need to check sched_getcpu is available, somehow? */ + if ((cpu = sched_getcpu()) < 0) + elog(ERROR, "sched_getcpu failed: %m"); + + node = numa_node_of_cpu(cpu); + } +#endif + + /* + * We should't get unexpected NUMA nodes, not considered when setting up the + * buffer partitions. It could happen if the allowed NUMA nodes get adjusted + * at runtime, but at this point we just create partitions for all existing + * nodes. We could plan for allowed partitions, but then what if those get + * disabled, and the user allows some other partitions? + */ + if ((node < 0) || (node > StrategyControl->num_nodes)) + elog(ERROR, "node out of range: %d > %u", node, StrategyControl->num_nodes); + + /* + * Calculate the partition index. Nodes have the same number of partitions, + * and we use the PID to pick one of those (for a given node). If there's + * only a single partition per node, we can ignore PID and use node directly. + */ + if (StrategyControl->num_partitions_per_node == 1) + { + /* fast-path */ + index = node; + } + else + { + /* use PID to pick one of node's partitions */ + index = (node * StrategyControl->num_partitions_per_node) + + (pid % StrategyControl->num_partitions_per_node); + } + + /* should have a valid partition index */ + Assert((index >= 0) && (index < StrategyControl->num_partitions)); + + return index; +} + +/* + * ChooseClockSweep + * pick a clocksweep partition based on NUMA node and PID + * + * Pick a partition mapped to the NUMA node the backend is currently running + * on, and use PID if there are multiple partitions per node. Without NUMA + * supported/enabled, use just PID. + * + * XXX Maybe we should do both the total and "per group" counts a power of + * two? That'd allow using shifts instead of divisions in the calculation, + * and that's cheaper. But how would that deal with odd number of nodes? + */ +static ClockSweep * +ChooseClockSweep(void) +{ + int index = ClockSweepPartitionIndex(); + + return &StrategyControl->sweeps[index]; } /* @@ -242,10 +375,37 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r * We count buffer allocation requests so that the bgwriter can estimate * the rate of buffer consumption. Note that buffers recycled by a * strategy object are intentionally not counted here. + * + * XXX It's not quite right we call ChooseClockSweep twice - now, and then + * a couple lines later (through ClockSweepTick). If the process moves + * between CPUs / NUMA nodes in between, these call may pick different + * partitions, confusing the logic a bit. */ - pg_atomic_fetch_add_u32(&StrategyControl->numBufferAllocs, 1); + pg_atomic_fetch_add_u32(&ChooseClockSweep()->numBufferAllocs, 1); - /* Use the "clock sweep" algorithm to find a free buffer */ + /* + * Use the "clock sweep" algorithm to find a free buffer + * + * XXX Note that ClockSweepTick() is NUMA-aware, i.e. it only looks at + * buffers from a single partition, aligned with the NUMA node. That means + * a process "sweeps" only a fraction of buffers, even if the other buffers + * are better candidates for eviction. Maybe there should be some logic to + * "steal" buffers from other partitions or other nodes? + * + * XXX This only searches a single partition, which can result in "no + * unpinned buffers available" even if there are buffers in other + * partitions. Needs to scan partitions if needed, as a fallback. + * + * XXX Would that also mean we should have multiple bgwriters, one for each + * node, or would one bgwriter still handle all nodes? + * + * XXX Also, the trycounter should not be set to NBuffers, but to buffer + * count for that one partition. In fact, this should not call ClockSweepTick + * for every iteration. The call is likely quite expensive (does a lot + * of stuff), and also may return a different partition on each call. + * We should just do it once, and then do the for(;;) loop. And then + * maybe advance to the next partition, until we scan through all of them. + */ trycounter = NBuffers; for (;;) { @@ -325,6 +485,48 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r } } +/* + * StrategySyncPrepare -- prepare for sync of all partitions + * + * Determine the number of clocksweep partitions, and calculate the recent + * buffers allocs (as a sum of all the partitions). This allows BgBufferSync + * to calculate average number of allocations per partition for the next + * sync cycle. + * + * In addition it returns the count of recent buffer allocs, which is a total + * summed from all partitions. The alloc counts are reset after being read, + * as the partitions are walked. + */ +void +StrategySyncPrepare(int *num_parts, uint32 *num_buf_alloc) +{ + *num_buf_alloc = 0; + *num_parts = StrategyControl->num_partitions; + + /* + * We lock the partitions one by one, so not exacly in sync, but that + * should be fine. We're only looking for heuristics anyway. + */ + for (int i = 0; i < StrategyControl->num_partitions; i++) + { + ClockSweep *sweep = &StrategyControl->sweeps[i]; + + /* XXX Do we need the lock, if we're only accessing atomics? Surely not. */ + /* XXX Are we ever calling this without num_buf_alloc? */ + SpinLockAcquire(&sweep->clock_sweep_lock); + if (num_buf_alloc) + { + uint32 allocs = pg_atomic_exchange_u32(&sweep->numBufferAllocs, 0); + + /* include the count in the running total */ + pg_atomic_fetch_add_u64(&sweep->numTotalAllocs, allocs); + + *num_buf_alloc += allocs; + } + SpinLockRelease(&sweep->clock_sweep_lock); + } +} + /* * StrategySyncStart -- tell BgBufferSync where to start syncing * @@ -332,37 +534,44 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r * BgBufferSync() will proceed circularly around the buffer array from there. * * In addition, we return the completed-pass count (which is effectively - * the higher-order bits of nextVictimBuffer) and the count of recent buffer - * allocs if non-NULL pointers are passed. The alloc count is reset after - * being read. + * the higher-order bits of nextVictimBuffer). + * + * This only considers a single clocksweep partition, as BgBufferSync looks + * at them one by one. */ int -StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc) +StrategySyncStart(int partition, uint32 *complete_passes, + int *first_buffer, int *num_buffers) { uint32 nextVictimBuffer; int result; + ClockSweep *sweep = &StrategyControl->sweeps[partition]; - SpinLockAcquire(&StrategyControl->buffer_strategy_lock); - nextVictimBuffer = pg_atomic_read_u32(&StrategyControl->nextVictimBuffer); - result = nextVictimBuffer % NBuffers; + Assert((partition >= 0) && (partition < StrategyControl->num_partitions)); + + SpinLockAcquire(&sweep->clock_sweep_lock); + nextVictimBuffer = pg_atomic_read_u32(&sweep->nextVictimBuffer); + result = nextVictimBuffer % sweep->numBuffers; + + *first_buffer = sweep->firstBuffer; + *num_buffers = sweep->numBuffers; if (complete_passes) { - *complete_passes = StrategyControl->completePasses; + *complete_passes = sweep->completePasses; /* * Additionally add the number of wraparounds that happened before * completePasses could be incremented. C.f. ClockSweepTick(). */ - *complete_passes += nextVictimBuffer / NBuffers; + *complete_passes += nextVictimBuffer / sweep->numBuffers; } + SpinLockRelease(&sweep->clock_sweep_lock); - if (num_buf_alloc) - { - *num_buf_alloc = pg_atomic_exchange_u32(&StrategyControl->numBufferAllocs, 0); - } - SpinLockRelease(&StrategyControl->buffer_strategy_lock); - return result; + /* XXX buffer IDs start at 1, we're calculating 0-based indexes */ + Assert(BufferIsValid(1 + sweep->firstBuffer + result)); + + return sweep->firstBuffer + result; } /* @@ -394,8 +603,14 @@ StrategyNotifyBgWriter(int bgwprocno) static void StrategyCtlShmemRequest(void *arg) { + int num_partitions; + + /* get the number of buffer partitions */ + BufferPartitionsCalculate(NULL, &num_partitions, NULL); + ShmemRequestStruct(.name = "Buffer Strategy Status", - .size = sizeof(BufferStrategyControl), + .size = offsetof(BufferStrategyControl, sweeps) + + mul_size(num_partitions, sizeof(ClockSweep)), .ptr = (void **) &StrategyControl ); } @@ -408,12 +623,42 @@ StrategyCtlShmemInit(void *arg) { SpinLockInit(&StrategyControl->buffer_strategy_lock); - /* Initialize the clock-sweep pointer */ - pg_atomic_init_u32(&StrategyControl->nextVictimBuffer, 0); + /* Remember the number of partitions */ + BufferPartitionsParams(&StrategyControl->num_nodes, + &StrategyControl->num_partitions, + &StrategyControl->num_partitions_per_node); + + /* Initialize the clock sweep pointers (for all partitions) */ + for (int i = 0; i < StrategyControl->num_partitions; i++) + { + int node, + num_buffers, + first_buffer, + last_buffer; + + SpinLockInit(&StrategyControl->sweeps[i].clock_sweep_lock); - /* Clear statistics */ - StrategyControl->completePasses = 0; - pg_atomic_init_u32(&StrategyControl->numBufferAllocs, 0); + pg_atomic_init_u32(&StrategyControl->sweeps[i].nextVictimBuffer, 0); + + /* get info about the buffer partition */ + BufferPartitionGet(i, &node, &num_buffers, + &first_buffer, &last_buffer); + + /* + * FIXME This may not quite right, because if NBuffers is not a + * perfect multiple of numBuffers, the last partition will have + * numBuffers set too high. buf_init handles this by tracking the + * remaining number of buffers, and not overflowing. + */ + StrategyControl->sweeps[i].node = node; + StrategyControl->sweeps[i].numBuffers = num_buffers; + StrategyControl->sweeps[i].firstBuffer = first_buffer; + + /* Clear statistics */ + StrategyControl->sweeps[i].completePasses = 0; + pg_atomic_init_u32(&StrategyControl->sweeps[i].numBufferAllocs, 0); + pg_atomic_init_u64(&StrategyControl->sweeps[i].numTotalAllocs, 0); + } /* No pending notification */ StrategyControl->bgwprocno = -1; @@ -777,3 +1022,23 @@ StrategyRejectBuffer(BufferAccessStrategy strategy, BufferDesc *buf, bool from_r return true; } + +void +ClockSweepPartitionGetInfo(int idx, + uint32 *complete_passes, uint32 *next_victim_buffer, + uint64 *buffer_total_allocs, uint32 *buffer_allocs) +{ + ClockSweep *sweep = &StrategyControl->sweeps[idx]; + + Assert((idx >= 0) && (idx < StrategyControl->num_partitions)); + + /* get the clocksweep stats */ + *complete_passes = sweep->completePasses; + *next_victim_buffer = pg_atomic_read_u32(&sweep->nextVictimBuffer); + + *buffer_allocs = pg_atomic_read_u32(&sweep->numBufferAllocs); + *buffer_total_allocs = pg_atomic_read_u64(&sweep->numTotalAllocs); + + /* calculate the actual buffer ID */ + *next_victim_buffer = sweep->firstBuffer + (*next_victim_buffer % sweep->numBuffers); +} diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h index 87d7d5124dae3..5b0cf6db2712f 100644 --- a/src/include/storage/buf_internals.h +++ b/src/include/storage/buf_internals.h @@ -599,7 +599,9 @@ extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy, extern bool StrategyRejectBuffer(BufferAccessStrategy strategy, BufferDesc *buf, bool from_ring); -extern int StrategySyncStart(uint32 *complete_passes, uint32 *num_buf_alloc); +extern void StrategySyncPrepare(int *num_parts, uint32 *num_buf_alloc); +extern int StrategySyncStart(int partition, uint32 *complete_passes, + int *first_buffer, int *num_buffers); extern void StrategyNotifyBgWriter(int bgwprocno); /* buf_table.c */ diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h index 1cf09e8fb7c5e..e0bb4cc1df1d4 100644 --- a/src/include/storage/bufmgr.h +++ b/src/include/storage/bufmgr.h @@ -410,6 +410,11 @@ extern int GetAccessStrategyBufferCount(BufferAccessStrategy strategy); extern int GetAccessStrategyPinLimit(BufferAccessStrategy strategy); extern void FreeAccessStrategy(BufferAccessStrategy strategy); +extern void ClockSweepPartitionGetInfo(int idx, + uint32 *complete_passes, + uint32 *next_victim_buffer, + uint64 *buffer_total_allocs, + uint32 *buffer_allocs); /* inline functions */ diff --git a/src/test/recovery/t/027_stream_regress.pl b/src/test/recovery/t/027_stream_regress.pl index ae97729784943..f68e08e569751 100644 --- a/src/test/recovery/t/027_stream_regress.pl +++ b/src/test/recovery/t/027_stream_regress.pl @@ -18,6 +18,11 @@ $node_primary->append_conf('postgresql.conf', 'max_prepared_transactions = 10'); +# The default is 1MB, which is not enough with clock-sweep partitioning. +# Increase to 32MB, so that we don't get "no unpinned buffers". +$node_primary->append_conf('postgresql.conf', + 'shared_buffers = 32MB'); + # Enable pg_stat_statements to force tests to do query jumbling. # pg_stat_statements.max should be large enough to hold all the entries # of the regression database. diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 2ad38118af8a4..fe9517815fae6 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -447,6 +447,7 @@ ClientCertName ClientConnectionInfo ClientData ClientSocket +ClockSweep ClonePtrType ClosePortalStmt ClosePtrType From f53d61089daaaea5e0a20ad5c991e6a9a97c5b5a Mon Sep 17 00:00:00 2001 From: Tomas Vondra Date: Tue, 2 Jun 2026 22:28:10 +0200 Subject: [PATCH 08/17] clock-sweep: balancing of allocations If backends only allocate buffers from the "home" partition, that may cause significant misbalance. Some partitions might be overused, while other partitions would be left unused. In other words, shared buffers would not be used efficiently. We want all partitions to be used about the same, i.e. serve about the same number of allocations. To achieve that, allocations from partitions that are "too busy" may get redirected to other partitions. The system counts allocations requested from each partition, calculates the "fair share" (average per partition), and then redirectsexcess allocations to other partitions. Each partition gets a set of coefficients determining the fraction of allocations to redirect to other partitions. The coefficients may be interpreted as a "budget" for each of the partition, i.e. the number of allocations to serve from that partition, before moving to the next partition (in a round-robin manner). All of this is tied to the partition where the allocation was requested. Each partition has a separate set of coefficients. We might also treat the coefficients as probabilities, and use PRNG to determine where to direct individual requests. But a PRNG seems fairly expensive, and the budget approach works well. We intentionally keep the "budget" fairly low, with the sum for a given partition 100. That means we get to the same partition after only 100 allocations, keeping it more balanced. It wouldn't be hard to make the budgets higher (e.g. matching the number of allocations per round), but it might also make the behavior less smooth (long period of allocations from each partition). This is very simple/cheap, and over many allocations it has the same effect. For periods of low activity it may diverge, but that does not matter much (we care about high-activity periods much more). --- .../pg_buffercache--1.7--1.8.sql | 5 +- contrib/pg_buffercache/pg_buffercache_pages.c | 44 +- src/backend/storage/buffer/bufmgr.c | 3 + src/backend/storage/buffer/freelist.c | 428 +++++++++++++++++- src/include/storage/buf_internals.h | 1 + src/include/storage/bufmgr.h | 12 +- 6 files changed, 472 insertions(+), 21 deletions(-) diff --git a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql index 92176fed7f89b..43d2e84f9d279 100644 --- a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql +++ b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql @@ -20,7 +20,10 @@ CREATE VIEW pg_buffercache_partitions AS num_passes bigint, -- clocksweep passes next_buffer integer, -- next victim buffer for clocksweep total_allocs bigint, -- handled allocs (running total) - num_allocs bigint); -- handled allocs (current cycle) + num_allocs bigint, -- handled allocs (current cycle) + total_req_allocs bigint, -- requested allocs (running total) + num_req_allocs bigint, -- handled allocs (current cycle) + weights int[]); -- balancing weights -- Don't want these to be available to public. REVOKE ALL ON FUNCTION pg_buffercache_partitions() FROM PUBLIC; diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c index b07fafda0d98b..c3cd24af097e9 100644 --- a/contrib/pg_buffercache/pg_buffercache_pages.c +++ b/contrib/pg_buffercache/pg_buffercache_pages.c @@ -15,8 +15,11 @@ #include "port/pg_numa.h" #include "storage/buf_internals.h" #include "storage/bufmgr.h" +#include "utils/array.h" +#include "utils/builtins.h" #include "utils/rel.h" #include "utils/tuplestore.h" +#include "utils/typcache.h" #define NUM_BUFFERCACHE_PAGES_MIN_ELEM 8 @@ -31,7 +34,7 @@ #define NUM_BUFFERCACHE_MARK_DIRTY_ALL_ELEM 3 #define NUM_BUFFERCACHE_OS_PAGES_ELEM 3 -#define NUM_BUFFERCACHE_PARTITIONS_ELEM 9 +#define NUM_BUFFERCACHE_PARTITIONS_ELEM 12 PG_MODULE_MAGIC_EXT( .name = "pg_buffercache", @@ -940,6 +943,8 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) if (SRF_IS_FIRSTCALL()) { + TypeCacheEntry *typentry = lookup_type_cache(INT4OID, 0); + funcctx = SRF_FIRSTCALL_INIT(); /* Switch context when allocating stuff to be used in later calls */ @@ -971,6 +976,12 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) INT8OID, -1, 0); TupleDescInitEntry(tupledesc, (AttrNumber) 9, "num_allocs", INT8OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 10, "total_req_allocs", + INT8OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 11, "num_req_allocs", + INT8OID, -1, 0); + TupleDescInitEntry(tupledesc, (AttrNumber) 12, "weigths", + typentry->typarray, -1, 0); funcctx->user_fctx = BlessTupleDesc(tupledesc); @@ -992,11 +1003,17 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) first_buffer, last_buffer; - uint64 buffer_total_allocs; + uint64 buffer_total_allocs, + buffer_total_req_allocs; uint32 complete_passes, next_victim_buffer, - buffer_allocs; + buffer_allocs, + buffer_req_allocs; + + int *weights; + Datum *dweights; + ArrayType *array; Datum values[NUM_BUFFERCACHE_PARTITIONS_ELEM]; bool nulls[NUM_BUFFERCACHE_PARTITIONS_ELEM]; @@ -1005,8 +1022,16 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) &first_buffer, &last_buffer); ClockSweepPartitionGetInfo(i, - &complete_passes, &next_victim_buffer, - &buffer_total_allocs, &buffer_allocs); + &complete_passes, &next_victim_buffer, + &buffer_total_allocs, &buffer_allocs, + &buffer_total_req_allocs, &buffer_req_allocs, + &weights); + + dweights = palloc_array(Datum, funcctx->max_calls); + for (int i = 0; i < funcctx->max_calls; i++) + dweights[i] = Int32GetDatum(weights[i]); + + array = construct_array_builtin(dweights, funcctx->max_calls, INT4OID); values[0] = Int32GetDatum(i); nulls[0] = false; @@ -1035,6 +1060,15 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) values[8] = Int64GetDatum(buffer_allocs); nulls[8] = false; + values[9] = Int64GetDatum(buffer_total_req_allocs); + nulls[9] = false; + + values[10] = Int64GetDatum(buffer_req_allocs); + nulls[10] = false; + + values[11] = PointerGetDatum(array); + nulls[11] = false; + /* Build and return the tuple. */ tuple = heap_form_tuple((TupleDesc) funcctx->user_fctx, values, nulls); result = HeapTupleGetDatum(tuple); diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c index f9f43e79e48b4..d58397f098099 100644 --- a/src/backend/storage/buffer/bufmgr.c +++ b/src/backend/storage/buffer/bufmgr.c @@ -4151,6 +4151,9 @@ BgBufferSync(WritebackContext *wb_context) /* assume we can hibernate, any partition can set to false */ bool hibernate = true; + /* trigger partition rebalancing first */ + StrategySyncBalance(); + /* get the number of clocksweep partitions, and total alloc count */ StrategySyncPrepare(&num_partitions, &recent_alloc); diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index 2d56579682e6d..a543fb12b2178 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -35,6 +35,26 @@ #define INT_ACCESS_ONCE(var) ((int)(*((volatile int *)&(var)))) +/* + * XXX We need to make ClockSweep fixed-size, so that we can have an array + * in shared memory. The easiest way is to pick a sufficiently high value + * that no system will actually need. 32 seems high enough. + * + * XXX We should enforce this in bufmgr.c, when initializing the partitions. + */ +#define MAX_BUFFER_PARTITIONS 32 + +/* + * Coefficient used to combine the old and new balance coefficients, using + * weighted average, so that we don't flap too much. The higher the value, the + * more the old value affects the result. + * + * XXX Doesn't this obscure the interpretation of weights as probabilities to + * allocate from a given partition? Does it still sum to 100%? I don't think + * so, it's just a fraction of allocations to go from a given partition. + */ +#define CLOCKSWEEP_HISTORY_COEFF 0.5 + /* * Information about one partition of the ClockSweep (on a subset of buffers). * @@ -68,9 +88,32 @@ typedef struct uint32 completePasses; /* Complete cycles of the clock-sweep */ pg_atomic_uint32 numBufferAllocs; /* Buffers allocated since last reset */ + /* + * Buffers that should have been allocated in this partition (but might + * have been redirected to keep allocations balanced). + */ + pg_atomic_uint32 numRequestedAllocs; + /* running total of allocs */ pg_atomic_uint64 numTotalAllocs; + pg_atomic_uint64 numTotalRequestedAllocs; + /* + * Weights to balance buffer allocations for all the partitions. Each + * partition gets a vector of weights 0-100, determining what fraction + * of buffers to allocate from that partition. So [75, 15, 5, 5] would + * mean 75% allocations should go from partition 0, 15% from partition + * 1, and 5% from partitions 2&3. Each partition gets a different vector + * of weights. + * + * Backends use the budget from it's "home" partition, so that a busy + * partitions (with a lot of processes on that NUMA node etc.) spread + * the allocations evenly. + * + * XXX Allocate a fixed-length array, to simplify working with array of + * the structs, etc. + */ + uint8 balance[MAX_BUFFER_PARTITIONS]; } ClockSweep; /* @@ -140,7 +183,66 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state); static void AddBufferToRing(BufferAccessStrategy strategy, BufferDesc *buf); -static ClockSweep *ChooseClockSweep(void); +static ClockSweep *ChooseClockSweep(bool balance); + +/* + * clocksweep allocation balancing + * + * To balance allocations from clocksweep partitions, each partition gets a + * budget for allocating buffers from other partitions. A process that + * "exhausts" a budget in it's home partition gets redirected to the other + * partitions, driven by the budgets. + * + * For example, a partition may have budget [25, 25, 25, 25], which means + * each of the 4 partitions should get 1/4 of allocations. Or the buget + * can be [50, 50, 0, 0], which means all allocations will go to the first + * two partitions (one of them being the "home" one); + * + * We could do that based on a random number generator, but for now we + * simply treat the values as a budget, i.e. a number of allocations to + * serve from other partitions, and move in round-robin way. + * + * This is very simple/cheap, and over many allocations it has the same + * effect. For periods of low activity it may diverge, but that does not + * matter much (we care about high-activity periods much more). + * + * We intentionally keep the "budget" fairly low, with the sum for a given + * partition 100. That means we get to the same partition after only 100 + * allocations, keeping it more balanced. We can make the budgets higher + * (say, to match the expected number of allocations, i.e. bout the average + * number of allocations from the past interval). Or maybe configurable. + * + * XXX We should always start allocating from the "home" partition, i.e. + * from from it, and only then redirect to other partitions. + * + * XXX It probably is not great all the processes from that "home" + * partition are coordinated, and move to between partitions at about the + * same time. Not sure what to do about this. + * + * XXX We should also prefer other partitions from the same NUMA node (if + * there are some). Probably by setting the budgets. + * + * FIXME Explain at which point are the budgets recalculated, by which + * process, and how that affects other processes allocating buffers. + */ + +/* + * The "optimal" clock-sweep partition. After a backend gets moved to a + * different NUMA node, we restart the balancing so that it uses the + * correct "budget" from the new home partition. + */ +static int clocksweep_partition_home = -1; + +/* + * The partition the backend is currently allocating from (either the + * home one, or one of the redirected ones). + */ +static int clocksweep_partition_current = -1; + +/* + * The number of buffers to allocate from the current partition. + */ +static int clocksweep_partition_budget = 0; /* * ClockSweepTick - Helper routine for StrategyGetBuffer() @@ -152,7 +254,7 @@ static inline uint32 ClockSweepTick(void) { uint32 victim; - ClockSweep *sweep = ChooseClockSweep(); + ClockSweep *sweep = ChooseClockSweep(true); /* * Atomically move hand ahead one buffer - if there's several processes @@ -300,11 +402,68 @@ ClockSweepPartitionIndex(void) * and that's cheaper. But how would that deal with odd number of nodes? */ static ClockSweep * -ChooseClockSweep(void) +ChooseClockSweep(bool balance) { + /* What's the "optimal" partition for this backend? */ int index = ClockSweepPartitionIndex(); + ClockSweep *sweep = &StrategyControl->sweeps[index]; + + /* + * Was the process migrated to a different NUMA node? If the home partition + * changed, we need to reset the budget and start over, so that we correctly + * prefer "nearby" partitions etc. + * + * XXX Could this be a problem when processes move all the time? I don't + * think so - if a process moves between many partitions, that alone will + * spread the allocations over partitions. Similarly, if there are many + * processes, that should make it even more even. + */ + if (clocksweep_partition_home != index) + { + clocksweep_partition_home = index; + clocksweep_partition_current = index; + clocksweep_partition_budget = sweep->balance[index]; + } + + /* we should have a valid partition */ + Assert(clocksweep_partition_home != -1); + Assert(clocksweep_partition_current != -1); + Assert(clocksweep_partition_budget >= 0); + + /* + * When balancing allocations, redirect the allocations to other partitions + * according to the budgets. We move through partitions in a round-robin way, + * after allocating the "budget" of allocations from the current one. + */ + if (balance) + { + /* + * Ran out of budget from the current partition? Move to the next one + * with non-zero budget. + */ + while (clocksweep_partition_budget == 0) + { + /* wrap around at the end */ + clocksweep_partition_current++; + if (clocksweep_partition_current >= StrategyControl->num_partitions) + clocksweep_partition_current = 0; + + clocksweep_partition_budget + = sweep->balance[clocksweep_partition_current]; + } - return &StrategyControl->sweeps[index]; + /* account for the current allocation */ + --clocksweep_partition_budget; + + /* + * Account for the allocation in the "home" partition, so that the next + * round of rebalancing (recalculating the budgets) knows about the + * allocation traffic in various partitions. + */ + pg_atomic_fetch_add_u32(&sweep->numRequestedAllocs, 1); + } + + return &StrategyControl->sweeps[clocksweep_partition_current]; } /* @@ -381,7 +540,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r * between CPUs / NUMA nodes in between, these call may pick different * partitions, confusing the logic a bit. */ - pg_atomic_fetch_add_u32(&ChooseClockSweep()->numBufferAllocs, 1); + pg_atomic_fetch_add_u32(&ChooseClockSweep(false)->numBufferAllocs, 1); /* * Use the "clock sweep" algorithm to find a free buffer @@ -485,6 +644,229 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r } } +/* + * StrategySyncBalance + * update partition budgets, to balance the buffer allocations + * + * We want to give preference to allocating buffers on the same NUMA node, + * but that might lead to imbalance - a single process would only use a + * fraction of shared buffers. We don't want that, we want to utilize the + * whole shared buffers. The number of allocations in each partition may + * also change over time, so we need to adapt to that. + * + * To allow this "adaptive balancing", each partition has a set of weights, + * determining what fraction of allocations to direct to other partitions. + * For simplicity the coefficients are integers 0-100, expressing the + * percentage of allocations redirected to that partition. + * + * Consider for example weights [50, 25, 25, 0] for one of 4 partitions. + * This means 50% of allocations will be redirected to partition 0, 25% + * to partitions 1 and 2, and no allocations will go to partition 3. + * + * This means an allocation may be requested in partition A (i.e. the + * home partition of the process requesting it), but end up allocating + * the buffer in partition B. We have a counter for both - the number of + * allocations requested in a partition, and the number of allocations + * actually handled by that partition. The former is used for calculating + * weights, the latter is used only for monitoring. + * + * The balancing happens in intervals - it adjusts future allocations + * based on stats about recent allocations, namely: + * + * - numBufferAllocs - number of allocations served by a partition + * + * - numRequestedAllocs - number of allocatios requested in a partition + * + * We're trying to smooth numBufferAllocs in the next interval, based on + * numRequestedAllocs measured in the last interval. + * + * The balancing algorithm works like this: + * + * - the target (average number of allocations per partition) is calculated + * from total number of allocations requested in the last intervaal + * + * - partitions get divided into two groups - those with more allocation + * requests than the target, and those with fewer requests + * + * - we "distribute" the delta (which is the same between the groups) + * between the groups (one has more, the other fewer) + * + * Partitions with (nallocs > avg_nallocs) redirect the extra allocations, + * with each target allocation getting a proportional part (with respect + * to the total delta). + * + * XXX Currently this does not give preference to other partitions on the + * same NUMA node (redirect to it first), but it could. + */ +void +StrategySyncBalance(void) +{ + /* snapshot of allocation requests for partitions */ + uint32 allocs[MAX_BUFFER_PARTITIONS]; + + uint32 total_allocs = 0, /* total number of allocations */ + avg_allocs, /* average allocations (per partition) */ + delta_allocs = 0; /* sum of allocs above average */ + + /* + * Collect the number of allocations requested in the past interval. + * While at it, reset the counter to start the new interval. + * + * XXX We lock the partitions one by one, so this is not a perfectly + * consistent snapshot of the counts, and the resets happen before we + * update the weights too. But we're only looking for heuristics, so + * this should be good enough. + * + * XXX A similar issue applies to the counter reset later - we haven't + * updated the weights yet, so some of the requests counted for the next + * interval will be redirected per current weights. Should be fine, it's + * just an approximate heuristics, and there should be very few requests in + * between. Alternatively, we could reset the request counters when setting + * the new weights, and just ignore the couple requests in between. + * + * XXX Does this need to worry about the completePasses too? + */ + for (int i = 0; i < StrategyControl->num_partitions; i++) + { + ClockSweep *sweep = &StrategyControl->sweeps[i]; + + /* no need for a spinlock */ + allocs[i] = pg_atomic_exchange_u32(&sweep->numRequestedAllocs, 0); + + /* add the allocs to running total */ + pg_atomic_fetch_add_u64(&sweep->numTotalRequestedAllocs, allocs[i]); + + total_allocs += allocs[i]; + } + + /* Calculate the "fair share" of allocations per partition. */ + avg_allocs = (total_allocs / StrategyControl->num_partitions); + + /* + * Calculate the "delta" from balanced state for each partition, i.e. how + * many more/fewer allocations it handled relative to the average. + */ + for (int i = 0; i < StrategyControl->num_partitions; i++) + { + if (allocs[i] > avg_allocs) + delta_allocs += (allocs[i] - avg_allocs); + } + + /* + * Skip rebalancing when there's not enough activity, and just keep the + * current weights. + * + * XXX The threshold of 100 allocation is pretty arbitrary. + * + * XXX Maybe a better strategy would be to slowly return to the default + * weights, with each partition allocation only from itself? + * + * XXX Maybe we shouldn't even reset the counters in this case? But it + * should not matter, if the activity is low. + */ + if (avg_allocs < 100) + { + elog(DEBUG1, "rebalance skipped: not enough allocations (allocs: %u)", + avg_allocs); + return; + } + + /* + * Likewise, skip rebalancing if the misbalance is not significant. We + * consider it acceptable if the amount of allocations we'd need to + * redistribute is less than 10% of the average. + * + * XXX Again, these threshold are rather arbitrary. And maybe we should + * do the rabalancing in this case anyway, it's likely cheap and on a big + * system 10% can be quite a lot. + */ + if (delta_allocs < (avg_allocs * 0.1)) + { + elog(DEBUG1, "rebalance skipped: delta within limit (delta: %u, threshold: %u)", + delta_allocs, (uint32) (avg_allocs * 0.1)); + return; + } + + /* + * The actual rebalancing + * + * Partition with fewer than average allocations, should not redirect any + * allocations to other partitions. So just use weights with a single + * non-zero weight for the partition itself. + * + * Partition with more than average allocations, should not receive any + * redirected allocations, and instead it should redirect excess allocations + * to other partitions. + * + * The redistribution is "proportional" - if the excess allocations of a + * partition represent 10% of the "delta", then each partition that + * needs more allocations will get 10% of the gap from it. + * + * XXX We should add hysteresis, so that it does not oscillate or something + * like that. Maybe CLOCKSWEEP_HISTORY_COEFF already does that? + * + * XXX Ideally, the alternative partitions to use first would be the other + * partitions for the same node (if any). + */ + for (int i = 0; i < StrategyControl->num_partitions; i++) + { + ClockSweep *sweep = &StrategyControl->sweeps[i]; + uint8 balance[MAX_BUFFER_PARTITIONS]; + + /* lock, we're going to modify the balance weights */ + SpinLockAcquire(&sweep->clock_sweep_lock); + + /* reset the weights to start from scratch */ + memset(balance, 0, sizeof(uint8) * MAX_BUFFER_PARTITIONS); + + /* does this partition has fewer or more than avg_allocs? */ + if (allocs[i] < avg_allocs) + { + /* fewer - don't redirect any allocations elsewhere */ + balance[i] = 100; + } + else + { + /* + * more - redistribute the excess allocations + * + * Each "target" partition (with less than avg_allocs) should get + * a fraction proportional to (excess/delta) from this one. + */ + + /* fraction of the "total" delta */ + double delta_frac = (allocs[i] - avg_allocs) * 1.0 / delta_allocs; + + /* keep just enough allocations to meet the target */ + balance[i] = (100.0 * avg_allocs / allocs[i]); + + /* redirect the extra allocations */ + for (int j = 0; j < StrategyControl->num_partitions; j++) + { + /* How many allocations to receive from i-th partition? */ + uint32 receive_allocs = delta_frac * (avg_allocs - allocs[j]); + + /* ignore partitions that don't need additional allocations */ + if (allocs[j] > avg_allocs) + continue; + + /* fraction to redirect */ + balance[j] = (100.0 * receive_allocs / allocs[i]) + 0.5; + } + } + + /* combine the old and new weights (hysteresis) */ + for (int j = 0; j < MAX_BUFFER_PARTITIONS; j++) + { + sweep->balance[j] + = CLOCKSWEEP_HISTORY_COEFF * sweep->balance[j] + + (1.0 - CLOCKSWEEP_HISTORY_COEFF) * balance[j]; + } + + SpinLockRelease(&sweep->clock_sweep_lock); + } +} + /* * StrategySyncPrepare -- prepare for sync of all partitions * @@ -657,7 +1039,21 @@ StrategyCtlShmemInit(void *arg) /* Clear statistics */ StrategyControl->sweeps[i].completePasses = 0; pg_atomic_init_u32(&StrategyControl->sweeps[i].numBufferAllocs, 0); + pg_atomic_init_u32(&StrategyControl->sweeps[i].numRequestedAllocs, 0); pg_atomic_init_u64(&StrategyControl->sweeps[i].numTotalAllocs, 0); + pg_atomic_init_u64(&StrategyControl->sweeps[i].numTotalRequestedAllocs, 0); + + /* + * Initialize the weights - start by allocating 100% buffers from + * the current node / partition. + */ + for (int j = 0; j < MAX_BUFFER_PARTITIONS; j++) + { + if (i == j) + StrategyControl->sweeps[i].balance[i] = 100; + else + StrategyControl->sweeps[i].balance[j] = 0; + } } /* No pending notification */ @@ -1025,8 +1421,10 @@ StrategyRejectBuffer(BufferAccessStrategy strategy, BufferDesc *buf, bool from_r void ClockSweepPartitionGetInfo(int idx, - uint32 *complete_passes, uint32 *next_victim_buffer, - uint64 *buffer_total_allocs, uint32 *buffer_allocs) + uint32 *complete_passes, uint32 *next_victim_buffer, + uint64 *buffer_total_allocs, uint32 *buffer_allocs, + uint64 *buffer_total_req_allocs, uint32 *buffer_req_allocs, + int **weights) { ClockSweep *sweep = &StrategyControl->sweeps[idx]; @@ -1034,11 +1432,21 @@ ClockSweepPartitionGetInfo(int idx, /* get the clocksweep stats */ *complete_passes = sweep->completePasses; + + /* calculate the actual buffer ID */ *next_victim_buffer = pg_atomic_read_u32(&sweep->nextVictimBuffer); + *next_victim_buffer = sweep->firstBuffer + (*next_victim_buffer % sweep->numBuffers); - *buffer_allocs = pg_atomic_read_u32(&sweep->numBufferAllocs); *buffer_total_allocs = pg_atomic_read_u64(&sweep->numTotalAllocs); + *buffer_allocs = pg_atomic_read_u32(&sweep->numBufferAllocs); - /* calculate the actual buffer ID */ - *next_victim_buffer = sweep->firstBuffer + (*next_victim_buffer % sweep->numBuffers); + *buffer_total_req_allocs = pg_atomic_read_u64(&sweep->numTotalRequestedAllocs); + *buffer_req_allocs = pg_atomic_read_u32(&sweep->numRequestedAllocs); + + /* return the weights in a newly allocated array */ + *weights = palloc_array(int, StrategyControl->num_partitions); + for (int i = 0; i < StrategyControl->num_partitions; i++) + { + (*weights)[i] = (int) sweep->balance[i]; + } } diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h index 5b0cf6db2712f..d2f69bfba6870 100644 --- a/src/include/storage/buf_internals.h +++ b/src/include/storage/buf_internals.h @@ -599,6 +599,7 @@ extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy, extern bool StrategyRejectBuffer(BufferAccessStrategy strategy, BufferDesc *buf, bool from_ring); +extern void StrategySyncBalance(void); extern void StrategySyncPrepare(int *num_parts, uint32 *num_buf_alloc); extern int StrategySyncStart(int partition, uint32 *complete_passes, int *first_buffer, int *num_buffers); diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h index e0bb4cc1df1d4..02833b19b0c63 100644 --- a/src/include/storage/bufmgr.h +++ b/src/include/storage/bufmgr.h @@ -411,11 +411,13 @@ extern int GetAccessStrategyPinLimit(BufferAccessStrategy strategy); extern void FreeAccessStrategy(BufferAccessStrategy strategy); extern void ClockSweepPartitionGetInfo(int idx, - uint32 *complete_passes, - uint32 *next_victim_buffer, - uint64 *buffer_total_allocs, - uint32 *buffer_allocs); - + uint32 *complete_passes, + uint32 *next_victim_buffer, + uint64 *buffer_total_allocs, + uint32 *buffer_allocs, + uint64 *buffer_total_req_allocs, + uint32 *buffer_req_allocs, + int **weights); /* inline functions */ From d41c47ec72f12e0e5f84ed5c138365bab3ace15b Mon Sep 17 00:00:00 2001 From: Tomas Vondra Date: Tue, 2 Jun 2026 22:39:38 +0200 Subject: [PATCH 09/17] clock-sweep: scan all partitions When looking for a free buffer, scan all clock-sweep partitions, not just the "home" one. All buffers in the home partition may be pinned, in which case we should not fail. Instead, advance to the next partition, in a round-robin way, and only fail after scanning through all of them. --- src/backend/storage/buffer/freelist.c | 83 ++++++++++++++++------- src/test/recovery/t/027_stream_regress.pl | 5 -- 2 files changed, 57 insertions(+), 31 deletions(-) diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index a543fb12b2178..1ac1e3e3490a4 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -184,6 +184,9 @@ static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy, static void AddBufferToRing(BufferAccessStrategy strategy, BufferDesc *buf); static ClockSweep *ChooseClockSweep(bool balance); +static BufferDesc *StrategyGetBufferPartition(ClockSweep *sweep, + BufferAccessStrategy strategy, + uint64 *buf_state); /* * clocksweep allocation balancing @@ -251,10 +254,9 @@ static int clocksweep_partition_budget = 0; * id of the buffer now under the hand. */ static inline uint32 -ClockSweepTick(void) +ClockSweepTick(ClockSweep *sweep) { uint32 victim; - ClockSweep *sweep = ChooseClockSweep(true); /* * Atomically move hand ahead one buffer - if there's several processes @@ -486,7 +488,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r { BufferDesc *buf; int bgwprocno; - int trycounter; + ClockSweep *sweep, + *sweep_start; /* starting clock-sweep partition */ *from_ring = false; @@ -545,33 +548,61 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r /* * Use the "clock sweep" algorithm to find a free buffer * - * XXX Note that ClockSweepTick() is NUMA-aware, i.e. it only looks at - * buffers from a single partition, aligned with the NUMA node. That means - * a process "sweeps" only a fraction of buffers, even if the other buffers - * are better candidates for eviction. Maybe there should be some logic to - * "steal" buffers from other partitions or other nodes? - * - * XXX This only searches a single partition, which can result in "no - * unpinned buffers available" even if there are buffers in other - * partitions. Needs to scan partitions if needed, as a fallback. - * - * XXX Would that also mean we should have multiple bgwriters, one for each - * node, or would one bgwriter still handle all nodes? - * - * XXX Also, the trycounter should not be set to NBuffers, but to buffer - * count for that one partition. In fact, this should not call ClockSweepTick - * for every iteration. The call is likely quite expensive (does a lot - * of stuff), and also may return a different partition on each call. - * We should just do it once, and then do the for(;;) loop. And then - * maybe advance to the next partition, until we scan through all of them. + * Start with the "preferred" partition, and then proceed in a round-robin + * manner. If we cycle back to the starting partition, it means none of the + * partitions has unpinned buffers. */ - trycounter = NBuffers; + sweep = ChooseClockSweep(true); + sweep_start = sweep; + for (;;) + { + buf = StrategyGetBufferPartition(sweep, strategy, buf_state); + + /* found a buffer in the "sweep" partition, we're done */ + if (buf != NULL) + return buf; + + /* + * Try advancing to the next partition, round-robin (if last partition, + * wrap around to the beginning). + * + * XXX This is a bit ugly, there must be a better way to advance to the + * next partition. + */ + if (sweep == &StrategyControl->sweeps[StrategyControl->num_partitions - 1]) + sweep = StrategyControl->sweeps; + else + sweep++; + + /* we've scanned all partitions */ + if (sweep == sweep_start) + break; + } + + /* we shouldn't get here if there are unpinned buffers */ + elog(ERROR, "no unpinned buffers available"); +} + +/* + * StrategyGetBufferPartition + * get a free buffer from a single clock-sweep partition + * + * Returns NULL if there are no free (unpinned) buffers in the partition. +*/ +static BufferDesc * +StrategyGetBufferPartition(ClockSweep *sweep, BufferAccessStrategy strategy, + uint64 *buf_state) +{ + BufferDesc *buf; + int trycounter; + + trycounter = sweep->numBuffers; for (;;) { uint64 old_buf_state; uint64 local_buf_state; - buf = GetBufferDescriptor(ClockSweepTick()); + buf = GetBufferDescriptor(ClockSweepTick(sweep)); /* * Check whether the buffer can be used and pin it if so. Do this @@ -599,7 +630,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r * one eventually, but it's probably better to fail than * to risk getting stuck in an infinite loop. */ - elog(ERROR, "no unpinned buffers available"); + return NULL; } break; } @@ -618,7 +649,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state, local_buf_state)) { - trycounter = NBuffers; + trycounter = sweep->numBuffers; break; } } diff --git a/src/test/recovery/t/027_stream_regress.pl b/src/test/recovery/t/027_stream_regress.pl index f68e08e569751..ae97729784943 100644 --- a/src/test/recovery/t/027_stream_regress.pl +++ b/src/test/recovery/t/027_stream_regress.pl @@ -18,11 +18,6 @@ $node_primary->append_conf('postgresql.conf', 'max_prepared_transactions = 10'); -# The default is 1MB, which is not enough with clock-sweep partitioning. -# Increase to 32MB, so that we don't get "no unpinned buffers". -$node_primary->append_conf('postgresql.conf', - 'shared_buffers = 32MB'); - # Enable pg_stat_statements to force tests to do query jumbling. # pg_stat_statements.max should be large enough to hold all the entries # of the regression database. From 600e6de0a30282b06258b497ccfcc216c5e1992b Mon Sep 17 00:00:00 2001 From: Jakub Wartak Date: Thu, 11 Jun 2026 12:38:43 +0200 Subject: [PATCH 10/17] Add parttioned clocksweep and NUMA goodies. 1. Add three clocksweep GUCs to allow manipulation of partitoned clocksweep in runtime. 2. Add pg_buffercache_set_weights(int, int[]) to alter partition allocs (with clocksweep_balance_recalc=off). 3. Add debug_numa_node GUC to pin to NUMA node. --- .../pg_buffercache--1.7--1.8.sql | 7 ++ contrib/pg_buffercache/pg_buffercache_pages.c | 37 ++++++++ src/backend/storage/buffer/freelist.c | 86 ++++++++++++++++++- src/backend/tcop/postgres.c | 57 ++++++++++++ src/backend/utils/misc/guc_parameters.dat | 31 +++++++ src/include/miscadmin.h | 1 + src/include/storage/bufmgr.h | 6 ++ 7 files changed, 224 insertions(+), 1 deletion(-) diff --git a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql index 43d2e84f9d279..9d8f4969555af 100644 --- a/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql +++ b/contrib/pg_buffercache/pg_buffercache--1.7--1.8.sql @@ -25,9 +25,16 @@ CREATE VIEW pg_buffercache_partitions AS num_req_allocs bigint, -- handled allocs (current cycle) weights int[]); -- balancing weights +-- Register the function to set clock-sweep balance weights. +CREATE FUNCTION pg_buffercache_set_partition(IN partition int, IN weights int[]) +RETURNS void +AS 'MODULE_PATHNAME', 'pg_buffercache_set_partition' +LANGUAGE C VOLATILE PARALLEL UNSAFE; + -- Don't want these to be available to public. REVOKE ALL ON FUNCTION pg_buffercache_partitions() FROM PUBLIC; REVOKE ALL ON pg_buffercache_partitions FROM PUBLIC; +REVOKE ALL ON FUNCTION pg_buffercache_set_partition(int, int[]) FROM PUBLIC; GRANT EXECUTE ON FUNCTION pg_buffercache_partitions() TO pg_monitor; GRANT SELECT ON pg_buffercache_partitions TO pg_monitor; diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c index c3cd24af097e9..0265241e139ac 100644 --- a/contrib/pg_buffercache/pg_buffercache_pages.c +++ b/contrib/pg_buffercache/pg_buffercache_pages.c @@ -82,6 +82,7 @@ PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty); PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_relation); PG_FUNCTION_INFO_V1(pg_buffercache_mark_dirty_all); PG_FUNCTION_INFO_V1(pg_buffercache_partitions); +PG_FUNCTION_INFO_V1(pg_buffercache_set_partition); /* Only need to touch memory once per backend process lifetime */ @@ -1078,3 +1079,39 @@ pg_buffercache_partitions(PG_FUNCTION_ARGS) else SRF_RETURN_DONE(funcctx); } + +/* + * Set the clock-sweep balance weights for a single partition. + */ +Datum +pg_buffercache_set_partition(PG_FUNCTION_ARGS) +{ + int partition = PG_GETARG_INT32(0); + ArrayType *array = PG_GETARG_ARRAYTYPE_P(1); + Datum *elems; + bool *nulls; + int nelems; + int *weights; + + if (ARR_NDIM(array) > 1) + ereport(ERROR, + (errcode(ERRCODE_INVALID_PARAMETER_VALUE), + errmsg("weights must be a one-dimensional array"))); + + deconstruct_array_builtin(array, INT4OID, &elems, &nulls, &nelems); + + weights = palloc_array(int, nelems); + for (int i = 0; i < nelems; i++) + { + if (nulls[i]) + ereport(ERROR, + (errcode(ERRCODE_INVALID_PARAMETER_VALUE), + errmsg("weights must not contain NULL values"))); + + weights[i] = DatumGetInt32(elems[i]); + } + + ClockSweepSetWeights(partition, weights, nelems); + + PG_RETURN_VOID(); +} diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index 1ac1e3e3490a4..e677c71e0b3d9 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -55,6 +55,26 @@ */ #define CLOCKSWEEP_HISTORY_COEFF 0.5 +/* + * GUCs controlling the NUMA-aware clock-sweep behavior. + * + * clocksweep_balance - when enabled, allocations may get redirected between + * clock-sweep partitions to keep them balanced (see StrategySyncBalance). + * + * clocksweep_balance_recalc - when enabled, the balance weights are + * periodically recalculated (see StrategySyncBalance). Disabling this keeps + * the current weights, e.g. ones configured manually using + * pg_buffercache_set_partition(), while still using them to balance + * allocations (if clocksweep_balance is enabled). + * + * clocksweep_scan_all_partitions - when enabled, looking for a free buffer + * scans all clock-sweep partitions (in a round-robin way), not just the + * backend's "home" partition. + */ +bool clocksweep_balance = true; +bool clocksweep_balance_recalc = true; +bool clocksweep_scan_all_partitions = true; + /* * Information about one partition of the ClockSweep (on a subset of buffers). * @@ -436,8 +456,11 @@ ChooseClockSweep(bool balance) * When balancing allocations, redirect the allocations to other partitions * according to the budgets. We move through partitions in a round-robin way, * after allocating the "budget" of allocations from the current one. + * + * Balancing can be disabled at runtime using the clocksweep_balance GUC, in + * which case allocations always stay in the backend's "home" partition. */ - if (balance) + if (balance && clocksweep_balance) { /* * Ran out of budget from the current partition? Move to the next one @@ -551,6 +574,10 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r * Start with the "preferred" partition, and then proceed in a round-robin * manner. If we cycle back to the starting partition, it means none of the * partitions has unpinned buffers. + * + * Scanning of the other partitions can be disabled at runtime using the + * clocksweep_scan_all_partitions GUC. In that case we only scan the + * backend's "home" partition, and fail if it has no unpinned buffers. */ sweep = ChooseClockSweep(true); sweep_start = sweep; @@ -562,6 +589,10 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_r if (buf != NULL) return buf; + /* don't look at other partitions unless allowed to */ + if (!clocksweep_scan_all_partitions) + break; + /* * Try advancing to the next partition, round-robin (if last partition, * wrap around to the beginning). @@ -739,6 +770,9 @@ StrategySyncBalance(void) avg_allocs, /* average allocations (per partition) */ delta_allocs = 0; /* sum of allocs above average */ + if (!clocksweep_balance || !clocksweep_balance_recalc) + return; + /* * Collect the number of allocations requested in the past interval. * While at it, reset the counter to start the new interval. @@ -1481,3 +1515,53 @@ ClockSweepPartitionGetInfo(int idx, (*weights)[i] = (int) sweep->balance[i]; } } + +/* + * ClockSweepSetWeights override the clock-sweep balance weights of a single + * partition. + */ +void +ClockSweepSetWeights(int partition, int *weights, int nweights) +{ + ClockSweep *sweep; + + /* + * Disallow manual weights while the automatic recalculation is enabled, as + * StrategySyncBalance would just recompute (and overwrite) them. + */ + if (clocksweep_balance_recalc) + ereport(ERROR, + (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE), + errmsg("cannot set clock-sweep weights while debug_clocksweep_balance_recalc is enabled"), + errhint("Set debug_clocksweep_balance_recalc to off before setting the weights manually."))); + + if ((partition < 0) || (partition >= StrategyControl->num_partitions)) + ereport(ERROR, + (errcode(ERRCODE_INVALID_PARAMETER_VALUE), + errmsg("clock-sweep partition %d out of range", partition), + errhint("There are %d clock-sweep partitions, numbered from 0.", + StrategyControl->num_partitions))); + + if (nweights != StrategyControl->num_partitions) + ereport(ERROR, + (errcode(ERRCODE_INVALID_PARAMETER_VALUE), + errmsg("number of weights (%d) does not match number of clock-sweep partitions (%d)", + nweights, StrategyControl->num_partitions))); + + for (int i = 0; i < nweights; i++) + { + if ((weights[i] < 0) || (weights[i] > 100)) + ereport(ERROR, + (errcode(ERRCODE_INVALID_PARAMETER_VALUE), + errmsg("clock-sweep weight %d out of range", + weights[i]), + errhint("Each weight must be between 0 and 100."))); + } + + sweep = &StrategyControl->sweeps[partition]; + + SpinLockAcquire(&sweep->clock_sweep_lock); + for (int j = 0; j < nweights; j++) + sweep->balance[j] = (uint8) weights[j]; + SpinLockRelease(&sweep->clock_sweep_lock); +} diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c index b6bdfe213feec..d8e0d48e0ce39 100644 --- a/src/backend/tcop/postgres.c +++ b/src/backend/tcop/postgres.c @@ -86,6 +86,7 @@ #include "utils/timeout.h" #include "utils/timestamp.h" #include "utils/varlena.h" +#include /* ---------------- * global variables @@ -110,6 +111,9 @@ int client_connection_check_interval = 0; /* flags for non-system relation kinds to restrict use */ int restrict_nonsystem_relation_kind; +/* NUMA node to pin the backend to at query start; -1 disables pinning */ +int debug_numa_node = -1; + /* * Include signal sender PID/UID in the server log when available * (SA_SIGINFO). The caller must supply the already-captured pid and uid @@ -1021,6 +1025,53 @@ pg_plan_queries(List *querytrees, const char *query_string, int cursorOptions, } +/* + * process_debug_numa_node + * + * If the debug_numa_node GUC is set (>= 0), pin this backend to run on the + * CPUs of the requested NUMA node. -1 disables it (default) + */ +static void +process_debug_numa_node(void) +{ +#ifdef USE_LIBNUMA + static int applied_numa_node = -1; + + if (debug_numa_node == applied_numa_node) + return; + + /* Nothing we can do if the kernel/library has no NUMA support. */ + if (numa_available() < 0) + { + applied_numa_node = debug_numa_node; + return; + } + + if (debug_numa_node < 0) + { + /* Pinning disabled: allow running on all nodes again. */ + numa_run_on_node_mask(numa_all_nodes_ptr); + } + else if (debug_numa_node > numa_max_node()) + { + ereport(WARNING, + (errmsg("debug_numa_node %d exceeds the highest available NUMA node %d, ignoring", + debug_numa_node, numa_max_node()))); + } + else if (numa_run_on_node(debug_numa_node) != 0) + { + ereport(WARNING, + (errmsg("could not pin backend to NUMA node %d: %m", + debug_numa_node))); + } + else + elog(DEBUG1, "pinned backend to NUMA node %d", debug_numa_node); + + applied_numa_node = debug_numa_node; +#endif +} + + /* * exec_simple_query * @@ -1045,6 +1096,9 @@ exec_simple_query(const char *query_string) pgstat_report_activity(STATE_RUNNING, query_string); + /* Pin the backend to the configured NUMA node, if requested. */ + process_debug_numa_node(); + TRACE_POSTGRESQL_QUERY_START(query_string); /* @@ -2214,6 +2268,9 @@ exec_execute_message(const char *portal_name, long max_rows) pgstat_report_activity(STATE_RUNNING, sourceText); + /* Pin the backend to the configured NUMA node, if requested. */ + process_debug_numa_node(); + foreach(lc, portal->stmts) { PlannedStmt *stmt = lfirst_node(PlannedStmt, lc); diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat index d0a6c77552dea..c2eac5f535025 100644 --- a/src/backend/utils/misc/guc_parameters.dat +++ b/src/backend/utils/misc/guc_parameters.dat @@ -632,6 +632,27 @@ boot_val => 'DEFAULT_ASSERT_ENABLED', }, +{ name => 'debug_clocksweep_balance', type => 'bool', context => 'PGC_USERSET', group => 'DEVELOPER_OPTIONS', + short_desc => 'Enables balancing of buffer allocations between clock-sweep partitions.', + flags => 'GUC_NOT_IN_SAMPLE', + variable => 'clocksweep_balance', + boot_val => 'true' +}, + +{ name => 'debug_clocksweep_balance_recalc', type => 'bool', context => 'PGC_USERSET', group => 'DEVELOPER_OPTIONS', + short_desc => 'Enables periodic recalculation of clock-sweep partition balance weights.', + flags => 'GUC_NOT_IN_SAMPLE', + variable => 'clocksweep_balance_recalc', + boot_val => 'true' +}, + +{ name => 'debug_clocksweep_scan_all_partitions', type => 'bool', context => 'PGC_USERSET', group => 'DEVELOPER_OPTIONS', + short_desc => 'Enables scanning all clock-sweep partitions when looking for a free buffer.', + flags => 'GUC_NOT_IN_SAMPLE', + variable => 'clocksweep_scan_all_partitions', + boot_val => 'true' +}, + { name => 'debug_copy_parse_plan_trees', type => 'bool', context => 'PGC_SUSET', group => 'DEVELOPER_OPTIONS', short_desc => 'Set this to force all parse and plan trees to be passed through copyObject(), to facilitate catching errors and omissions in copyObject().', flags => 'GUC_NOT_IN_SAMPLE', @@ -684,6 +705,16 @@ options => 'debug_logical_replication_streaming_options', }, +{ name => 'debug_numa_node', type => 'int', context => 'PGC_USERSET', group => 'DEVELOPER_OPTIONS', + short_desc => 'Pins the backend to the given NUMA node at query start.', + long_desc => '-1 (the default) disables pinning.', + flags => 'GUC_NOT_IN_SAMPLE', + variable => 'debug_numa_node', + boot_val => '-1', + min => '-1', + max => 'INT_MAX', +}, + { name => 'debug_parallel_query', type => 'enum', context => 'PGC_USERSET', group => 'DEVELOPER_OPTIONS', short_desc => 'Forces the planner\'s use parallel query nodes.', long_desc => 'This can be useful for testing the parallel query infrastructure by forcing the planner to generate plans that contain nodes that perform tuple communication between workers and the main process.', diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h index ef72549ebc66e..bd520356231db 100644 --- a/src/include/miscadmin.h +++ b/src/include/miscadmin.h @@ -215,6 +215,7 @@ extern PGDLLIMPORT bool MyDatabaseHasLoginEventTriggers; extern PGDLLIMPORT bool shmem_populate; extern PGDLLIMPORT bool shmem_interleave; +extern PGDLLIMPORT int debug_numa_node; /* diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h index 02833b19b0c63..b5c3873acae68 100644 --- a/src/include/storage/bufmgr.h +++ b/src/include/storage/bufmgr.h @@ -190,6 +190,11 @@ extern PGDLLIMPORT int bgwriter_lru_maxpages; extern PGDLLIMPORT double bgwriter_lru_multiplier; extern PGDLLIMPORT bool track_io_timing; +/* in freelist.c */ +extern PGDLLIMPORT bool clocksweep_balance; +extern PGDLLIMPORT bool clocksweep_balance_recalc; +extern PGDLLIMPORT bool clocksweep_scan_all_partitions; + #define DEFAULT_EFFECTIVE_IO_CONCURRENCY 16 #define DEFAULT_MAINTENANCE_IO_CONCURRENCY 16 extern PGDLLIMPORT int effective_io_concurrency; @@ -418,6 +423,7 @@ extern void ClockSweepPartitionGetInfo(int idx, uint64 *buffer_total_req_allocs, uint32 *buffer_req_allocs, int **weights); +extern void ClockSweepSetWeights(int partition, int *weights, int nweights); /* inline functions */ From d043b8cdfe23f00e50fcd26c61e8a8559d5fcb89 Mon Sep 17 00:00:00 2001 From: Jakub Wartak Date: Tue, 30 Jun 2026 14:22:02 +0200 Subject: [PATCH 11/17] clock-sweep: cached CPU/NUMA node and more locality-aware balancing Enhancements on top of 0001-0007, to have sligthly better NUMA locality and perfromance. 1. Cache numa_node_of_cpu()/sched_getcpu() per backend in ClockSweepPartitionIndex(), refreshing every CLOCKSWEEP_CPU_NODE_REFRESH allocations rather than on every call (visible hot buffer path in perf) 2. CLOCKSWEEP_BALANCE_THRESHOLD - make it less likely to redirect on any surplus of allocations (so scatter buffers LESS onto remote nodes). With this, it redirects its allocations to other (remote?) partitions when the allocation exceeds the per-partition average allocation rate by this percentage factor . 3. Avoid redirects to "idle" partitions: a redirect partition target must have some traffic which is at least 2x our demand. This elimnates cold partitions, but we can still reach them using scan-all-partitions fallback. --- src/backend/storage/buffer/freelist.c | 85 +++++++++++++++++++++++---- 1 file changed, 74 insertions(+), 11 deletions(-) diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index e677c71e0b3d9..d64c2c67eb6d5 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -55,6 +55,9 @@ */ #define CLOCKSWEEP_HISTORY_COEFF 0.5 +/* How often backend should re-fetch the CPU/node on which it is running on? */ +#define CLOCKSWEEP_CPU_NODE_REFRESH 128 + /* * GUCs controlling the NUMA-aware clock-sweep behavior. * @@ -70,6 +73,7 @@ * clocksweep_scan_all_partitions - when enabled, looking for a free buffer * scans all clock-sweep partitions (in a round-robin way), not just the * backend's "home" partition. + * */ bool clocksweep_balance = true; bool clocksweep_balance_recalc = true; @@ -368,13 +372,29 @@ ClockSweepPartitionIndex(void) #ifdef USE_LIBNUMA if (shared_buffers_numa) { - int cpu; + /* + * Cache the CPU/NUMA node, refreshing only every CLOCKSWEEP_CPU_NODE_REFRESH + * allocations. It appears that sched_getcpu()/numa_node_of_cpu() are not free. + * On some platforms it take price of full system call, or the rest (x86_64?) + * is can be use VDSO optimization. The backend rarely migrates between NUMA + * nodes, and the balance logic only needs to notice migration after some time, + * so an occasional refresh is good enough. + */ + static int cached_node = -1; + static uint32 refresh_counter = 0; + + if (cached_node < 0 || (refresh_counter++ % CLOCKSWEEP_CPU_NODE_REFRESH) == 0) + { + int cpu; - /* XXX do we need to check sched_getcpu is available, somehow? */ - if ((cpu = sched_getcpu()) < 0) + /* XXX do we need to check sched_getcpu is available, somehow? */ + if ((cpu = sched_getcpu()) < 0) elog(ERROR, "sched_getcpu failed: %m"); - node = numa_node_of_cpu(cpu); + /* XXX/JW: use libnuma wrapper for this */ + cached_node = numa_node_of_cpu(cpu); + } + node = cached_node; } #endif @@ -768,7 +788,8 @@ StrategySyncBalance(void) uint32 total_allocs = 0, /* total number of allocations */ avg_allocs, /* average allocations (per partition) */ - delta_allocs = 0; /* sum of allocs above average */ + delta_allocs = 0, /* sum of allocs above average */ + redirect_cutoff; /* redirect only above this many allocs */ if (!clocksweep_balance || !clocksweep_balance_recalc) return; @@ -852,6 +873,20 @@ StrategySyncBalance(void) return; } + /* + * A partition only redirects allocations to other partitions when it + * exceeds the average by more than some threshold percent. + * Below this cutoff we keep allocations local, to preserve NUMA locality. + * + * TODO: maybe better value is possible. On 4s with 25 I've got good results, + * but with value of 50 I've got slight degradation. Maybe it should + * be equal to 100/numa_nodes ? + * + */ +#define CLOCKSWEEP_CUTOFF_THRESHOLD 25 + redirect_cutoff = avg_allocs + + (uint32) ((uint64) avg_allocs * CLOCKSWEEP_CUTOFF_THRESHOLD / 100); + /* * The actual rebalancing * @@ -884,10 +919,15 @@ StrategySyncBalance(void) /* reset the weights to start from scratch */ memset(balance, 0, sizeof(uint8) * MAX_BUFFER_PARTITIONS); - /* does this partition has fewer or more than avg_allocs? */ - if (allocs[i] < avg_allocs) + /* + * Does this partition exceed its fair share by more than the + * threshold? If not, keep all allocations local - redirecting them + * would push memory onto remote NUMA nodes for no real benefit when + * the load is already close to balanced. + */ + if (allocs[i] <= redirect_cutoff) { - /* fewer - don't redirect any allocations elsewhere */ + /* near fair share (or below) - keep allocations local */ balance[i] = 100; } else @@ -902,22 +942,45 @@ StrategySyncBalance(void) /* fraction of the "total" delta */ double delta_frac = (allocs[i] - avg_allocs) * 1.0 / delta_allocs; - /* keep just enough allocations to meet the target */ - balance[i] = (100.0 * avg_allocs / allocs[i]); + /* how much we keep local; we hand out the rest below */ + int kept = 100; /* redirect the extra allocations */ for (int j = 0; j < StrategyControl->num_partitions; j++) { /* How many allocations to receive from i-th partition? */ uint32 receive_allocs = delta_frac * (avg_allocs - allocs[j]); + int w; + + /* do not redirect to ourselves */ + if (j == i) + continue; /* ignore partitions that don't need additional allocations */ if (allocs[j] > avg_allocs) continue; + /* + * Only use other partitions that actually have demand of + * their own (avoid idle). If we fail, there's always the + * scan-all-partitions fallback. + * + * TODO:: just guessing,heuristics + */ + if (allocs[j] < (avg_allocs / 2)) + continue; + /* fraction to redirect */ - balance[j] = (100.0 * receive_allocs / allocs[i]) + 0.5; + w = (int) ((100.0 * receive_allocs / allocs[i]) + 0.5); + balance[j] = w; + kept -= w; } + + /* avoid negative balances */ + if (kept > 0) + balance[i] = kept; + else + balance[i] = 1; } /* combine the old and new weights (hysteresis) */ From 9b70542d60582413007a87064f8a65d4bf13acd8 Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Sun, 4 Oct 2026 10:39:06 -0400 Subject: [PATCH 12/17] Replace the usage_count clock sweep with a cooling-stage evictor Replace the 0..5 usage_count buffer-replacement policy with a cooling-stage clock (the LeanStore / 2Q-A1 model): a buffer is either HOT (recently used) or COOL (an eviction candidate), with "pinned" being the existing refcount. There is no per-buffer access counter. - A demand-loaded page is admitted COOL (probationary), not HOT. A second access via PinBuffer promotes it COOL -> HOT (the rescue). A page touched once -- a sequential scan -- therefore fills and drains the COOL stage and is evicted from it without displacing the HOT working set. - The clock sweep in StrategyGetBuffer() reclaims an already-COOL, unpinned buffer, pinning it with a CAS so a racing PinBuffer always wins. When it passes a HOT buffer it demotes it to COOL and keeps scanning, so a HOT buffer survives the visit that cools it and is only reclaimed if it is still COOL when the hand comes around again. That gives every buffer one full sweep of grace in which a new access can promote it back to HOT. - A strategy (ring) access deliberately does not promote, which is the cooling-state form of the existing rule that ring buffers must not evict others from the pool: a buffer the ring keeps recycling stays COOL and so stays reusable by GetBufferFromRing(), while an access from outside the ring promotes it to HOT and thereby removes it from the ring's reuse set. GetBufferFromRing() tests that state directly, replacing the stock "usage_count > 1 means someone else touched it" test. The usage_count field is reinterpreted in place as the one-bit cooling state. The 64-bit buffer-state layout -- refcount, flag and lock offsets and their StaticAsserts -- is unchanged; only the meaning of the field and the instructions that touch it change. BM_MAX_USAGE_COUNT becomes BUF_COOLSTATE_HOT (1), so the pin fast path saturates at HOT. Local (temp-table) buffers get the same two-state treatment. contrib/pg_buffercache reports the cooling state in its usagecount column (0 = COOL, 1 = HOT), and pg_buffercache_summary and pg_buffercache_usage_counts() follow, so the usage-count histogram now has two populated buckets instead of six. This is a replacement-policy change only. It does not alter the background writer's pacing or write limits, and it keeps BufferAccessStrategy rings. --- contrib/pg_buffercache/pg_buffercache_pages.c | 6 +- src/backend/storage/buffer/bufmgr.c | 50 ++++++++++------- src/backend/storage/buffer/freelist.c | 37 ++++++++----- src/backend/storage/buffer/localbuf.c | 13 +++-- src/include/storage/buf_internals.h | 55 +++++++++++++++---- 5 files changed, 107 insertions(+), 54 deletions(-) diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c index 0265241e139ac..e3aa16884844d 100644 --- a/contrib/pg_buffercache/pg_buffercache_pages.c +++ b/contrib/pg_buffercache/pg_buffercache_pages.c @@ -167,7 +167,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS) reldatabase = bufHdr->tag.dbOid; forknum = BufTagGetForkNum(&bufHdr->tag); blocknum = bufHdr->tag.blockNum; - usagecount = BUF_STATE_GET_USAGECOUNT(buf_state); + usagecount = BUF_STATE_GET_COOLSTATE(buf_state); pinning_backends = BUF_STATE_GET_REFCOUNT(buf_state); if (buf_state & BM_DIRTY) @@ -611,7 +611,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS) if (buf_state & BM_VALID) { buffers_used++; - usagecount_total += BUF_STATE_GET_USAGECOUNT(buf_state); + usagecount_total += BUF_STATE_GET_COOLSTATE(buf_state); if (buf_state & BM_DIRTY) buffers_dirty++; @@ -661,7 +661,7 @@ pg_buffercache_usage_counts(PG_FUNCTION_ARGS) CHECK_FOR_INTERRUPTS(); - usage_count = BUF_STATE_GET_USAGECOUNT(buf_state); + usage_count = BUF_STATE_GET_COOLSTATE(buf_state); usage_counts[usage_count]++; if (buf_state & BM_DIRTY) diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c index d58397f098099..36f2d48cb912f 100644 --- a/src/backend/storage/buffer/bufmgr.c +++ b/src/backend/storage/buffer/bufmgr.c @@ -2335,7 +2335,13 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum, * checkpoints, except for their "init" forks, which need to be treated * just like permanent relations. */ - set_bits |= BM_TAG_VALID | BUF_USAGECOUNT_ONE; + set_bits |= BM_TAG_VALID; + + /* + * Admit the newly loaded page COOL (probation); a second access via + * PinBuffer promotes it to HOT. This is what makes a one-touch scan + * self-evicting -- see the cooling-state notes in buf_internals.h. + */ if (relpersistence == RELPERSISTENCE_PERMANENT || forkNum == INIT_FORKNUM) set_bits |= BM_PERMANENT; @@ -3004,7 +3010,12 @@ ExtendBufferedRelShared(BufferManagerRelation bmr, victim_buf_hdr->tag = tag; - set_bits |= BM_TAG_VALID | BUF_USAGECOUNT_ONE; + set_bits |= BM_TAG_VALID; + + /* + * Admit COOL (probation); see the comment at the other admission + * site and the cooling-state notes in buf_internals.h. + */ if (bmr.relpersistence == RELPERSISTENCE_PERMANENT || fork == INIT_FORKNUM) set_bits |= BM_PERMANENT; @@ -3334,21 +3345,22 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy, /* increase refcount */ buf_state += BUF_REFCOUNT_ONE; - if (strategy == NULL) - { - /* Default case: increase usagecount unless already max. */ - if (BUF_STATE_GET_USAGECOUNT(buf_state) < BM_MAX_USAGE_COUNT) - buf_state += BUF_USAGECOUNT_ONE; - } - else - { - /* - * Ring buffers shouldn't evict others from pool. Thus we - * don't make usagecount more than 1. - */ - if (BUF_STATE_GET_USAGECOUNT(buf_state) == 0) - buf_state += BUF_USAGECOUNT_ONE; - } + /* + * Accessing a resident buffer promotes it to HOT (the 2Q rescue): a + * page admitted COOL on probation joins the hot working set on its + * second touch. The cooling state saturates at BUF_COOLSTATE_HOT, + * so this never overflows the field. + * + * A strategy (ring) access deliberately does not promote, which is + * the cooling-state form of the stock rule that ring buffers must + * not evict others from the pool: a buffer the ring keeps recycling + * stays COOL and so stays reusable by GetBufferFromRing(), while an + * access from outside the ring promotes it to HOT and thereby takes + * it out of the ring's reuse set. + */ + if (strategy == NULL && + BUF_STATE_GET_COOLSTATE(buf_state) < BUF_COOLSTATE_HOT) + buf_state += BUF_COOLSTATE_ONE; if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state, buf_state)) @@ -4216,7 +4228,7 @@ BgBufferSync(WritebackContext *wb_context) * Returns a bitmask containing the following flag bits: * BUF_WRITTEN: we wrote the buffer. * BUF_REUSABLE: buffer is available for replacement, ie, it has - * pin count 0 and usage count 0. + * pin count 0 and is COOL (an eviction candidate). * * (BUF_WRITTEN could be set in error if FlushBuffer finds the buffer clean * after locking it, but we don't care all that much.) @@ -4245,7 +4257,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context) buf_state = LockBufHdr(bufHdr); if (BUF_STATE_GET_REFCOUNT(buf_state) == 0 && - BUF_STATE_GET_USAGECOUNT(buf_state) == 0) + BUF_STATE_GET_COOLSTATE(buf_state) == BUF_COOLSTATE_COOL) { result |= BUF_REUSABLE; } diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index d64c2c67eb6d5..6e41a56c5b977 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -664,11 +664,7 @@ StrategyGetBufferPartition(ClockSweep *sweep, BufferAccessStrategy strategy, { local_buf_state = old_buf_state; - /* - * If the buffer is pinned or has a nonzero usage_count, we cannot - * use it; decrement the usage_count (unless pinned) and keep - * scanning. - */ + /* If the buffer is pinned we cannot use it; keep scanning. */ if (BUF_STATE_GET_REFCOUNT(local_buf_state) != 0) { @@ -693,9 +689,17 @@ StrategyGetBufferPartition(ClockSweep *sweep, BufferAccessStrategy strategy, continue; } - if (BUF_STATE_GET_USAGECOUNT(local_buf_state) != 0) + if (BUF_STATE_GET_COOLSTATE(local_buf_state) != BUF_COOLSTATE_COOL) { - local_buf_state -= BUF_USAGECOUNT_ONE; + /* + * HOT buffer: cool it in place this tick and keep scanning. We + * do NOT claim it now -- a demoted buffer only becomes a victim + * on a later tick, so a HOT buffer always survives the pass that + * cools it and gets a full sweep of grace in which a new access + * can promote it back to HOT. Cooling is progress toward a + * victim, so reset trycounter. + */ + local_buf_state &= ~BUF_USAGECOUNT_MASK; /* HOT -> COOL */ if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state, local_buf_state)) @@ -706,7 +710,7 @@ StrategyGetBufferPartition(ClockSweep *sweep, BufferAccessStrategy strategy, } else { - /* pin the buffer if the CAS succeeds */ + /* COOL and unpinned: claim it. Pin if the CAS succeeds. */ local_buf_state += BUF_REFCOUNT_ONE; if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state, @@ -1433,14 +1437,17 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state) /* * If the buffer is pinned we cannot use it under any circumstances. * - * If usage_count is 0 or 1 then the buffer is fair game (we expect 1, - * since our own previous usage of the ring element would have left it - * there, but it might've been decremented by clock-sweep since then). - * A higher usage_count indicates someone else has touched the buffer, - * so we shouldn't re-use it. + * If it is unpinned but has been promoted to HOT, an access from + * outside the ring touched it since we last cycled past this slot, so + * it has joined the working set and we must not recycle it -- tell the + * caller to get a fresh victim from the sweep instead. This replaces + * the stock "usage_count > 1 means someone else touched it" test: a + * slot the ring keeps reusing stays COOL (ring accesses do not + * promote), and an out-of-ring PinBuffer() is exactly what makes it + * HOT. */ - if (BUF_STATE_GET_REFCOUNT(local_buf_state) != 0 - || BUF_STATE_GET_USAGECOUNT(local_buf_state) > 1) + if (BUF_STATE_GET_REFCOUNT(local_buf_state) != 0 || + BUF_STATE_GET_COOLSTATE(local_buf_state) != BUF_COOLSTATE_COOL) break; /* See equivalent code in PinBuffer() */ diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c index 4870c8e13d010..f243411292746 100644 --- a/src/backend/storage/buffer/localbuf.c +++ b/src/backend/storage/buffer/localbuf.c @@ -167,7 +167,7 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum, buf_state = pg_atomic_read_u64(&bufHdr->state); buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK); - buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE; + buf_state |= BM_TAG_VALID; /* admit COOL (probation) */ pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state); *foundPtr = false; @@ -248,9 +248,10 @@ GetLocalVictimBuffer(void) { uint64 buf_state = pg_atomic_read_u64(&bufHdr->state); - if (BUF_STATE_GET_USAGECOUNT(buf_state) > 0) + if (BUF_STATE_GET_COOLSTATE(buf_state) != BUF_COOLSTATE_COOL) { - buf_state -= BUF_USAGECOUNT_ONE; + /* HOT: demote it to COOL and keep scanning. */ + buf_state &= ~BUF_USAGECOUNT_MASK; pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state); trycounter = NLocBuffer; } @@ -454,7 +455,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr, victim_buf_hdr->tag = tag; - buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE; + buf_state |= BM_TAG_VALID; /* admit COOL (probation) */ pg_atomic_unlocked_write_u64(&victim_buf_hdr->state, buf_state); @@ -839,9 +840,9 @@ PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount) NLocalPinnedBuffers++; buf_state += BUF_REFCOUNT_ONE; if (adjust_usagecount && - BUF_STATE_GET_USAGECOUNT(buf_state) < BM_MAX_USAGE_COUNT) + BUF_STATE_GET_COOLSTATE(buf_state) < BUF_COOLSTATE_HOT) { - buf_state += BUF_USAGECOUNT_ONE; + buf_state += BUF_COOLSTATE_ONE; } pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state); diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h index d2f69bfba6870..6c53ecc5ed0bd 100644 --- a/src/include/storage/buf_internals.h +++ b/src/include/storage/buf_internals.h @@ -67,6 +67,36 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS + BUF_L #define BUF_USAGECOUNT_ONE \ (UINT64CONST(1) << BUF_REFCOUNT_BITS) +/* + * Cooling state (LeanStore / 2Q-A1 cooling-stage clock sweep). + * + * The field historically used for the 0..5 usage_count now holds a single + * cooling-state bit: HOT (recently accessed, not an eviction candidate) or + * COOL (an eviction candidate). We reuse BUF_USAGECOUNT_ONE as the unit so + * the buffer-state bit geography -- refcount, flag, and lock offsets, and the + * 64-bit StaticAsserts -- is unchanged; only the meaning of the field and the + * instructions that touch it change. + * + * A demand-loaded page is admitted COOL (probation); a second access promotes + * it to HOT (the rescue). The sweep reclaims an already-COOL buffer and + * demotes a HOT one to COOL as it passes, so a HOT buffer survives the visit + * that cools it and is only reclaimed if it is still COOL when the hand comes + * around again. A page touched once -- a sequential scan -- therefore fills + * and drains the COOL stage without displacing the HOT working set: scan + * resistance intrinsic to the replacement algorithm. + */ +#define BUF_COOLSTATE_COOL 0 +#define BUF_COOLSTATE_HOT 1 +#define BUF_COOLSTATE_ONE BUF_USAGECOUNT_ONE + +/* + * The cooling state is one bit, so the field must be at least that wide. + * Assert it here so a future change to BUF_USAGECOUNT_BITS cannot silently + * narrow the field out from under the cooling state. + */ +StaticAssertDecl(BUF_USAGECOUNT_BITS >= 1, + "cooling state needs at least one bit in the usagecount field"); + /* flags related definitions */ #define BUF_FLAG_SHIFT \ (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS) @@ -86,11 +116,16 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS + BUF_L ((((uint64) MAX_BACKENDS) << BM_LOCK_SHIFT) | BM_LOCK_VAL_SHARE_EXCLUSIVE | BM_LOCK_VAL_EXCLUSIVE) -/* Get refcount and usagecount from buffer state */ +/* Get refcount and cooling state from buffer state */ #define BUF_STATE_GET_REFCOUNT(state) \ ((uint32)((state) & BUF_REFCOUNT_MASK)) -#define BUF_STATE_GET_USAGECOUNT(state) \ - ((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)) + +/* + * Cooling state (HOT/COOL) from buffer state. The field holds only + * BUF_COOLSTATE_COOL or BUF_COOLSTATE_HOT. + */ +#define BUF_STATE_GET_COOLSTATE(state) \ + ((uint32) (((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)) /* * Flags for buffer descriptors @@ -134,17 +169,15 @@ StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2), /* - * The maximum allowed value of usage_count represents a tradeoff between - * accuracy and speed of the clock-sweep buffer management algorithm. A - * large value (comparable to NBuffers) would approximate LRU semantics. - * But it can take as many as BM_MAX_USAGE_COUNT+1 complete cycles of the - * clock-sweep hand to find a free buffer, so in practice we don't want the - * value to be very large. + * The cooling state is a single bit (HOT/COOL); the maximum value stored in + * the field is therefore BUF_COOLSTATE_HOT. Retained under the historical + * name BM_MAX_USAGE_COUNT so the pin fast path ("promote unless already at + * max") reads naturally. */ -#define BM_MAX_USAGE_COUNT 5 +#define BM_MAX_USAGE_COUNT BUF_COOLSTATE_HOT StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS), - "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits"); + "cooling state doesn't fit in BUF_USAGECOUNT_BITS bits"); /* * Buffer tag identifies which disk block the buffer contains. From b18c97f4c8c0e920dab91f7cfe9e9a152c120aa1 Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Sun, 4 Oct 2026 12:58:46 -0400 Subject: [PATCH 13/17] Bound clock-sweep work by claiming a HOT buffer under pressure The cooling stage only produces victims if a buffer demoted to COOL is still COOL when the clock hand comes back to it. With a pass time of T, a page must therefore go roughly T without an access to become reclaimable. T grows with NBuffers, so on a large buffer pool a workload that touches much of the pool more often than once per pass re-promotes buffers faster than the hand demotes them. The population of COOL buffers collapses, and StrategyGetBuffer() can then cool indefinitely without ever finding a victim: cooling counts as progress and resets trycounter, so there is no bound on the work one allocation performs. The same failure mode exists with the 0..5 usage_count, and is worse there, because a page must go ~5T unaccessed rather than ~T. Bound it. A sweep that has demoted BUF_COOL_CLAIM_THRESHOLD buffers without finding a single COOL one has established, by observation rather than by guesswork, that the pool is hotter than the hand can grind down. From that point in the same call it claims the next unpinned HOT buffer directly instead of merely demoting it. The cooling state is cleared when a victim is reused (InvalidateVictimBuffer), so claiming a HOT buffer needs no separate demotion step. The cost is evicting a buffer that had not finished its probation, which is the right trade only when the alternative is an unbounded scan; the threshold is therefore set well above what a healthy workload reaches. Under no pressure the counter never reaches it and behaviour is exactly as before. The threshold is derived from the cache line size rather than being a new tunable. StrategyControl gains numCoolClaims, counting allocations that had to resort to this. It is the signal that the pool is undersized for its access rate, and a later commit uses it to drive background pre-cooling so that the foreground sweep stops reaching this path at all. --- src/backend/storage/buffer/freelist.c | 76 +++++++++++++++++++++++---- src/include/storage/buf_internals.h | 17 ++++++ 2 files changed, 84 insertions(+), 9 deletions(-) diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index 6e41a56c5b977..d88132d724938 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -122,6 +122,17 @@ typedef struct pg_atomic_uint64 numTotalAllocs; pg_atomic_uint64 numTotalRequestedAllocs; + /* + * Allocations in this partition that had to claim a still-HOT buffer + * because the sweep could not find a COOL one (see + * StrategyGetBufferPartition). A non-zero and growing value means this + * partition is hotter than its own clock hand can keep up with, which is + * the condition the background pre-cooling in BgBufferSync() exists to + * prevent. This is per-partition because each partition has its own + * hand: one node can starve while another sits idle. + */ + pg_atomic_uint64 numCoolClaims; + /* * Weights to balance buffer allocations for all the partitions. Each * partition gets a vector of weights 0-100, determining what fraction @@ -646,6 +657,7 @@ StrategyGetBufferPartition(ClockSweep *sweep, BufferAccessStrategy strategy, { BufferDesc *buf; int trycounter; + int cooled = 0; /* HOT buffers this call demoted (the governor) */ trycounter = sweep->numBuffers; for (;;) @@ -692,20 +704,65 @@ StrategyGetBufferPartition(ClockSweep *sweep, BufferAccessStrategy strategy, if (BUF_STATE_GET_COOLSTATE(local_buf_state) != BUF_COOLSTATE_COOL) { /* - * HOT buffer: cool it in place this tick and keep scanning. We - * do NOT claim it now -- a demoted buffer only becomes a victim - * on a later tick, so a HOT buffer always survives the pass that - * cools it and gets a full sweep of grace in which a new access - * can promote it back to HOT. Cooling is progress toward a - * victim, so reset trycounter. + * HOT buffer. Normally we cool it in place and keep + * scanning: a demoted buffer only becomes a victim on a later + * tick, so a HOT buffer survives the pass that cools it and + * gets a full sweep of grace in which a new access can promote + * it back. Cooling is progress toward a victim, so reset + * trycounter. + * + * That grace period is the right default, but it is not free. + * A buffer is only evictable if it is still COOL when the hand + * returns, so with a pass time of T a page must go roughly T + * without an access to be reclaimable. T grows with NBuffers, + * so on a large pool under a workload that touches much of the + * pool more often than that, buffers are re-promoted before + * the hand comes back, the supply of COOL victims collapses, + * and this loop can cool indefinitely without finding one -- + * unbounded work for a single allocation, because cooling + * resets trycounter. + * + * So we give up the grace period when, and only when, the + * evidence says it is unaffordable: once this call has cooled + * BUF_COOL_CLAIM_THRESHOLD buffers without finding a single + * COOL one, the pool is demonstrably hotter than the hand can + * grind down, and we claim the next unpinned HOT buffer + * instead of merely cooling it. Under no pressure the counter + * never reaches the threshold and behaviour is unchanged. + */ + if (cooled < BUF_COOL_CLAIM_THRESHOLD) + { + local_buf_state &= ~BUF_USAGECOUNT_MASK; /* HOT -> COOL */ + + if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state, + local_buf_state)) + { + cooled++; + trycounter = sweep->numBuffers; + break; + } + continue; + } + + /* + * Pressure case: claim this HOT buffer. The cooling state is + * cleared when the victim is reused (InvalidateVictimBuffer), + * so there is no need to demote it first. */ - local_buf_state &= ~BUF_USAGECOUNT_MASK; /* HOT -> COOL */ + pg_atomic_fetch_add_u64(&sweep->numCoolClaims, 1); + local_buf_state += BUF_REFCOUNT_ONE; if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state, local_buf_state)) { - trycounter = sweep->numBuffers; - break; + /* Found a usable buffer */ + if (strategy != NULL) + AddBufferToRing(strategy, buf); + *buf_state = local_buf_state; + + TrackNewBufferPin(BufferDescriptorGetBuffer(buf)); + + return buf; } } else @@ -1174,6 +1231,7 @@ StrategyCtlShmemInit(void *arg) pg_atomic_init_u32(&StrategyControl->sweeps[i].numRequestedAllocs, 0); pg_atomic_init_u64(&StrategyControl->sweeps[i].numTotalAllocs, 0); pg_atomic_init_u64(&StrategyControl->sweeps[i].numTotalRequestedAllocs, 0); + pg_atomic_init_u64(&StrategyControl->sweeps[i].numCoolClaims, 0); /* * Initialize the weights - start by allocating 100% buffers from diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h index 6c53ecc5ed0bd..cb5ee06afdd95 100644 --- a/src/include/storage/buf_internals.h +++ b/src/include/storage/buf_internals.h @@ -89,6 +89,23 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS + BUF_L #define BUF_COOLSTATE_HOT 1 #define BUF_COOLSTATE_ONE BUF_USAGECOUNT_ONE +/* + * How many HOT buffers a single StrategyGetBuffer() call will demote before it + * gives up on the probation rule and claims a still-HOT buffer outright. + * + * The cooling stage only yields victims if a demoted buffer is still COOL when + * the clock hand returns to it. When the pool is accessed faster than the hand + * can traverse it, buffers are re-promoted before that happens, the supply of + * COOL buffers collapses, and a sweep can cool indefinitely without ever + * finding a victim. Claiming a HOT buffer after this many fruitless demotions + * bounds the work of one allocation; the cost is evicting a buffer that had not + * finished its probation, so the threshold wants to be high enough that healthy + * workloads never reach it. One cache line of buffer descriptors' worth of + * demotions is a cheap, hardware-derived choice. + */ +#define BUF_COOL_CLAIM_THRESHOLD \ + (PG_CACHE_LINE_SIZE / sizeof(uint32)) + /* * The cooling state is one bit, so the field must be at least that wide. * Assert it here so a future change to BUF_USAGECOUNT_BITS cannot silently From 7310921419c2d7f29d2780d758e1c661308ed376 Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Sat, 10 Oct 2026 23:46:31 -0400 Subject: [PATCH 14/17] Fix -Wtype-limits in the partitioned ClockSweepTick assert `victim` is uint32, so `victim >= 0` is always true and GCC rejects the comparison under -Werror=type-limits. Keep the half of the assertion that can actually fail: that the computed index is inside the partition. Needed so every commit in the series builds clean under -Werror. --- src/backend/storage/buffer/freelist.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index d88132d724938..a9a8bcc51c143 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -349,7 +349,7 @@ ClockSweepTick(ClockSweep *sweep) * Make sure we've calculated a buffer in the range of the partition. Buffer * IDs are 1-based, we're calculating 0-based indexes. */ - Assert((victim >= 0) && (victim < sweep->numBuffers)); + Assert(victim < sweep->numBuffers); Assert(BufferIsValid(1 + sweep->firstBuffer + victim)); return sweep->firstBuffer + victim; From 16c8995044a652d4929fd184c92ab4b328117250 Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Sun, 4 Oct 2026 13:02:11 -0400 Subject: [PATCH 15/17] Stage eviction candidates in the background writer when the sweep starves The previous commit bounds the work of an allocation that cannot find a COOL buffer, but it pays for the bound by evicting a buffer that had not finished its probation, and the allocating backend still performs the search. Both costs land on the critical path of a query. Move the work off that path. The background writer's cleaning scan already walks buffers ahead of the clock hand and already takes each buffer header lock, so demoting an unpinned HOT buffer as it passes costs one masked store on work being done anyway, and it needs no CAS. A buffer staged this way is a victim the foreground sweep will not have to demote itself. The scan stages candidates only when the foreground is demonstrably starving. StrategyCoolClaims() reports the running count of allocations that had to claim a still-HOT buffer; BgBufferSync() compares it against the previous cycle and enables staging only while that count is rising. This matters: cooling buffers that nobody is waiting for would shorten every buffer's probation for no benefit, which is precisely the mistake of a fixed background cooler. Under a workload the hand can keep up with, the count does not move, staging stays off, and the full probation period is preserved. Staged buffers count toward reusable_buffers, so the scan stops once it has built enough supply for the predicted next-cycle demand. Pre-cooling is thus bounded by the same demand estimate that already bounds cleaning, and cannot run away and cool the entire pool. A concurrent PinBuffer() promotes a staged buffer back to HOT, which is the behaviour we want: a buffer being accessed should not be evicted, and the staging is advisory rather than a commitment. The checkpointer's call site passes cool_if_hot = false; checkpointing is not replacement and has no reason to alter replacement state. --- src/backend/storage/buffer/bufmgr.c | 82 +++++++++++++++++++++++++-- src/backend/storage/buffer/freelist.c | 17 ++++++ src/include/storage/buf_internals.h | 1 + 3 files changed, 96 insertions(+), 4 deletions(-) diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c index 36f2d48cb912f..f2057cef37928 100644 --- a/src/backend/storage/buffer/bufmgr.c +++ b/src/backend/storage/buffer/bufmgr.c @@ -83,6 +83,7 @@ /* Bits in SyncOneBuffer's return value */ #define BUF_WRITTEN 0x01 #define BUF_REUSABLE 0x02 +#define BUF_COOLED 0x04 #define RELS_BSEARCH_THRESHOLD 20 @@ -636,7 +637,7 @@ static void PinBuffer_Locked(BufferDesc *buf); static void UnpinBuffer(BufferDesc *buf); static void UnpinBufferNoOwner(BufferDesc *buf); static void BufferSync(int flags); -static int SyncOneBuffer(int buf_id, bool skip_recently_used, +static int SyncOneBuffer(int buf_id, bool skip_recently_used, bool cool_if_hot, WritebackContext *wb_context); static void WaitIO(BufferDesc *buf); static void AbortBufferIO(Buffer buffer); @@ -3799,7 +3800,7 @@ BufferSync(int flags) */ if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED) { - if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN) + if (SyncOneBuffer(buf_id, false, false, &wb_context) & BUF_WRITTEN) { TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id); PendingCheckpointerStats.buffers_written++; @@ -3904,6 +3905,15 @@ BgBufferSyncPartition(WritebackContext *wb_context, int num_partitions, int num_to_scan; int num_written; int reusable_buffers; + bool cool_if_hot; + uint64 cool_claims; + + /* + * Cool-claim count as of our previous cycle, to get a rate not a total. + * Per-partition: each partition has its own hand and can starve + * independently, so staging is decided for each one separately. + */ + static uint64 *prev_cool_claims = NULL; /* Variables for final smoothed_density update */ long new_strategy_delta; @@ -4077,11 +4087,42 @@ BgBufferSyncPartition(WritebackContext *wb_context, int num_partitions, num_written = 0; reusable_buffers = reusable_buffers_est; + /* + * Decide whether to stage eviction candidates as we go. + * + * The foreground sweep only finds a victim if some buffer is COOL when the + * clock hand reaches it. StrategyGetBuffer() reports, via + * StrategyCoolClaims(), how many allocations had to claim a still-HOT + * buffer because no COOL one could be found. A non-zero rate means the + * pool is being re-promoted faster than the hand demotes it, and that + * backends are paying for the search on the critical path. + * + * When that happens we demote HOT buffers as the cleaning scan passes + * them. The scan is already walking these buffers and already holds each + * header lock, so staging costs a masked store on work we are doing + * anyway, and it moves the demotion off the allocating backend. When the + * foreground is not starving we do not touch the cooling state at all, so + * a healthy workload keeps the full probation period and is unaffected. + */ + if (prev_cool_claims == NULL) + { + /* Same reason as saved_info below: a variable-length static. */ + prev_cool_claims = calloc(num_partitions, sizeof(uint64)); + if (prev_cool_claims == NULL) + ereport(ERROR, + (errcode(ERRCODE_OUT_OF_MEMORY), + errmsg("out of memory"))); + } + + cool_claims = StrategyCoolClaims(partition); + cool_if_hot = (cool_claims > prev_cool_claims[partition]); + prev_cool_claims[partition] = cool_claims; + /* Execute the LRU scan */ while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est) { int sync_state = SyncOneBuffer(saved->next_to_clean, true, - wb_context); + cool_if_hot, wb_context); if (++saved->next_to_clean >= (first_buffer + num_buffers)) { @@ -4101,6 +4142,15 @@ BgBufferSyncPartition(WritebackContext *wb_context, int num_partitions, } else if (sync_state & BUF_REUSABLE) reusable_buffers++; + + /* + * A buffer we demoted is a victim the sweep will not have to demote + * itself. Count it toward the supply we are building so the scan + * stops once it has staged enough for the predicted demand, rather + * than cooling the whole pool. + */ + if (sync_state & BUF_COOLED) + reusable_buffers++; } PendingBgWriterStats.buf_written_clean += num_written; @@ -4225,16 +4275,25 @@ BgBufferSync(WritebackContext *wb_context) * If skip_recently_used is true, we don't write currently-pinned buffers, nor * buffers marked recently used, as these are not replacement candidates. * + * If cool_if_hot is true, an unpinned HOT buffer is demoted to COOL as we pass + * it, staging it as an eviction candidate so the foreground clock sweep finds a + * victim without having to demote buffers itself. We hold the buffer header + * lock, so the demotion is a plain masked store rather than a CAS, and a + * concurrent PinBuffer() simply promotes the buffer back to HOT -- the intended + * behaviour, since a buffer that is being accessed should not be evicted. + * * Returns a bitmask containing the following flag bits: * BUF_WRITTEN: we wrote the buffer. * BUF_REUSABLE: buffer is available for replacement, ie, it has * pin count 0 and is COOL (an eviction candidate). + * BUF_COOLED: we demoted this buffer from HOT to COOL. * * (BUF_WRITTEN could be set in error if FlushBuffer finds the buffer clean * after locking it, but we don't care all that much.) */ static int -SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context) +SyncOneBuffer(int buf_id, bool skip_recently_used, bool cool_if_hot, + WritebackContext *wb_context) { BufferDesc *bufHdr = GetBufferDescriptor(buf_id); int result = 0; @@ -4256,6 +4315,21 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context) */ buf_state = LockBufHdr(bufHdr); + /* + * Stage an eviction candidate if asked. An unpinned HOT buffer is demoted + * to COOL here so the foreground sweep can claim it on its next visit. We + * already hold the header lock, so this costs one masked store and no + * atomic retry. + */ + if (cool_if_hot && + BUF_STATE_GET_REFCOUNT(buf_state) == 0 && + BUF_STATE_GET_COOLSTATE(buf_state) != BUF_COOLSTATE_COOL) + { + UnlockBufHdrExt(bufHdr, buf_state, 0, BUF_USAGECOUNT_MASK, 0); + buf_state = LockBufHdr(bufHdr); + result |= BUF_COOLED; + } + if (BUF_STATE_GET_REFCOUNT(buf_state) == 0 && BUF_STATE_GET_COOLSTATE(buf_state) == BUF_COOLSTATE_COOL) { diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c index a9a8bcc51c143..9e0cae44ff266 100644 --- a/src/backend/storage/buffer/freelist.c +++ b/src/backend/storage/buffer/freelist.c @@ -1145,6 +1145,23 @@ StrategySyncStart(int partition, uint32 *complete_passes, return sweep->firstBuffer + result; } +/* + * StrategyCoolClaims -- how many allocations had to claim a HOT buffer + * + * Returns a monotonically increasing count of allocations that could not find + * a COOL victim and claimed a still-HOT buffer instead. The background writer + * compares successive readings: a rising count means the foreground sweep is + * starving for eviction candidates, which is the condition under which staging + * candidates in the background pays for itself. + */ +uint64 +StrategyCoolClaims(int partition) +{ + Assert((partition >= 0) && (partition < StrategyControl->num_partitions)); + + return pg_atomic_read_u64(&StrategyControl->sweeps[partition].numCoolClaims); +} + /* * StrategyNotifyBgWriter -- set or clear allocation notification latch * diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h index cb5ee06afdd95..db77730b3020d 100644 --- a/src/include/storage/buf_internals.h +++ b/src/include/storage/buf_internals.h @@ -654,6 +654,7 @@ extern void StrategySyncPrepare(int *num_parts, uint32 *num_buf_alloc); extern int StrategySyncStart(int partition, uint32 *complete_passes, int *first_buffer, int *num_buffers); extern void StrategyNotifyBgWriter(int bgwprocno); +extern uint64 StrategyCoolClaims(int partition); /* buf_table.c */ extern uint32 BufTableHashCode(BufferTag *tagPtr); From 431f9b60d06f4b0390256546c417b26f9618f2d0 Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Sun, 11 Oct 2026 01:13:09 -0400 Subject: [PATCH 16/17] Use the measured claim threshold of 128 The cache-line-derived value (32) was a placeholder chosen before the trade-off was measured. The threshold sensitivity curve and the premature-eviction harm measurement both land on 128: it bounds the tail at threshold+2 advances while displacing no hot pages and costing no measurable hit ratio, where 32 costs 0.03 points and 4 costs 1.8 points. --- src/include/storage/buf_internals.h | 15 +++++++++++---- 1 file changed, 11 insertions(+), 4 deletions(-) diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h index db77730b3020d..557afd14d5236 100644 --- a/src/include/storage/buf_internals.h +++ b/src/include/storage/buf_internals.h @@ -100,11 +100,18 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS + BUF_L * finding a victim. Claiming a HOT buffer after this many fruitless demotions * bounds the work of one allocation; the cost is evicting a buffer that had not * finished its probation, so the threshold wants to be high enough that healthy - * workloads never reach it. One cache line of buffer descriptors' worth of - * demotions is a cheap, hardware-derived choice. + * workloads never reach it. + * + * Both sides have been measured. The worst-case advances an allocation + * performs is almost exactly this threshold plus two, so the tail is a direct + * function of it. The cost is premature eviction: on a workload whose hot set + * most queries re-read, a threshold of 4 leaves 4,519 hot pages non-resident + * and costs 1.8 points of hot-set hit ratio, 32 leaves 11 pages and costs 0.03 + * points, and at 128 nothing hot is displaced and the cost is not measurable. + * 128 therefore buys a bound four orders of magnitude below an unbounded scan + * for no measurable loss of replacement quality. */ -#define BUF_COOL_CLAIM_THRESHOLD \ - (PG_CACHE_LINE_SIZE / sizeof(uint32)) +#define BUF_COOL_CLAIM_THRESHOLD 128 /* * The cooling state is one bit, so the field must be at least that wide. From 467c3472fc6d562e32371b053a35588a6af05a91 Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Sun, 11 Oct 2026 05:26:11 -0400 Subject: [PATCH 17/17] doc: four-arm measurement of the series on the NUMA base Measured on bare metal (z1d.metal, Fedora 44, XFS on RAID-0 instance-local NVMe at 429k random read IOPS), four arms isolating each change: pristine master, + the Vondra/Wartak NUMA partitioning, + the 1-bit HOT/COOL evictor, and + the governor and bgwriter cooling stage. The two mechanisms separate cleanly. The 1-bit replacement state cuts amplification 1.66-1.86x (advances per victim 7.46 -> 4.01 at 4GiB, 6.99 -> 4.20 at 32GiB) and does nothing for the tail. The governor collapses the tail 44-189x and, unlike every other arm, holds it flat in pool size: 424, 371 and 297 advances at 4, 8 and 32GiB. Throughput moves +0.03 to +0.32% and hit ratio by at most 0.03 points, better at 32GiB. Worth reporting upstream: partitioning on its own makes the tail 6.7-8.7x WORSE than unpartitioned master, because each backend sweeps only its own quarter of the pool and can starve while other partitions hold candidates. The governor removes that regression and lands well below unpartitioned stock. Caveat recorded in .agent/kiwi-2026-10-11/RESULTS.md: this run's pool fill was 87-95% rather than 100%, so its *stock* tails are much lower than the earlier O(NBuffers) measurement and must not be read as contradicting it. All four arms shared the same regime here, so the comparison between them holds. --- paper/data_fourarm_numa.dat | 13 +++++++++++++ 1 file changed, 13 insertions(+) create mode 100644 paper/data_fourarm_numa.dat diff --git a/paper/data_fourarm_numa.dat b/paper/data_fourarm_numa.dat new file mode 100644 index 0000000000000..ae2c51e6a6db1 --- /dev/null +++ b/paper/data_fourarm_numa.dat @@ -0,0 +1,13 @@ +# pool nbuffers arm tps ampl tail tail_over_n claims hit_pct runs +4GB 524288 stock 664342 7.560 6099 0.011633 0 98.9993 3 +4GB 524288 numa 659593 7.461 43334 0.082653 0 98.9957 3 +4GB 524288 hotcool 661096 4.012 35512 0.067734 0 98.9682 3 +4GB 524288 full 659792 3.971 424 0.000809 283 98.9677 3 +8GB 1048576 stock 664504 7.458 5168 0.004929 0 99.0667 3 +8GB 1048576 numa 660483 7.144 45072 0.042984 0 99.0583 3 +8GB 1048576 hotcool 657902 4.069 16664 0.015892 0 99.0451 3 +8GB 1048576 full 662584 4.017 371 0.000354 223 99.0456 3 +32GB 4194304 stock 663698 6.117 13654 0.003255 0 99.1402 6 +32GB 4194304 numa 657391 6.986 90986 0.021693 0 99.0542 3 +32GB 4194304 hotcool 657270 4.200 56174 0.013393 0 99.1533 3 +32GB 4194304 full 659361 4.403 297 0.000071 299 99.1483 3