Repository navigation
Segfault: postgres session + SET ROLE authenticated + call to a function with EXECUTE revoked (17.6.1.104 / 17.6.1.106) #2377
Description
Activity
No repro environment on my end, so this is static analysis of the image, not a reproduction -- but it may narrow the bisect.
Upstream differential: vanilla PostgreSQL handles this path cleanly and has for a very long time. Function-EXECUTE denial goes
exec_check_permissions->pg_proc_aclcheck->aclcheck_error-> clean 42501 ereport; there is no plausible upstream segfault there, which matches your PostgREST-path observation (same target role, same denial, clean error). So the crash almost certainly lives in one of the image's hooks rather than in core's ACL machinery.One thing your preload list omits: the image also sets
session_preload_libraries = 'supautils'(seedocker/pgctld/postgresql.conf.tmplin this repo). supautils hooks utility processing specifically to constrain what the privilegedpostgresrole can do -- privilege escalation guards, SET ROLE / SET SESSION AUTHORIZATION interception, object-owner restrictions. That is exactly the code path where your two cases diverge: a session that started aspostgrestakes supautils' privileged-session branch when switching toauthenticated, while theauthenticatorpath never enters it. An interaction between that hook branch and the subsequent function-ACL denial (or between it and pgaudit's object-access logging of the same denial) is the shape of bug I would expect to produce signal 11 only in the postgres-originated case.Suggested bisect order, cheapest discrimination first:
# 1. prime suspect: privileged-session switch guard session_preload_libraries = '' # if still crashing: # 2. add shared libs back without pgaudit (object-access hook sits # directly in the denial path) shared_preload_libraries = 'pg_stat_statements, plpgsql, plpgsql_check, pg_cron, pg_net, auto_explain, pg_tle, plan_filter, supabase_vault' # 3. then restore pgaudit alone with everything else off, etc.
If step 1 alone stops the crash, the report belongs partly upstream in supautils (
supabase/supautils) as well as here, since the defect would be in its SET ROLE handling for privileged sessions rather than in core or the packaging.One cross-reference worth noting since it affects how often users will hit this: the "revoke EXECUTE from authenticated" hardening pattern is currently taught by an example in
supabase/agent-skills(security-rls-performance.md, tracked in supabase/agent-skills#390, fix PR #409 open) -- agents following it against local stacks built from these images will hit whatever happens here instead of the expected 42501, so the guidance fix and this crash are likely to surface together.Reproduced independently on
public.ecr.aws/supabase/postgres:17.6.1.104(aarch64,PostgreSQL 17.6 ... gcc 15.2.0), provisioned bysupabase startwith CLI 2.117.0 and a linked project on the same version.Same shape as the report, with the
anonrole as well asauthenticated:create function public.f() returns int language sql as 'select 1'; revoke all on function public.f() from public, anon; begin; set local role anon; select public.f(); rollback; -- server closed the connection unexpectedly -- LOG: server process (PID 2013) was terminated by signal 11: Segmentation fault -- DETAIL: Failed process was running: SELECT public.f();
Also observed here: table-level permission denials and ordinary errors (
select 1/0) under the sameset local roleraise normally;set local pgaudit.log = 'none',plan_filter.statement_cost_limit = 0andauto_explain.log_min_duration = -1before the role switch do not change the outcome.Cross-references that may be useful for whoever triages this: supabase/cli#6094 reports the same crash on
17.6.1.111and a cleanpermission denied for functionon17.6.1.121, citing supabase/supautils#196, supabase/supautils#200 and supautils PR #190 (v3.2.2).Thanks for the independent aarch64 reproduction and the additional
anoncase. That rules out a report-specific role or architecture quirk and makes the version boundary insupabase/supautils#196/#200and PR #190 especially useful for triage. The clean result on17.6.1.121versus the crash on17.6.1.111should give maintainers a concrete regression window; I agree this belongs in the image/supautils release path rather than in the application’s RLS policies.Thanks for the independent reproduction and the version cross-reference. The clean behavior reported on 17.6.1.121 / supautils 3.2.2 makes this look like the regression is fixed in the newer image line. For triage, it would be useful to record the exact bundled supautils version and whether the same minimal
SET LOCAL ROLE+ revoked function test passes on 17.6.1.121, including bothauthenticatedandanon. If confirmed, this issue can point users to the fixed image while the older tags remain affected.Still reproduces on
supabase/postgres:17.6.1.111(newer than the.104/.106in the report), on a local CLI stack (x86_64). A hosted project on the same patch level runsaarch64; I did not test the hosted one, for reasons that will be obvious below.One correction to the original report, and it changes the severity
The report states it does not reproduce through PostgREST. That is not true for the
anonrole on.111.Setup, a trivial function with no attributes at all:
create or replace function public.repro_rest(x int) returns int language sql as $$ select x $$; revoke all on function public.repro_rest(int) from public; revoke execute on function public.repro_rest(int) from anon; notify pgrst, 'reload schema';
Three anonymous HTTP calls with the project's publishable (
anon) key:POST /rest/v1/rpc/repro_rest -> 503 {"code":"PGRST001","details":"no connection to the server\n",...} POST /rest/v1/rpc/repro_rest -> 503 {"code":"PGRST001","details":"no connection to the server\n",...} POST /rest/v1/rpc/repro_rest -> 503 {"code":"PGRST001","details":"no connection to the server\n",...}Segfault count over that window: before 16, after 19, delta +3. One crash per call. The server log names PostgREST's own wrapper query as the crashing process:
LOG: server process (PID ...) was terminated by signal 11: Segmentation fault DETAIL: Failed process was running: WITH pgrst_source AS (SELECT pgrst_call.pgrst_scalar FROM (SELECT $1 AS json_data) pgrst_payload, LATERAL (SELECT "x" FROM json_to_record(pgrst_payload.json_data) AS _("x" integer) LIMIT 1) pgrst_body , LATERAL (SELECT "public"."repro_rest"("x" := pgrst_body."x") pgrst_scalar) pgrst_call) SELECT null::bigint AS total_result_set, ... LOG: terminating any other active server processes LOG: all server processes terminated; reinitializingI tested
anononly, notauthenticated, over HTTP. It is possible the original observation holds forauthenticatedand not foranon.Why this matters for severity
A crash reachable only from a privileged
postgressession is a developer annoyance. This one is reachable by anyone, with no account: theanon/ publishable key is public by design and ships inside the site's JavaScript bundle. Anypublic-schema function withEXECUTErevoked fromanontherefore becomes a remote, unauthenticated denial of service against the whole cluster, in a loop, withcurl.On our project four functions were in exactly that state, three of them for several weeks, without anyone noticing: revoking
EXECUTEis the natural, documented way to lock an RPC down, so the safest-looking configuration was the vulnerable one. No incident occurred (24 h of production logs show no crash and no function permission error), but the exposure was real.Attributes are irrelevant
All of these crash identically under
set local role anon:function definition result language sql, no attributessegfault language plpgsql, no attributessegfault security definersegfault security invoker set search_path = ''segfault security definer set search_path = ''segfault security definer set search_path = pg_catalogsegfault Controls that hold
So the failure really is specific to function-
EXECUTEdenial, and not to permission denial in general:- a permission denial on a table returns cleanly, no crash;
- a function the role is allowed to execute runs normally;
- after the crash, a fresh connection works again (the cluster recovers in about a second, dropping every other connection with it).
Workaround, in case it helps others
The crash is on the privilege-denial path, so we removed the denial rather than the function:
- where the function is genuinely called over HTTP by an authenticated admin, we granted
EXECUTEback toanonand let the existing in-body guard do the refusing. It was already the first statement, so the authorization outcome is unchanged; only the failure mode moves from "cluster restarts" to a clean401/42501; - where nothing calls the function over HTTP, we moved it out of the exposed schema entirely. Strictly safer for functions that write.
We verified this by behaviour, not by reading
has_function_privilege: the privilege check was correct all along and protected nothing, because the server died before it ever got to the refusal.We hit the same crash on 17.6.1.104 and bisected it to supautils' hint_roles (supautils.hint_roles = 'anon, authenticated, service_role', loaded via session_preload_libraries).
- Same repro as yours (postgres session → SET ROLE anon|authenticated|service_role → call a function whose EXECUTE is revoked → signal 11, all backends reinitialized).
- With supautils not loaded, or with PGOPTIONS='-c supautils.hint_roles=' (empty), the same call returns a clean 42501 permission denied for function, no crash.
- Denials on tables, views or schemas don't crash; only function EXECUTE denials do, which points at the hint hook formatting the GRANT … ON FUNCTION … TO … hint.
- PostgREST is unaffected because authenticator has session_preload_libraries=safeupdate (supautils not loaded).
This should be resolved in service versions and images >= 17.6.1.143, supautils 3.2.2
See:Reproduced on 17.6.1.104 aarch64 (Apple Silicon, Docker Desktop, Rosetta off): any function the role lacks EXECUTE on, with either anon or authenticated, crashes the backend; a table-level denial in the same session returns 42501 normally. Two hosted projects on PostgreSQL 17.6 (aarch64 and x86_64) return 42501 cleanly for the identical statement.
Edited. My original comment attributed this to pgAudit's
ExecutorCheckPerms_hookand restated findings that were already in the thread. Both were wrong on my part: @kchebani had already bisected it to supautilshint_roles(loaded viasession_preload_libraries), and theanoncase and the table-vs-function asymmetry were reported earlier by @EmireHaouas and @MRecinosOrg. I had checkedshared_preload_libraries, seen no supautils there, and concluded from the wrong parameter.Leaving only the part that does not seem to be in the thread yet:
The crash takes down every open connection, not just the caller
LOG: server process (PID 74894) was terminated by signal 11: Segmentation fault DETAIL: Failed process was running: SELECT public.f_crash_repro() LOG: terminating any other active server processes LOG: all server processes terminated; reinitializing LOG: database system was not properly shut down; automatic recovery in progress LOG: database system is ready to accept connectionsRecovery is fast (~690 ms), so on an interactive workload it looks like a blip. On a test suite it is much worse than it appears: a single test asserting "role X cannot execute function Y" — the natural way to cover a
REVOKE EXECUTEpolicy — kills every connection in flight.Concretely, in one CI run of our RLS/isolation suites against
17.6.1.104: 58 test files lost, each failing withthe database system is not accepting connections, none of them having anything to do with privileges. They were simply connecting during the recovery window. We had to teach our gate to recognize the crash signature and drop those files from the verdict, otherwise the noise is attributed to whoever opened the pull request.Worth flagging for anyone landing here from a CI failure: the symptom you see is almost never at the test that caused it.
Confirmed on
x86_64, imagesupabase/postgres:17.6.1.104. Thanks @kostasb for the pointer to >=17.6.1.143/ supautils 3.2.2 — upgrading the image is the path for us.Confirmed fixed on
17.6.1.178@kostasb's pointer to >=
17.6.1.143holds up in practice — verifying here since the thread so far only had the expectation, not a test.Upgraded the image from
17.6.1.104to17.6.1.178on an existing volume (same major, no re-init needed) and re-ran the exact repro:EXECUTE -> anon=false authenticated=false (the bug's precondition) anon: [42501] permission denied for function f_crash_repro authenticated: [42501] permission denied for function f_crash_repro server still up: YESClean
42501on both roles, no segfault, no recovery cycle. Before the upgrade, the same call on the same database killed the backend every single time.Worth noting for anyone who, like me, went looking in the wrong place:
pgauditis still loaded inshared_preload_librariesafter the upgrade and nothing crashes. What changed issupautils, which lives insession_preload_libraries— a different parameter. I checked only the former early on and built a whole (wrong) hypothesis around pgAudit before @kchebani's bisect set it straight.session_preload_libraries: supautils shared_preload_libraries: pg_stat_statements, pgaudit, plpgsql, plpgsql_check, pg_cron, pg_net, pgsodium, auto_explain, pg_tle, plan_filter, supabase_vaultMeasured impact of the fix, on a real workload
Our RLS/isolation test suites, same hardware, same database, before and after:
17.6.1.10417.6.1.178test files lost to server crashes 34–58 per run (varying) 0 passing cases 1,753 1,789 The variability is the part that hurt most: the set of casualties was decided by which files happened to be connecting during the ~690 ms recovery window, so the same commit produced different results run to run. That is what kept us from making the suite a required check.
Thanks to everyone who reproduced and bisected this — it saved us from chasing the wrong library.
Summary
On
supabase/postgres:17.6.1.104and17.6.1.106, a session connected as thepostgresrole that doesSET ROLE authenticatedand then calls any function on which EXECUTE has been revoked fromauthenticatedcrashes the backend with signal 11 (Segmentation fault) instead of returningpermission denied for function. The whole cluster goes into recovery for ~1 s.It reproduces with a trivial
select 1function,LANGUAGE sqlorplpgsql, scalar or set-returning, inside or outside aDOblock, withSET ROLEorSET LOCAL ROLE.It does not reproduce through the PostgREST path (
authenticator→SET ROLE authenticated→ call): that returns a cleanERROR: permission denied for function f. Permission errors on tables (RLS 42501, revoked TRUNCATE) do not crash either — only the function-EXECUTE denial from apostgressession does.Minimal repro (bare image, no user data)
Same result with
supabase/postgres:17.6.1.106(the imagesupabase startpins with CLI 2.90.0). Also reproduced inside a fullsupabase startstack.show shared_preload_librariesin the image:pg_stat_statements, pgaudit, plpgsql, plpgsql_check, pg_cron, pg_net, pgsodium, auto_explain, pg_tle, plan_filter, supabase_vault.Expected
ERROR: permission denied for function f(SQLSTATE 42501), as it happens viaauthenticator.Why it matters
The
postgresrole is the one used by migration runners, CI test harnesses (psqlas postgres +SET LOCAL ROLE authenticatedto exercise RLS/grants) and the Supabase MCPexecute_sql. A test that legitimately checks "authenticated cannot execute this function" takes the database down — in CI it kills the job; against a hosted project it would restart the primary. We hit it in CI first, then confirmed on both images. Postgres logs on our hosted project (same 17.6 line) show no segfaults, so the PostgREST/client path looks unaffected; we are only reporting the privileged-session path.Happy to provide a core dump / more details if useful.