Skip to content

Segfault: postgres session + SET ROLE authenticated + call to a function with EXECUTE revoked (17.6.1.104 / 17.6.1.106) #2377

Description

@Anthony-atic

Summary

On supabase/postgres:17.6.1.104 and 17.6.1.106, a session connected as the postgres role that does SET ROLE authenticated and then calls any function on which EXECUTE has been revoked from authenticated crashes the backend with signal 11 (Segmentation fault) instead of returning permission denied for function. The whole cluster goes into recovery for ~1 s.

It reproduces with a trivial select 1 function, LANGUAGE sql or plpgsql, scalar or set-returning, inside or outside a DO block, with SET ROLE or SET LOCAL ROLE.

It does not reproduce through the PostgREST path (authenticator → SET ROLE authenticated → call): that returns a clean ERROR: permission denied for function f. Permission errors on tables (RLS 42501, revoked TRUNCATE) do not crash either — only the function-EXECUTE denial from a postgres session does.

Minimal repro (bare image, no user data)

docker run --rm -d --name pgrepro -e POSTGRES_PASSWORD=postgres -p 54334:5432 supabase/postgres:17.6.1.104
# wait for pg_isready
docker exec pgrepro psql -U postgres -d postgres -c "
  create function public.f() returns int language sql as \$f\$ select 1 \$f\$;
  revoke execute on function public.f() from public, authenticated;"
docker exec pgrepro psql -U postgres -d postgres -c "set role authenticated; select public.f();"
#   server closed the connection unexpectedly
docker logs pgrepro 2>&1 | grep 'terminated by signal'
#   LOG:  server process (PID 238) was terminated by signal 11: Segmentation fault

Same result with supabase/postgres:17.6.1.106 (the image supabase start pins with CLI 2.90.0). Also reproduced inside a full supabase start stack.

show shared_preload_libraries in the image: pg_stat_statements, pgaudit, plpgsql, plpgsql_check, pg_cron, pg_net, pgsodium, auto_explain, pg_tle, plan_filter, supabase_vault.

Expected

ERROR: permission denied for function f (SQLSTATE 42501), as it happens via authenticator.

Why it matters

The postgres role is the one used by migration runners, CI test harnesses (psql as postgres + SET LOCAL ROLE authenticated to exercise RLS/grants) and the Supabase MCP execute_sql. A test that legitimately checks "authenticated cannot execute this function" takes the database down — in CI it kills the job; against a hosted project it would restart the primary. We hit it in CI first, then confirmed on both images. Postgres logs on our hosted project (same 17.6 line) show no segfaults, so the PostgREST/client path looks unaffected; we are only reporting the privileged-session path.

Happy to provide a core dump / more details if useful.

Activity

  1. cekuu35 commented on Aug 22, 2026

    @cekuu35

    No repro environment on my end, so this is static analysis of the image, not a reproduction -- but it may narrow the bisect.

    Upstream differential: vanilla PostgreSQL handles this path cleanly and has for a very long time. Function-EXECUTE denial goes exec_check_permissions -> pg_proc_aclcheck -> aclcheck_error -> clean 42501 ereport; there is no plausible upstream segfault there, which matches your PostgREST-path observation (same target role, same denial, clean error). So the crash almost certainly lives in one of the image's hooks rather than in core's ACL machinery.

    One thing your preload list omits: the image also sets session_preload_libraries = 'supautils' (see docker/pgctld/postgresql.conf.tmpl in this repo). supautils hooks utility processing specifically to constrain what the privileged postgres role can do -- privilege escalation guards, SET ROLE / SET SESSION AUTHORIZATION interception, object-owner restrictions. That is exactly the code path where your two cases diverge: a session that started as postgres takes supautils' privileged-session branch when switching to authenticated, while the authenticator path never enters it. An interaction between that hook branch and the subsequent function-ACL denial (or between it and pgaudit's object-access logging of the same denial) is the shape of bug I would expect to produce signal 11 only in the postgres-originated case.

    Suggested bisect order, cheapest discrimination first:

    # 1. prime suspect: privileged-session switch guard
    session_preload_libraries = ''
    
    # if still crashing:
    # 2. add shared libs back without pgaudit (object-access hook sits
    #    directly in the denial path)
    shared_preload_libraries = 'pg_stat_statements, plpgsql, plpgsql_check, pg_cron, pg_net, auto_explain, pg_tle, plan_filter, supabase_vault'
    
    # 3. then restore pgaudit alone with everything else off, etc.

    If step 1 alone stops the crash, the report belongs partly upstream in supautils (supabase/supautils) as well as here, since the defect would be in its SET ROLE handling for privileged sessions rather than in core or the packaging.

    One cross-reference worth noting since it affects how often users will hit this: the "revoke EXECUTE from authenticated" hardening pattern is currently taught by an example in supabase/agent-skills (security-rls-performance.md, tracked in supabase/agent-skills#390, fix PR #409 open) -- agents following it against local stacks built from these images will hit whatever happens here instead of the expected 42501, so the guidance fix and this crash are likely to surface together.

  2. EmireHaouas commented on Sep 16, 2026

    @EmireHaouas

    Reproduced independently on public.ecr.aws/supabase/postgres:17.6.1.104 (aarch64, PostgreSQL 17.6 ... gcc 15.2.0), provisioned by supabase start with CLI 2.117.0 and a linked project on the same version.

    Same shape as the report, with the anon role as well as authenticated:

    create function public.f() returns int language sql as 'select 1';
    revoke all on function public.f() from public, anon;
    begin; set local role anon; select public.f(); rollback;
    -- server closed the connection unexpectedly
    -- LOG:  server process (PID 2013) was terminated by signal 11: Segmentation fault
    -- DETAIL:  Failed process was running: SELECT public.f();

    Also observed here: table-level permission denials and ordinary errors (select 1/0) under the same set local role raise normally; set local pgaudit.log = 'none', plan_filter.statement_cost_limit = 0 and auto_explain.log_min_duration = -1 before the role switch do not change the outcome.

    Cross-references that may be useful for whoever triages this: supabase/cli#6094 reports the same crash on 17.6.1.111 and a clean permission denied for function on 17.6.1.121, citing supabase/supautils#196, supabase/supautils#200 and supautils PR #190 (v3.2.2).

  3. cekuu35 commented on Sep 16, 2026

    @cekuu35

    Thanks for the independent aarch64 reproduction and the additional anon case. That rules out a report-specific role or architecture quirk and makes the version boundary in supabase/supautils#196/#200 and PR #190 especially useful for triage. The clean result on 17.6.1.121 versus the crash on 17.6.1.111 should give maintainers a concrete regression window; I agree this belongs in the image/supautils release path rather than in the application’s RLS policies.

  4. cekuu35 commented on Sep 16, 2026

    @cekuu35

    Thanks for the independent reproduction and the version cross-reference. The clean behavior reported on 17.6.1.121 / supautils 3.2.2 makes this look like the regression is fixed in the newer image line. For triage, it would be useful to record the exact bundled supautils version and whether the same minimal SET LOCAL ROLE + revoked function test passes on 17.6.1.121, including both authenticated and anon. If confirmed, this issue can point users to the fixed image while the older tags remain affected.

  5. stormz85ia commented on Sep 18, 2026

    @stormz85ia

    Still reproduces on supabase/postgres:17.6.1.111 (newer than the .104 / .106 in the report), on a local CLI stack (x86_64). A hosted project on the same patch level runs aarch64; I did not test the hosted one, for reasons that will be obvious below.

    One correction to the original report, and it changes the severity

    The report states it does not reproduce through PostgREST. That is not true for the anon role on .111.

    Setup, a trivial function with no attributes at all:

    create or replace function public.repro_rest(x int) returns int language sql as $$ select x $$;
    revoke all     on function public.repro_rest(int) from public;
    revoke execute on function public.repro_rest(int) from anon;
    notify pgrst, 'reload schema';

    Three anonymous HTTP calls with the project's publishable (anon) key:

    POST /rest/v1/rpc/repro_rest -> 503 {"code":"PGRST001","details":"no connection to the server\n",...}
    POST /rest/v1/rpc/repro_rest -> 503 {"code":"PGRST001","details":"no connection to the server\n",...}
    POST /rest/v1/rpc/repro_rest -> 503 {"code":"PGRST001","details":"no connection to the server\n",...}
    

    Segfault count over that window: before 16, after 19, delta +3. One crash per call. The server log names PostgREST's own wrapper query as the crashing process:

    LOG:  server process (PID ...) was terminated by signal 11: Segmentation fault
    DETAIL:  Failed process was running: WITH pgrst_source AS (SELECT pgrst_call.pgrst_scalar
      FROM (SELECT $1 AS json_data) pgrst_payload, LATERAL (SELECT "x" FROM
      json_to_record(pgrst_payload.json_data) AS _("x" integer) LIMIT 1) pgrst_body ,
      LATERAL (SELECT "public"."repro_rest"("x" := pgrst_body."x") pgrst_scalar) pgrst_call)
      SELECT null::bigint AS total_result_set, ...
    LOG:  terminating any other active server processes
    LOG:  all server processes terminated; reinitializing
    

    I tested anon only, not authenticated, over HTTP. It is possible the original observation holds for authenticated and not for anon.

    Why this matters for severity

    A crash reachable only from a privileged postgres session is a developer annoyance. This one is reachable by anyone, with no account: the anon / publishable key is public by design and ships inside the site's JavaScript bundle. Any public-schema function with EXECUTE revoked from anon therefore becomes a remote, unauthenticated denial of service against the whole cluster, in a loop, with curl.

    On our project four functions were in exactly that state, three of them for several weeks, without anyone noticing: revoking EXECUTE is the natural, documented way to lock an RPC down, so the safest-looking configuration was the vulnerable one. No incident occurred (24 h of production logs show no crash and no function permission error), but the exposure was real.

    Attributes are irrelevant

    All of these crash identically under set local role anon:

    function definition result
    language sql, no attributes segfault
    language plpgsql, no attributes segfault
    security definer segfault
    security invoker set search_path = '' segfault
    security definer set search_path = '' segfault
    security definer set search_path = pg_catalog segfault

    Controls that hold

    So the failure really is specific to function-EXECUTE denial, and not to permission denial in general:

    • a permission denial on a table returns cleanly, no crash;
    • a function the role is allowed to execute runs normally;
    • after the crash, a fresh connection works again (the cluster recovers in about a second, dropping every other connection with it).

    Workaround, in case it helps others

    The crash is on the privilege-denial path, so we removed the denial rather than the function:

    • where the function is genuinely called over HTTP by an authenticated admin, we granted EXECUTE back to anon and let the existing in-body guard do the refusing. It was already the first statement, so the authorization outcome is unchanged; only the failure mode moves from "cluster restarts" to a clean 401 / 42501;
    • where nothing calls the function over HTTP, we moved it out of the exposed schema entirely. Strictly safer for functions that write.

    We verified this by behaviour, not by reading has_function_privilege: the privilege check was correct all along and protected nothing, because the server died before it ever got to the refusal.

  6. kchebani commented on Sep 23, 2026

    @kchebani

    We hit the same crash on 17.6.1.104 and bisected it to supautils' hint_roles (supautils.hint_roles = 'anon, authenticated, service_role', loaded via session_preload_libraries).

    • Same repro as yours (postgres session → SET ROLE anon|authenticated|service_role → call a function whose EXECUTE is revoked → signal 11, all backends reinitialized).
    • With supautils not loaded, or with PGOPTIONS='-c supautils.hint_roles=' (empty), the same call returns a clean 42501 permission denied for function, no crash.
    • Denials on tables, views or schemas don't crash; only function EXECUTE denials do, which points at the hint hook formatting the GRANT … ON FUNCTION … TO … hint.
    • PostgREST is unaffected because authenticator has session_preload_libraries=safeupdate (supautils not loaded).
  7. MRecinosOrg commented on Sep 29, 2026

    @MRecinosOrg

    Reproduced on 17.6.1.104 aarch64 (Apple Silicon, Docker Desktop, Rosetta off): any function the role lacks EXECUTE on, with either anon or authenticated, crashes the backend; a table-level denial in the same session returns 42501 normally. Two hosted projects on PostgreSQL 17.6 (aarch64 and x86_64) return 42501 cleanly for the identical statement.

  8. Edilson76 commented on Oct 7, 2026

    @Edilson76

    Edited. My original comment attributed this to pgAudit's ExecutorCheckPerms_hook and restated findings that were already in the thread. Both were wrong on my part: @kchebani had already bisected it to supautils hint_roles (loaded via session_preload_libraries), and the anon case and the table-vs-function asymmetry were reported earlier by @EmireHaouas and @MRecinosOrg. I had checked shared_preload_libraries, seen no supautils there, and concluded from the wrong parameter.

    Leaving only the part that does not seem to be in the thread yet:

    The crash takes down every open connection, not just the caller

    LOG:  server process (PID 74894) was terminated by signal 11: Segmentation fault
    DETAIL:  Failed process was running: SELECT public.f_crash_repro()
    LOG:  terminating any other active server processes
    LOG:  all server processes terminated; reinitializing
    LOG:  database system was not properly shut down; automatic recovery in progress
    LOG:  database system is ready to accept connections
    

    Recovery is fast (~690 ms), so on an interactive workload it looks like a blip. On a test suite it is much worse than it appears: a single test asserting "role X cannot execute function Y" — the natural way to cover a REVOKE EXECUTE policy — kills every connection in flight.

    Concretely, in one CI run of our RLS/isolation suites against 17.6.1.104: 58 test files lost, each failing with the database system is not accepting connections, none of them having anything to do with privileges. They were simply connecting during the recovery window. We had to teach our gate to recognize the crash signature and drop those files from the verdict, otherwise the noise is attributed to whoever opened the pull request.

    Worth flagging for anyone landing here from a CI failure: the symptom you see is almost never at the test that caused it.

    Confirmed on x86_64, image supabase/postgres:17.6.1.104. Thanks @kostasb for the pointer to >= 17.6.1.143 / supautils 3.2.2 — upgrading the image is the path for us.

  9. Edilson76 commented on Oct 9, 2026

    @Edilson76

    Confirmed fixed on 17.6.1.178

    @kostasb's pointer to >= 17.6.1.143 holds up in practice — verifying here since the thread so far only had the expectation, not a test.

    Upgraded the image from 17.6.1.104 to 17.6.1.178 on an existing volume (same major, no re-init needed) and re-ran the exact repro:

    EXECUTE -> anon=false  authenticated=false     (the bug's precondition)
    
    anon:          [42501] permission denied for function f_crash_repro
    authenticated: [42501] permission denied for function f_crash_repro
    
    server still up: YES
    

    Clean 42501 on both roles, no segfault, no recovery cycle. Before the upgrade, the same call on the same database killed the backend every single time.

    Worth noting for anyone who, like me, went looking in the wrong place: pgaudit is still loaded in shared_preload_libraries after the upgrade and nothing crashes. What changed is supautils, which lives in session_preload_libraries — a different parameter. I checked only the former early on and built a whole (wrong) hypothesis around pgAudit before @kchebani's bisect set it straight.

    session_preload_libraries: supautils
    shared_preload_libraries:  pg_stat_statements, pgaudit, plpgsql, plpgsql_check, pg_cron,
                               pg_net, pgsodium, auto_explain, pg_tle, plan_filter, supabase_vault
    

    Measured impact of the fix, on a real workload

    Our RLS/isolation test suites, same hardware, same database, before and after:

    17.6.1.104 17.6.1.178
    test files lost to server crashes 34–58 per run (varying) 0
    passing cases 1,753 1,789

    The variability is the part that hurt most: the set of casualties was decided by which files happened to be connecting during the ~690 ms recovery window, so the same commit produced different results run to run. That is what kept us from making the suite a required check.

    Thanks to everyone who reproduced and bisected this — it saved us from chasing the wrong library.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions