Ephemeral, containerised, self-hosted GitHub Actions runners for private repositories. One checkout runs a pool per repository, on as many machines as you like.
Need runs-on: windows-latest? The containers here are Linux only and will
never match a windows-* job. This repo also ships a second, native (non-containerised)
Windows fleet for that -- windows/README.md -- driven by
windows-pools.conf and windows-startRunners.ps1, and runnable side by side with the
Linux pools on the same host. The rest of this file is the Linux/container fleet.
As I have a lot of private repos, and although you have some free action runner minutes for private repos, unlimited for free repos, I burn through them quickly. You can "self-host" runners and GitHub will trigger them.
See below
Quick start -- if you just want it running, start there.
Windows (native) runners -- the separate fleet for
runs-on: [self-hosted, windows, x64] jobs, in its own README.
Everything else in this file
Have these three ready before starting:
- Docker Desktop (or Docker Engine on Linux), set to start on login
- GitHub CLI, authenticated --
gh auth login - A GitHub personal access token (PAT) -- create one at github.com/settings/tokens/new: give it a name, tick the
reposcope checkbox, click Generate token, and copy the value it shows you. GitHub only displays it this once.
Windows: run every command below through Git Bash, not raw PowerShell, and not
bare bash -- see Host prerequisites if a command below says
gh: command not found even though gh works fine elsewhere. That single gotcha costs
more time than everything else in this file combined.
1. Clone it:
git clone https://github.com/leonarduk/local-github-runner
cd local-github-runner2. Save the PAT from above -- gitignored, never committed:
# macOS / Linux / Git Bash
printf '%s' 'ghp_your_token_here' > pat.secret# Windows PowerShell
[System.IO.File]::WriteAllText("$PWD\pat.secret", 'ghp_your_token_here')3. Name this machine:
cp .env.example .env
# then edit .env and set RUNNER_HOST_LABEL to something like "bedroom" or "office-desktop"4. List the repo(s) this machine should serve, and bring them up:
cp pools.conf.example pools.conf
# edit pools.conf -- replace its example lines with your own, one per repo:
# jobtrack owner/jobtrack 2
# "owner" is your GitHub username or org, not a literal word -- for the
# author that's "leonarduk", e.g. leonarduk/jobtrack; use your own instead.
./startRunners.sh5. Confirm it actually registered -- a started container is not the same fact as an online runner:
./pools.sh listDocker Desktop shows the same pools grouped by project, each expandable to its
runner-1/runner-2 containers -- this is what several healthy pools on one
host look like:
GitHub's own Settings -> Actions -> Runners page on the target repo is the
other side of the same check -- each container above is one row here, Idle
once it has registered:
6. Point a workflow at it. A runner sits idle until a job asks for it -- edit the target repo's workflow file:
- runs-on: ubuntu-latest
+ runs-on: [self-hosted, linux, x64]Every job in that workflow needs the label, or the untouched ones keep failing to start for the original reason. See Point a workflow at it for the full version of this step, including why to keep the old line commented rather than deleted.
Done.
Adding a repo later means adding a line to pools.conf and running
./startRunners.sh again. ./pools.sh up <name> <owner/repo> [count] is the
one-off primitive underneath it, for a single pool with no config file at all.
Setting up a second machine, or want repos to opt themselves in instead of
being listed here by hand? See Several machines.
These repos moved from all-public to an open-core split: the framework and plumbing stay public, the parts worth something move to private repos, either to sell directly or at least to stop someone else forking public work and profiting from it before the author does.
GitHub-hosted Actions minutes are metered per account, not per repo, so that split had a cost nothing about it made obvious upfront: every new private repo adds its own CI, and the minutes compound rather than add. The free tier was gone by day 20 of the month; GitHub Pro's larger allotment was gone by day 5. Paying more moved the wall, it didn't remove it.
Self-hosted runners are free to use with GitHub Actions: you supply and maintain the machine, and none of that time counts against Actions minutes. That's the actual fix here — not a workaround for an outage, but a way to stop paying GitHub for compute already sitting idle on a machine at home.
It began inside one repo, as a fix for that repo's minutes problem, and was pointed at four more within a day. It lives here because a tool that serves five repositories should not be a subdirectory of the first one that needed it.
A single-host tool, deliberately. It is a Dockerfile and a compose file that keep a handful of ephemeral runners alive on one machine, and the interesting part is not the container -- it is the collection of failure modes documented here, most of which cost somebody an afternoon to find.
If you need something else, use something else:
- Kubernetes, autoscaling across a cluster, or more than one team -> actions-runner-controller, which is the supported answer and does all of that properly. Resizing one host's pools to their queued jobs is covered: see Autoscaling a host's pools.
- A public repository -> nothing here, and see the next section. This is not a limitation to work around; it is the one hard rule.
- macOS jobs -> not covered by anything in this repo.
- Jobs needing
container:,services:, or Docker-based actions -> those need Docker-in-Docker, which this deliberately does not provide.
Windows jobs are covered, just not by a container: see windows/README.md for ephemeral runners as plain Windows processes instead.
Never point this at a public repository. On a public repo, anyone can open a pull request, and a workflow triggered by that PR runs their code on your machine. (GitHub defaults public repos to requiring approval for a first-time contributor, so it is not quite a drive-by -- but that bar is one trivial merged PR high, and it is a setting that can drift.) GitHub's own documentation
is unusually blunt about this: "Self-hosted runners should almost never be used
for public repositories on GitHub, because any user can open pull requests
against the repository and compromise the environment." The container boundary
here is not a security boundary either — jobs get passwordless sudo inside the
container (see below).
Point it only at repos where you control who can trigger a run. If one is ever made public, tear its pool down first.
The sharpest consequence is the PAT. entrypoint.sh needs a token that can mint a
registration token -- a classic PAT with repo, which reaches every repository on
the account -- and Compose keeps it mounted at /run/secrets/github_pat for the
container's whole life, not only during registration. Jobs run with passwordless
sudo, so any job can read it. GitHub's fork protections do not cover this: they
withhold workflow secrets from fork pull requests and make GITHUB_TOKEN
read-only, but this PAT is a file on the runner, not a workflow secret. Prefer a
fine-grained PAT
limited to the single target repository, which bounds what a
leaked token can reach.
Two things reduce the blast radius even so:
- Every container serves exactly one job, then exits (
--ephemeral). Nothing a job leaves behind — files, processes, a poisoned pip cache — is visible to the next one. - No volumes. The workspace lives in the container's own writable layer, which is discarded on exit. Adding a named volume for
_workwould quietly undo the isolation.
By hand, per machine, with docker compose. There is no control plane and
nothing to install besides Docker: a host that should serve runners gets a
clone, a PAT, and one compose up per pool. A machine that should stop serving
them gets a compose down.
This repo did carry a Jenkins pipeline to do the same thing. It was removed:
for a single host it wrapped two environment variables and a compose up while
adding a stateful service holding the Docker socket -- root on the host -- and
it was never actually run end to end. Per-machine setup is cheap enough that
the orchestration was not buying anything.
pools.sh wraps the operations below for a host running several pools:
./pools.sh up <name> <owner/repo> [count] [label] [mem] [pids] # bring a pool up
./pools.sh down <name> # tear one down
./pools.sh reset <name> <owner/repo> [count] [label] [mem] [pids] # down, then up fresh
./pools.sh start <name> # bring one up exactly as pools.conf declares it
./pools.sh stop <name> [--force] # tear one down, unless a runner is busy
./pools.sh restart <name> [--force] # stop, then start, unless a runner is busy
./pools.sh restart-runner <container> [--force] # restart one runner container, unless it is busy
./pools.sh scale <name> <count> [--force] # resize a pools.conf pool and its count there, unless shrinking would kill a busy runner
./pools.sh declare <name> <owner/repo> [count] # add a pool to pools.conf, starting nothing
./pools.sh sync [--dry-run] # rewrite pools.conf to match the pools running here
./pools.sh list [--json [<name>...]] # every pool on this host, and what GitHub actually sees
./pools.sh autoscale [--once] [--dry-run] [--interval <seconds>] # resize autoscale.conf's pools to their queued jobs<name> is just the label in the project name (gh-runner-<name>); it need
not match the repo. list is the one worth knowing about even if you never
use the others: it cross-checks each pool's containers against GitHub's own
actions/runners API, which is the only way to catch a pool that looks fine
in docker ps but registered nothing.
start, stop, restart, restart-runner, scale and list --json are
for driving pools from something else -- a dashboard, a cron job:
start <name>takes nothing but the name. The repo, size, label and limits come from that pool'spools.confline, so a caller can't bring up a pool that file doesn't describe.stop <name>asks GitHub first and exits3instead of stopping if any of the pool's runners is busy -- adownmid-job cancels the job. It also exits3when GitHub can't be asked, or when no runner on GitHub matches any of the pool's containers, since "unknown" is not "idle".--forceskips the check. It stops any pool running here, declared inpools.confor not.restart <name>isstopthenstart, with the same busy check, so it only works on a poolpools.confdeclares.restart-runner <container>restarts one runner container, named asdocker psshows it (gh-runner-jobtrack-runner-1): it deregisters and comes straight back as a fresh runner. It refuses if that runner is busy, or if GitHub can't be asked. A container with no runner registered -- one stuck failing to register, say -- is restarted, unless no container in its pool matches a runner either, which would mean the matching is broken.scale <name> <count>resizes a poolpools.confdeclares in place, instead of tearing it down and starting fresh likerestartdoes. Growing -- or bringing up a pool with nothing running -- is justupwith the new count, so it never refuses; there's nothing running yet that scaling up could hurt. Shrinking gets the same busy check asstopandrestart, becausedocker compose up --scaledown can't be told which containers to kill, and a busy one dying cancels its job.--forceskips that check. Once the pool is resized,<count>is written into itspools.confline -- nothing else in the file changes -- so a laterstartorrestartbrings it back at the new size instead of undoing the scale. A refused or failed scale leavespools.confalone.declare <name> <owner/repo> [count]adds a line for a poolpools.confdoesn't have yet -- 1 runner unless[count]says otherwise, and the default label and limits -- sostart,restartandscalecan bring it up. It starts nothing, and refuses a name that's already declared.syncgoes the other way fromstart: it rewritespools.confto describe the pools on this host. Each declared pool's count becomes the number of containers it has (running or not -- its compose scale), and eachgh-runner-*project the file doesn't declare gets a line, with the repo, extra label and mem/pids limits read off its containers, so a laterstartorrestartbrings it back the same. Declared pools with no containers, comments and every other column are left alone, and the old file is kept aspools.conf.bak.--dry-runprints the changes without writing them. It only asks docker, never GitHub.list --jsonprints one JSON array with every pool inpools.confplus everygh-runner-*project running here thatpools.confdoesn't declare ("managed": false) -- pools started by hand, whichstopRunners.shnever touches. Itsrunnerscounts only the GitHub runners registered by that pool's own containers (matched on the container ID in each runner's name), so two pools serving one repo -- a CI pool and a dedicatedissue-wormpool, say -- are told apart.runnersisnullwhen GitHub couldn't be asked.membershas one entry per container: its docker state and status, and the GitHub runner it registered (nullif none, or if GitHub couldn't be asked).list --json issue-wormlimits it to the named pools: every pool costs a GitHub API call, so a whole-host listing with a dozen pools takes tens of seconds, which matters for anything that polls.
[{"name":"issue-worm","project":"gh-runner-issue-worm","repo":"leonarduk/issue-worm-pro",
"managed":true,"desired":2,"label":"issue-worm",
"containers":{"total":2,"running":2},"runners":{"online":2,"busy":0},
"members":[{"container":"gh-runner-issue-worm-runner-1","id":"f4c4835f30e8",
"state":"running","status":"Up 8 hours",
"runner":{"name":"bedroom-f4c4835f30e8-1","status":"online","busy":false}}]}]Runners registered by other hosts serving the same repo are not counted: a host can only see and control its own containers.
stop, restart, restart-runner, a shrinking scale and list --json
read repos/<owner>/<repo>/actions/runners through the host's own gh
login, which needs admin access to the repo -- a classic token with repo,
or a fine-grained one with Administration: read. Without it, the busy
checks refuse and list --json reports "runners": null.
tests/pools_test.sh exercises start/stop/restart/restart-runner/scale/sync/list --json/autoscale against stub
docker and gh commands, so it needs neither a Docker daemon nor GitHub.
A fixed count is always wrong for somebody: too many runners holding memory
while nothing is queued, or too few while jobs wait behind each other.
./pools.sh autoscale resizes the pools listed in autoscale.conf to what is
actually queued:
cp autoscale.conf.example autoscale.conf # gitignored; one "<name> <min> <max> [idle_minutes]" line per pool
./pools.sh autoscale --once --dry-run # what it would do right now, changing nothing
./pools.sh autoscale # every 60s until stopped; --interval to change thatEvery pass, for each listed pool:
- Demand is the jobs queued in its repo -- in queued runs, and in runs
already in progress with jobs still waiting -- whose
runs-onlabels this pool's runners carry, including this host'sRUNNER_HOST_LABEL. A job more than onepools.confpool could take counts against the one with the fewest labels, so a pool reserved by an extra label (issue-worm) doesn't grow for the generic jobs its repo's plain CI pool is there for. - Growing: to busy runners plus queued jobs, capped at
<max>, when that is more containers than the pool has, and never below<min>. Containers still starting count as capacity, so a job that stays queued while its runner boots doesn't grow the pool a second time. - Shrinking: to
<min>once the pool has had no busy runner and no queued job for[idle_minutes](10 by default), and only after the same busy checkstopmakes. The idle clock is kept in.autoscale-state, so it survives from one pass to the next. A pool with any busy runner is not shrunk at all, becausedocker compose up --scalecan't be told which containers to remove. - A pool whose repo GitHub can't be asked about is left as it is.
It resizes with the same docker compose up --scale as scale, and never
writes pools.conf: that file's count is still what start and restart
use, and the next pass corrects it. Each pass costs, per autoscaled repo, one call
for its runners, one per page of queued and in-progress runs, and one per
such run for its jobs -- so a handful a minute for a quiet repo, but a repo
with 20 runs in flight costs over 20 a minute. A token gets 5,000 an hour;
raise --interval if several busy repos are autoscaled.
Two hosts serving the same repo each see the same queued job, so each may
add a runner for it; the spare one sits idle and goes again after
[idle_minutes]. Nothing here sizes memory or CPU yet: a pool still gets
its pools.conf [mem]/[pids] limits. The native Windows fleet has the
same thing as windows-pools.ps1 autoscale; see
windows/README.md.
Keep it running the way you keep anything else on the host running -- a systemd unit, or at the least:
nohup ./pools.sh autoscale >> autoscale.log 2>&1 &[label], [mem] and [pids] are optional, trailing, and positional --
pass - for one you want to leave at its default so a later one still lands
in the right slot. They let a pool carry an extra runner label and its own
mem_limit/pids_limit, separate from every other pool on the host; see
Dedicated pools for a specific kind of job.
You need the PAT described next.
Checklist first, story below for whichever line actually bites:
- Docker installed, engine set to start on login
- Enough memory allocated in Docker Desktop for however many containers you run
- Sleep disabled on this host while it serves runners (both idle timeout and lid-close -- see below, this one is sneaky)
- GitHub CLI installed and authenticated
- On Windows, scripts are invoked via Git Bash, not raw
bash(which may silently be WSL) - On Windows, an execution policy that runs unsigned local scripts, if you use the
.ps1entry points (startRunners.ps1,stopRunners.ps1) -- see below
The pool is only as reliable as the machine under it, and two of these are the kind of thing you debug for an afternoon before suspecting them.
Windows execution policy. Nothing in this repo is code-signed, so a host on
AllSigned refuses every .ps1 here before any of this code runs, naming the
file rather than the policy: "File ...\startRunners.ps1 cannot be loaded. The
file ... is not digitally signed." It reads like a broken script and is not.
Get-ExecutionPolicy -List shows which scope is in force -- typically
AllSigned on LocalMachine with CurrentUser left Undefined -- and
Set-ExecutionPolicy -Scope CurrentUser -ExecutionPolicy RemoteSigned fixes it
for your account without admin rights or weakening the machine-wide setting. It
affects only the PowerShell entry points: the Git Bash path (startRunners.sh,
pools.sh) is untouched by this, which is worth remembering when one fleet
starts and the other refuses. Full detail, including why RemoteSigned rather
than Bypass, is in windows/README.md.
Docker. Docker Desktop on Windows or macOS, Docker Engine on Linux. On Windows use the WSL2 backend; the Hyper-V backend works but is slower at the bind mounts this uses.
Set the engine to start on login — Docker Desktop → Settings → General →
Start Docker Desktop when you sign in. restart: always brings a pool back
after a reboot, but only once the engine is running. Without it the jobs simply
queue with no runner, which looks like CI is broken rather than a stopped
Docker engine.
Give it enough memory. Each runner declares mem_limit: 1g by default, so N
runners across all your pools can ask for N GiB, against whatever ceiling
Docker Desktop is set to in Settings → Resources. Over-committing does not
error — it shows up as jobs mysteriously crawling when several repos build at
once. Count the containers, not the pools — and a pool with a raised
mem_limit (see Dedicated pools for a specific kind of job)
counts for more than one container's worth per replica.
Sleep. This one silently cancels jobs. A host that suspends mid-job stops the runner's heartbeat, and GitHub cancels the job server-side. The signature is unmistakable once you know it, and it shows up two different ways depending on whether a job had been claimed yet when the host went under.
A job already running dies about ten minutes after the host goes under —
that interval is GitHub's own patience with a runner that has stopped
answering, not any timeout on your machine, so it is the same ten minutes
whatever sent the host to sleep. Do not read it as pointing at a ten-minute
power setting; that coincidence cost us an afternoon here. The tell is
that its logs keep going past its own recorded completion: the container had
the step suspended, not finished, and uploads the rest on wake into a job record
that is already closed. One observed here completed at 13:42:49Z with log
lines running to 14:47:09Z — 64 minutes after it supposedly ended. A job that
finished before its own logs were written is not a race condition; it is a
sleeping host, and it is a cheap thing to grep for.
A job still queued simply is not claimed, because no runner is awake to take
it. It then starts at the moment of wake and runs at completely normal speed —
in the same incident, a run queued at 13:32:45Z started its jobs at
14:47:26Z and finished them in 8 to 109 seconds, all green. Nothing was slow;
the machine was absent. This one is easy to misread as contention, which is what
makes the wake timestamp worth checking: jobs across different repositories
resuming within the same few seconds is one host waking, not several flakes.
Either way it is intermittent — jobs shorter than the outage finish fine — so it reads as a flaky test suite rather than a host problem.
There are two separate ways a host goes under, and fixing one does nothing about the other. On Windows both are recorded: Kernel-Power event 42 carries a reason, where 7 is an idle timeout and 0 is the lid or the power button. Check which you actually have before fixing anything.
# Which kind of sleep, and when -- reason 7 = idle, reason 0 = lid/button
Get-WinEvent -FilterHashtable @{LogName='System';
ProviderName='Microsoft-Windows-Kernel-Power'; Id=42} |
ForEach-Object { '{0:u} reason={1}' -f $_.TimeCreated, $_.Properties[2].Value }
# Idle timeout: STANDBYIDLE's AC index, in seconds
powercfg /query SCHEME_CURRENT SUB_SLEEP
powercfg /change standby-timeout-ac 0Closing the lid is the one that catches people, because a laptop on a desk gets shut without anyone thinking of it as powering the machine down, and no idle timeout protects against it. Worse, the setting that governs it is hidden from the Power Options UI on some machines, so it cannot be found by clicking through Windows settings — it has to be set by GUID, elevated:
# Lid close -> do nothing, on AC. SUB_BUTTONS / LIDACTION.
powercfg /setacvalueindex SCHEME_CURRENT `
4f971e89-eebd-4455-a8de-9e59040e7347 5ca83367-6e45-459f-a27b-476b1d01c936 0
powercfg /setactive SCHEME_CURRENTThat binds to the active power scheme only, so switching schemes silently reverts it. And a closed lid under sustained build load is a real thermal question on a laptop — pair it with the machine being docked and ventilated rather than setting it blind.
# macOS
sudo pmset -c sleep 0
# Linux -- or set logind's IdleAction to ignore
systemd-inhibit --what=idle:sleep --who=gh-runner --why=CI sleep infinity &Disabling sleep on battery too is usually the wrong trade on a laptop; if the
machine is a CI host on AC, standby-timeout-ac 0 is enough. Display sleep is
harmless — it is system standby that kills jobs.
Registering a runner by hand, without any of the tooling in this repo, is
documented directly by GitHub
— Settings -> Actions -> Runners -> New self-hosted runner on the target repo.
Everything below is the same registration, done by entrypoint.sh inside a
disposable container instead of by hand on a long-lived machine.
GitHub CLI. pools.sh, startRunners.sh/stopRunners.sh and discoverPools.sh all shell out to gh on the host to list repos and check which runners are actually online -- a separate installation from the gh baked into the runner container, which only helps workflows running inside a job. Install it and authenticate before using any of these scripts:
winget install --id GitHub.cli # Windows; see https://cli.github.com for other platforms
gh auth login # in a new shell, so PATH picks up the installOn Windows, invoke every script here with an explicit path to Git Bash rather
than bare bash -- if WSL is installed (likely, since Docker Desktop's
recommended backend needs it), bash on PATH resolves to
C:\WINDOWS\system32\bash.exe, WSL's launcher into a separate Linux
environment with its own PATH. gh installed on Windows via winget is
invisible there, so a script fails with a plain gh: command not found that
has nothing to do with whether gh is actually installed. Confirm which one
bash means with Get-Command bash, and if it points into
system32 or WindowsApps, use the real path instead:
& 'C:\Program Files\Git\bin\bash.exe' ./discoverPools.sh leonardukWithout it, these scripts fail fast with gh: command not found rather than
silently doing nothing -- but on Windows that error can end up wherever
stderr goes for whatever launched bash, so it is easy to miss if you are not
looking for it.
You need a personal access token that can register runners:
- Classic PAT —
reposcope, or - Fine-grained PAT — the target repositories,
Administration: read & write
One PAT covering several repositories can serve several pools. The container exchanges it for a short-lived registration token at start-up; the PAT itself is never baked into the image.
Quick start already walked through this via pools.sh -- what follows is the same thing one level down, with the raw docker compose command pools.sh up runs for you, for when something needs debugging directly.
git clone https://github.com/leonarduk/local-github-runner
cd local-github-runner
printf '%s' 'ghp_your_token_here' > pat.secret # gitignored
cp .env.example .env # set RUNNER_HOST_LABEL
GITHUB_REPOSITORY=owner/repo docker compose up -d --build
docker compose logs -fA healthy start looks like:
runner-entrypoint: requesting a registration token for owner/repo
runner-entrypoint: configuring bedroom-4f2c1a9b3e77-1
runner-entrypoint: waiting for a job
The runner then appears under the repo's Settings → Actions → Runners as idle. compose.yaml declares two of them; override that for a one-off with --scale runner=3.
GITHUB_REPOSITORY has no default and the up fails without it. That is deliberate: it used to default to the one repo this was written for, which is exactly the kind of thing that survives a copy to a new machine and quietly registers runners against the wrong repository.
docker compose up returning 0 means the containers started, not that any
runner registered. Registration happens later, inside the container, and can
fail on its own -- a PAT without the right scope, a typo in
GITHUB_REPOSITORY, a repo the token cannot see -- leaving containers that
restart forever while GitHub shows nothing. Ask GitHub rather than Docker:
gh api repos/OWNER/REPO/actions/runners \
--jq '.runners[] | "\(.name) \(.status) busy=\(.busy)"'Expect one online line per container, each prefixed with the host label from
.env. Fewer than you scaled to means some are still registering or some are
failing; docker compose logs -f says which. Runners listed offline with no
container behind them are strays from a container that was killed rather than
stopped -- clear them with
gh api -X DELETE repos/OWNER/REPO/actions/runners/ID.
The compose project name is the only thing separating one pool from another. Two pools sharing it are the same pool — an up for the second repo silently reconfigures the first repo's containers rather than adding to them. So set it alongside the repository:
GITHUB_REPOSITORY=leonarduk/jobtrack COMPOSE_PROJECT_NAME=gh-runner-jobtrack docker compose up -dOnly put per-machine facts in .env — RUNNER_HOST_LABEL, and GITHUB_REPOSITORY if this host serves exactly one repo. Every pool started from this directory reads that same file, so anything that differs per pool belongs on the command line.
Check what is running, across all pools, with docker ps --filter name=gh-runner.
Docker Desktop groups them the same way, by compose project -- each expandable
row is one pool, runner-1/runner-2 its containers (see the screenshot in
Quick start step 5).
Nothing is host-specific except two gitignored files, so a second machine is a clone plus those:
git clonethis repo.- Write the PAT to
pat.secret— it never travels through git. cp .env.example .envand setRUNNER_HOST_LABELto that machine's name, e.g.bedroom.- Bring up whichever pools that machine should serve. Two ways:
cp pools.conf.example pools.conf, edit it to list the repos this host serves, then./startRunners.sh.- Or
./discoverPools.sh <owner>, which needs no local config at all -- it finds every private repo under<owner>carrying a.local_runnermarker file and brings each one up. Either way,./pools.sh up <name> <owner/repo> [count]remains the one-off, no-config primitive underneath both.
RUNNER_HOST_LABEL becomes both a runner label and the runner-name prefix, so the GitHub runner list and gh api .../actions/runners say which box a runner is on — container hostnames are random hex and no help. It also lets a workflow pin a job to one machine with runs-on: [self-hosted, bedroom].
Runners on different machines serving the same repo simply join the same pool: GitHub hands each queued job to whichever is free. There is no coordination between hosts and none is needed -- the screenshot in Quick start step 5 is one repo with runners from two machines, bedroom and steves-big-laptop, sitting in the same list.
The image is built per machine (--build). There is no registry involved, so a second host costs one ~1.4GB build rather than any shared infrastructure.
A runner sits idle until a job asks for it. In the target repo's workflow:
- runs-on: ubuntu-latest
+ runs-on: [self-hosted, linux, x64]Every job in the workflow needs the label, or the untouched ones keep exhausting the account's Actions minutes as before. Keep the ubuntu-latest line commented directly above each replacement, so reverting to hosted runners is a one-line edit at the point of use rather than an archaeology exercise.
Running the same tests on Windows instead (or as well)? windows/README.md covers runs-on: [self-hosted, windows, x64], set up the same way but without a container.
Last verified: 2026-09-05. Re-run and update this date when making changes that could affect runner behavior.
Verified end to end against leonarduk/spring-professional-udemy-practice-tests: both its jobs ran on a pool from this image and passed, in 7s and 32s, having previously failed in about 2s without starting.
GitHub's own reference for this syntax and the default label set is Choosing the runner for a job.
Every pool so far shares one shape: a repo's normal CI, where a runner picks
up a job, runs it in minutes, and frees up again. Some jobs do not behave
like that -- a long-running pipeline that itself opens pull requests is the
one that surfaced this. issue-worm-pro's pipeline clones a repo, builds a
venv and runs pytest inside a container job, which already needs more than
the 1g/512-pid defaults below are sized for. Worse, the PR it opens needs
runners of its own, for python-ci.yml and the review workflows, and a
Changes Requested requeue loop waits on those. Point enough concurrent worm
jobs at the same pool that serves that repo's CI, and they can occupy every
runner while the CI they are blocked on queues behind them -- a
self-inflicted deadlock, not an external one.
The fix is a separate pool with its own label, never the CI pool: worm
jobs run where python-ci.yml and the review workflows cannot be queued
behind them, because they are physically on different runners.
RUNNER_EXTRA_LABELS (compose.yaml) is what makes a pool reachable that way.
It appends one or more comma-separated labels after the usual
self-hosted,linux,x64,docker,<host> set, so a workflow can target
runs-on: [self-hosted, issue-worm] and land only on pools carrying that
label -- CI's runs-on: [self-hosted, linux, x64] never matches it, and
nothing else needs to change. Unset, it reproduces today's label set exactly
-- no trailing comma, nothing appended.
POOL_MEM_LIMIT / POOL_PIDS_LIMIT (compose.yaml) override that pool's
mem_limit / pids_limit away from the 1g/512 defaults, which a clone +
venv + pytest build is liable to exceed. Unset, both stay at today's
values.
All three are per-pool, so they are set the same way as [count]: as
trailing, optional, positional arguments to ./pools.sh up/reset, or as
the fourth through sixth columns of a pools.conf line (- skips one to
reach a later column). They are deliberately not .env settings -- .env
holds facts about the machine, and these differ per pool on the same
machine.
# one-off, via pools.sh
./pools.sh up issue-worm leonarduk/issue-worm-pro 2 issue-worm 2g 1024
# or in pools.conf, alongside the plain CI pool for the same repo
# name owner/repo count label mem pids
issue-worm-pro leonarduk/issue-worm-pro 4
issue-worm leonarduk/issue-worm-pro 2 issue-worm 2g 1024Then point the worm workflow(s) at the label, leaving every other job in that repo (and every other repo's CI) targeting the plain pool:
- runs-on: ubuntu-latest
+ runs-on: [self-hosted, issue-worm]This mitigates the starvation risk; it does not eliminate queuing outright.
Two pools drawing from the same finite host memory still compete for it, and
undersizing either pool's [count] just moves the queue rather than removing
it -- size each pool to the concurrency it actually needs, the same way as
any other pool (see Host prerequisites, "Give it
enough memory").
| Base | ubuntu:24.04 |
| Runner | v2.337.0, SHA256-verified at build, --disableupdate |
| Arch | linux/amd64 and linux/arm64 (via TARGETARCH) |
| Tooling | git, curl, jq, tar, unzip, zstd, sudo, python3 3.12, uuidgen, pkill |
gh |
v2.100.0, SHA256-verified at build — preinstalled on GitHub-hosted images, so workflows assume it |
| JS actions | bundled Node 20 and Node 24, verified working |
| Tool cache | /opt/hostedtoolcache, writable — actions/setup-python needs this |
Two deliberate choices worth knowing about, because both look like oversights:
sudois installed, passwordless. Workflows install tools withsudo mv, matching what a GitHub-hosted runner allows. Removing sudo, or settingno-new-privilegesin compose, breaks those steps.ghis baked in rather than installed per job. Ubuntu's own package is too old for the--jsonflags workflows use, so repos had started curling the release tarball at job time without verifying it. Installing it here, checksummed, removes that. Workflows that guard onghalready being onPATHbecome no-ops; the binary carries no credentials, so each workflow still supplies its ownGH_TOKEN.RUNNER_MANUALLY_TRAP_SIG=1. Makes the runner handleSIGTERMitself and finish or cancel the running job cleanly. Without it,docker stopmid-job leaves the job hung until GitHub times it out.
Bumping the runner version means editing RUNNER_VERSION and the two checksums in the Dockerfile together — they are per-version, and a mismatched pair fails the build rather than silently skipping verification. The SHAs are published in each release's notes at https://github.com/actions/runner/releases.
--disableupdate keeps the pinned version honest. If GitHub starts refusing jobs from a version this old, drop that flag and rebuild the image to match.
A version bump has to reach every machine, since each builds its own image. Nothing enforces that; docker images local-github-runner:local on each host is how you tell.
COMPOSE_PROJECT_NAME=gh-runner-repo docker compose downEach container deregisters itself on the way out, so nothing should be left behind — a down of a two-runner pool completes in about five seconds with both registrations gone.
That relies on entrypoint.sh signalling Runner.Listener directly rather than only run.sh: a signal to the wrapper alone never reaches the process RUNNER_MANUALLY_TRAP_SIG=1 arms, and the container is then SIGKILLed with deregister() unreached. If a container is killed outright anyway — docker kill, a host crash — its runner lingers as "offline" under Settings → Actions → Runners and needs removing there, or with gh api -X DELETE repos/OWNER/REPO/actions/runners/ID. The container itself recovers unaided: the leftover .runner from its previous life is cleared at start-up before it reconfigures.
- No Docker-in-Docker. Adding it means mounting the host's Docker socket, which hands any job root on the host — do not do that on the strength of this README alone.
actions/cachehas no backing store, so cache steps are no-ops that cost a little time. Worth knowing if a workflow starts depending on a warm cache.pip installonly works afteractions/setup-python. The base is Ubuntu 24.04, whose system interpreter refuses installs under PEP 668 (externally-managed-environment).setup-pythonputs its own interpreter on PATH first — but a workflow that drops that step keeps working on a GitHub-hosted runner and fails here.mem_limit: 1g/pids_limit: 512are conservative defaults. If a build is OOM-killed, raise them per pool withPOOL_MEM_LIMIT/POOL_PIDS_LIMITrather than removing the limits outright -- see Dedicated pools for a specific kind of job.- This container image is Linux only. A workflow pinned to
windows-*will never match it -- see windows/README.md for the separate, non-containerised path that does.macos-*is not covered by anything in this repo. - Pool names use the repo name, not
owner/repo. Two repos of the same name under different owners would collide.
Everything above is GitHub's own supported mechanism -- a self-hosted runner registered the normal way -- wrapped in a container and a compose file for convenience. None of it depends on this repo; every piece is documented independently, and it is worth reading the source rather than taking this README's word for the security-critical parts:
- About self-hosted runners — what one is, and the three levels (repository, organization, enterprise).
- Adding self-hosted runners — the manual registration flow
entrypoint.shautomates. - Secure use reference — the authoritative source for the public-repository warning above. Read this one regardless of whether you read anything else here.
- Choosing the runner for a job —
runs-on:syntax and how label matching works. - Managing your personal access tokens — classic vs. fine-grained, and how to scope the one this needs.
- actions-runner-controller — GitHub's own answer once one host and a compose file stop being enough.

