Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

126 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

local-github-runner

WHAT

Ephemeral, containerised, self-hosted GitHub Actions runners for private repositories. One checkout runs a pool per repository, on as many machines as you like.

Need runs-on: windows-latest? The containers here are Linux only and will never match a windows-* job. This repo also ships a second, native (non-containerised) Windows fleet for that -- windows/README.md -- driven by windows-pools.conf and windows-startRunners.ps1, and runnable side by side with the Linux pools on the same host. The rest of this file is the Linux/container fleet.

WHY

As I have a lot of private repos, and although you have some free action runner minutes for private repos, unlimited for free repos, I burn through them quickly. You can "self-host" runners and GitHub will trigger them.

HOW

See below

Quick start -- if you just want it running, start there.

Windows (native) runners -- the separate fleet for runs-on: [self-hosted, windows, x64] jobs, in its own README.

Everything else in this file

Quick start

Have these three ready before starting:

Windows: run every command below through Git Bash, not raw PowerShell, and not bare bash -- see Host prerequisites if a command below says gh: command not found even though gh works fine elsewhere. That single gotcha costs more time than everything else in this file combined.

1. Clone it:

git clone https://github.com/leonarduk/local-github-runner
cd local-github-runner

2. Save the PAT from above -- gitignored, never committed:

# macOS / Linux / Git Bash
printf '%s' 'ghp_your_token_here' > pat.secret
# Windows PowerShell
[System.IO.File]::WriteAllText("$PWD\pat.secret", 'ghp_your_token_here')

3. Name this machine:

cp .env.example .env
# then edit .env and set RUNNER_HOST_LABEL to something like "bedroom" or "office-desktop"

4. List the repo(s) this machine should serve, and bring them up:

cp pools.conf.example pools.conf
# edit pools.conf -- replace its example lines with your own, one per repo:
#   jobtrack   owner/jobtrack   2
# "owner" is your GitHub username or org, not a literal word -- for the
# author that's "leonarduk", e.g. leonarduk/jobtrack; use your own instead.
./startRunners.sh

5. Confirm it actually registered -- a started container is not the same fact as an online runner:

./pools.sh list

Docker Desktop shows the same pools grouped by project, each expandable to its runner-1/runner-2 containers -- this is what several healthy pools on one host look like:

Several gh-runner-* compose projects, each expandable to its runner containers, in Docker Desktop

GitHub's own Settings -> Actions -> Runners page on the target repo is the other side of the same check -- each container above is one row here, Idle once it has registered:

Settings -> Actions -> Runners for one repo, showing self-hosted Linux X64 runners, all Idle

6. Point a workflow at it. A runner sits idle until a job asks for it -- edit the target repo's workflow file:

-    runs-on: ubuntu-latest
+    runs-on: [self-hosted, linux, x64]

Every job in that workflow needs the label, or the untouched ones keep failing to start for the original reason. See Point a workflow at it for the full version of this step, including why to keep the old line commented rather than deleted.

Done.

Adding a repo later means adding a line to pools.conf and running ./startRunners.sh again. ./pools.sh up <name> <owner/repo> [count] is the one-off primitive underneath it, for a single pool with no config file at all. Setting up a second machine, or want repos to opt themselves in instead of being listed here by hand? See Several machines.

Why this exists

These repos moved from all-public to an open-core split: the framework and plumbing stay public, the parts worth something move to private repos, either to sell directly or at least to stop someone else forking public work and profiting from it before the author does.

GitHub-hosted Actions minutes are metered per account, not per repo, so that split had a cost nothing about it made obvious upfront: every new private repo adds its own CI, and the minutes compound rather than add. The free tier was gone by day 20 of the month; GitHub Pro's larger allotment was gone by day 5. Paying more moved the wall, it didn't remove it.

Self-hosted runners are free to use with GitHub Actions: you supply and maintain the machine, and none of that time counts against Actions minutes. That's the actual fix here — not a workaround for an outage, but a way to stop paying GitHub for compute already sitting idle on a machine at home.

It began inside one repo, as a fix for that repo's minutes problem, and was pointed at four more within a day. It lives here because a tool that serves five repositories should not be a subdirectory of the first one that needed it.

What this is not

A single-host tool, deliberately. It is a Dockerfile and a compose file that keep a handful of ephemeral runners alive on one machine, and the interesting part is not the container -- it is the collection of failure modes documented here, most of which cost somebody an afternoon to find.

If you need something else, use something else:

  • Kubernetes, autoscaling across a cluster, or more than one team -> actions-runner-controller, which is the supported answer and does all of that properly. Resizing one host's pools to their queued jobs is covered: see Autoscaling a host's pools.
  • A public repository -> nothing here, and see the next section. This is not a limitation to work around; it is the one hard rule.
  • macOS jobs -> not covered by anything in this repo.
  • Jobs needing container:, services:, or Docker-based actions -> those need Docker-in-Docker, which this deliberately does not provide.

Windows jobs are covered, just not by a container: see windows/README.md for ephemeral runners as plain Windows processes instead.

⚠️ Private repositories only

Never point this at a public repository. On a public repo, anyone can open a pull request, and a workflow triggered by that PR runs their code on your machine. (GitHub defaults public repos to requiring approval for a first-time contributor, so it is not quite a drive-by -- but that bar is one trivial merged PR high, and it is a setting that can drift.) GitHub's own documentation is unusually blunt about this: "Self-hosted runners should almost never be used for public repositories on GitHub, because any user can open pull requests against the repository and compromise the environment." The container boundary here is not a security boundary either — jobs get passwordless sudo inside the container (see below).

Point it only at repos where you control who can trigger a run. If one is ever made public, tear its pool down first.

The sharpest consequence is the PAT. entrypoint.sh needs a token that can mint a registration token -- a classic PAT with repo, which reaches every repository on the account -- and Compose keeps it mounted at /run/secrets/github_pat for the container's whole life, not only during registration. Jobs run with passwordless sudo, so any job can read it. GitHub's fork protections do not cover this: they withhold workflow secrets from fork pull requests and make GITHUB_TOKEN read-only, but this PAT is a file on the runner, not a workflow secret. Prefer a fine-grained PAT limited to the single target repository, which bounds what a leaked token can reach.

Two things reduce the blast radius even so:

  • Every container serves exactly one job, then exits (--ephemeral). Nothing a job leaves behind — files, processes, a poisoned pip cache — is visible to the next one.
  • No volumes. The workspace lives in the container's own writable layer, which is discarded on exit. Adding a named volume for _work would quietly undo the isolation.

How it is run

By hand, per machine, with docker compose. There is no control plane and nothing to install besides Docker: a host that should serve runners gets a clone, a PAT, and one compose up per pool. A machine that should stop serving them gets a compose down.

This repo did carry a Jenkins pipeline to do the same thing. It was removed: for a single host it wrapped two environment variables and a compose up while adding a stateful service holding the Docker socket -- root on the host -- and it was never actually run end to end. Per-machine setup is cheap enough that the orchestration was not buying anything.

pools.sh wraps the operations below for a host running several pools:

./pools.sh up    <name> <owner/repo> [count] [label] [mem] [pids]  # bring a pool up
./pools.sh down  <name>                                            # tear one down
./pools.sh reset <name> <owner/repo> [count] [label] [mem] [pids]  # down, then up fresh
./pools.sh start <name>                                             # bring one up exactly as pools.conf declares it
./pools.sh stop  <name> [--force]                                   # tear one down, unless a runner is busy
./pools.sh restart <name> [--force]                                 # stop, then start, unless a runner is busy
./pools.sh restart-runner <container> [--force]                     # restart one runner container, unless it is busy
./pools.sh scale <name> <count> [--force]                           # resize a pools.conf pool and its count there, unless shrinking would kill a busy runner
./pools.sh declare <name> <owner/repo> [count]                      # add a pool to pools.conf, starting nothing
./pools.sh sync  [--dry-run]                                        # rewrite pools.conf to match the pools running here
./pools.sh list  [--json [<name>...]]                               # every pool on this host, and what GitHub actually sees
./pools.sh autoscale [--once] [--dry-run] [--interval <seconds>]    # resize autoscale.conf's pools to their queued jobs

<name> is just the label in the project name (gh-runner-<name>); it need not match the repo. list is the one worth knowing about even if you never use the others: it cross-checks each pool's containers against GitHub's own actions/runners API, which is the only way to catch a pool that looks fine in docker ps but registered nothing.

start, stop, restart, restart-runner, scale and list --json are for driving pools from something else -- a dashboard, a cron job:

  • start <name> takes nothing but the name. The repo, size, label and limits come from that pool's pools.conf line, so a caller can't bring up a pool that file doesn't describe.
  • stop <name> asks GitHub first and exits 3 instead of stopping if any of the pool's runners is busy -- a down mid-job cancels the job. It also exits 3 when GitHub can't be asked, or when no runner on GitHub matches any of the pool's containers, since "unknown" is not "idle". --force skips the check. It stops any pool running here, declared in pools.conf or not.
  • restart <name> is stop then start, with the same busy check, so it only works on a pool pools.conf declares.
  • restart-runner <container> restarts one runner container, named as docker ps shows it (gh-runner-jobtrack-runner-1): it deregisters and comes straight back as a fresh runner. It refuses if that runner is busy, or if GitHub can't be asked. A container with no runner registered -- one stuck failing to register, say -- is restarted, unless no container in its pool matches a runner either, which would mean the matching is broken.
  • scale <name> <count> resizes a pool pools.conf declares in place, instead of tearing it down and starting fresh like restart does. Growing -- or bringing up a pool with nothing running -- is just up with the new count, so it never refuses; there's nothing running yet that scaling up could hurt. Shrinking gets the same busy check as stop and restart, because docker compose up --scale down can't be told which containers to kill, and a busy one dying cancels its job. --force skips that check. Once the pool is resized, <count> is written into its pools.conf line -- nothing else in the file changes -- so a later start or restart brings it back at the new size instead of undoing the scale. A refused or failed scale leaves pools.conf alone.
  • declare <name> <owner/repo> [count] adds a line for a pool pools.conf doesn't have yet -- 1 runner unless [count] says otherwise, and the default label and limits -- so start, restart and scale can bring it up. It starts nothing, and refuses a name that's already declared.
  • sync goes the other way from start: it rewrites pools.conf to describe the pools on this host. Each declared pool's count becomes the number of containers it has (running or not -- its compose scale), and each gh-runner-* project the file doesn't declare gets a line, with the repo, extra label and mem/pids limits read off its containers, so a later start or restart brings it back the same. Declared pools with no containers, comments and every other column are left alone, and the old file is kept as pools.conf.bak. --dry-run prints the changes without writing them. It only asks docker, never GitHub.
  • list --json prints one JSON array with every pool in pools.conf plus every gh-runner-* project running here that pools.conf doesn't declare ("managed": false) -- pools started by hand, which stopRunners.sh never touches. Its runners counts only the GitHub runners registered by that pool's own containers (matched on the container ID in each runner's name), so two pools serving one repo -- a CI pool and a dedicated issue-worm pool, say -- are told apart. runners is null when GitHub couldn't be asked. members has one entry per container: its docker state and status, and the GitHub runner it registered (null if none, or if GitHub couldn't be asked). list --json issue-worm limits it to the named pools: every pool costs a GitHub API call, so a whole-host listing with a dozen pools takes tens of seconds, which matters for anything that polls.
[{"name":"issue-worm","project":"gh-runner-issue-worm","repo":"leonarduk/issue-worm-pro",
  "managed":true,"desired":2,"label":"issue-worm",
  "containers":{"total":2,"running":2},"runners":{"online":2,"busy":0},
  "members":[{"container":"gh-runner-issue-worm-runner-1","id":"f4c4835f30e8",
    "state":"running","status":"Up 8 hours",
    "runner":{"name":"bedroom-f4c4835f30e8-1","status":"online","busy":false}}]}]

Runners registered by other hosts serving the same repo are not counted: a host can only see and control its own containers.

stop, restart, restart-runner, a shrinking scale and list --json read repos/<owner>/<repo>/actions/runners through the host's own gh login, which needs admin access to the repo -- a classic token with repo, or a fine-grained one with Administration: read. Without it, the busy checks refuse and list --json reports "runners": null.

tests/pools_test.sh exercises start/stop/restart/restart-runner/scale/sync/list --json/autoscale against stub docker and gh commands, so it needs neither a Docker daemon nor GitHub.

Autoscaling a host's pools

A fixed count is always wrong for somebody: too many runners holding memory while nothing is queued, or too few while jobs wait behind each other. ./pools.sh autoscale resizes the pools listed in autoscale.conf to what is actually queued:

cp autoscale.conf.example autoscale.conf   # gitignored; one "<name> <min> <max> [idle_minutes]" line per pool
./pools.sh autoscale --once --dry-run      # what it would do right now, changing nothing
./pools.sh autoscale                       # every 60s until stopped; --interval to change that

Every pass, for each listed pool:

  • Demand is the jobs queued in its repo -- in queued runs, and in runs already in progress with jobs still waiting -- whose runs-on labels this pool's runners carry, including this host's RUNNER_HOST_LABEL. A job more than one pools.conf pool could take counts against the one with the fewest labels, so a pool reserved by an extra label (issue-worm) doesn't grow for the generic jobs its repo's plain CI pool is there for.
  • Growing: to busy runners plus queued jobs, capped at <max>, when that is more containers than the pool has, and never below <min>. Containers still starting count as capacity, so a job that stays queued while its runner boots doesn't grow the pool a second time.
  • Shrinking: to <min> once the pool has had no busy runner and no queued job for [idle_minutes] (10 by default), and only after the same busy check stop makes. The idle clock is kept in .autoscale-state, so it survives from one pass to the next. A pool with any busy runner is not shrunk at all, because docker compose up --scale can't be told which containers to remove.
  • A pool whose repo GitHub can't be asked about is left as it is.

It resizes with the same docker compose up --scale as scale, and never writes pools.conf: that file's count is still what start and restart use, and the next pass corrects it. Each pass costs, per autoscaled repo, one call for its runners, one per page of queued and in-progress runs, and one per such run for its jobs -- so a handful a minute for a quiet repo, but a repo with 20 runs in flight costs over 20 a minute. A token gets 5,000 an hour; raise --interval if several busy repos are autoscaled.

Two hosts serving the same repo each see the same queued job, so each may add a runner for it; the spare one sits idle and goes again after [idle_minutes]. Nothing here sizes memory or CPU yet: a pool still gets its pools.conf [mem]/[pids] limits. The native Windows fleet has the same thing as windows-pools.ps1 autoscale; see windows/README.md.

Keep it running the way you keep anything else on the host running -- a systemd unit, or at the least:

nohup ./pools.sh autoscale >> autoscale.log 2>&1 &

[label], [mem] and [pids] are optional, trailing, and positional -- pass - for one you want to leave at its default so a later one still lands in the right slot. They let a pool carry an extra runner label and its own mem_limit/pids_limit, separate from every other pool on the host; see Dedicated pools for a specific kind of job.

You need the PAT described next.

Host prerequisites

Checklist first, story below for whichever line actually bites:

  • Docker installed, engine set to start on login
  • Enough memory allocated in Docker Desktop for however many containers you run
  • Sleep disabled on this host while it serves runners (both idle timeout and lid-close -- see below, this one is sneaky)
  • GitHub CLI installed and authenticated
  • On Windows, scripts are invoked via Git Bash, not raw bash (which may silently be WSL)
  • On Windows, an execution policy that runs unsigned local scripts, if you use the .ps1 entry points (startRunners.ps1, stopRunners.ps1) -- see below

The pool is only as reliable as the machine under it, and two of these are the kind of thing you debug for an afternoon before suspecting them.

Windows execution policy. Nothing in this repo is code-signed, so a host on AllSigned refuses every .ps1 here before any of this code runs, naming the file rather than the policy: "File ...\startRunners.ps1 cannot be loaded. The file ... is not digitally signed." It reads like a broken script and is not. Get-ExecutionPolicy -List shows which scope is in force -- typically AllSigned on LocalMachine with CurrentUser left Undefined -- and Set-ExecutionPolicy -Scope CurrentUser -ExecutionPolicy RemoteSigned fixes it for your account without admin rights or weakening the machine-wide setting. It affects only the PowerShell entry points: the Git Bash path (startRunners.sh, pools.sh) is untouched by this, which is worth remembering when one fleet starts and the other refuses. Full detail, including why RemoteSigned rather than Bypass, is in windows/README.md.

Docker. Docker Desktop on Windows or macOS, Docker Engine on Linux. On Windows use the WSL2 backend; the Hyper-V backend works but is slower at the bind mounts this uses.

Set the engine to start on login — Docker Desktop → Settings → General → Start Docker Desktop when you sign in. restart: always brings a pool back after a reboot, but only once the engine is running. Without it the jobs simply queue with no runner, which looks like CI is broken rather than a stopped Docker engine.

Give it enough memory. Each runner declares mem_limit: 1g by default, so N runners across all your pools can ask for N GiB, against whatever ceiling Docker Desktop is set to in Settings → Resources. Over-committing does not error — it shows up as jobs mysteriously crawling when several repos build at once. Count the containers, not the pools — and a pool with a raised mem_limit (see Dedicated pools for a specific kind of job) counts for more than one container's worth per replica.

Sleep. This one silently cancels jobs. A host that suspends mid-job stops the runner's heartbeat, and GitHub cancels the job server-side. The signature is unmistakable once you know it, and it shows up two different ways depending on whether a job had been claimed yet when the host went under.

A job already running dies about ten minutes after the host goes under — that interval is GitHub's own patience with a runner that has stopped answering, not any timeout on your machine, so it is the same ten minutes whatever sent the host to sleep. Do not read it as pointing at a ten-minute power setting; that coincidence cost us an afternoon here. The tell is that its logs keep going past its own recorded completion: the container had the step suspended, not finished, and uploads the rest on wake into a job record that is already closed. One observed here completed at 13:42:49Z with log lines running to 14:47:09Z — 64 minutes after it supposedly ended. A job that finished before its own logs were written is not a race condition; it is a sleeping host, and it is a cheap thing to grep for.

A job still queued simply is not claimed, because no runner is awake to take it. It then starts at the moment of wake and runs at completely normal speed — in the same incident, a run queued at 13:32:45Z started its jobs at 14:47:26Z and finished them in 8 to 109 seconds, all green. Nothing was slow; the machine was absent. This one is easy to misread as contention, which is what makes the wake timestamp worth checking: jobs across different repositories resuming within the same few seconds is one host waking, not several flakes.

Either way it is intermittent — jobs shorter than the outage finish fine — so it reads as a flaky test suite rather than a host problem.

There are two separate ways a host goes under, and fixing one does nothing about the other. On Windows both are recorded: Kernel-Power event 42 carries a reason, where 7 is an idle timeout and 0 is the lid or the power button. Check which you actually have before fixing anything.

# Which kind of sleep, and when -- reason 7 = idle, reason 0 = lid/button
Get-WinEvent -FilterHashtable @{LogName='System';
  ProviderName='Microsoft-Windows-Kernel-Power'; Id=42} |
  ForEach-Object { '{0:u} reason={1}' -f $_.TimeCreated, $_.Properties[2].Value }

# Idle timeout: STANDBYIDLE's AC index, in seconds
powercfg /query SCHEME_CURRENT SUB_SLEEP
powercfg /change standby-timeout-ac 0

Closing the lid is the one that catches people, because a laptop on a desk gets shut without anyone thinking of it as powering the machine down, and no idle timeout protects against it. Worse, the setting that governs it is hidden from the Power Options UI on some machines, so it cannot be found by clicking through Windows settings — it has to be set by GUID, elevated:

# Lid close -> do nothing, on AC. SUB_BUTTONS / LIDACTION.
powercfg /setacvalueindex SCHEME_CURRENT `
  4f971e89-eebd-4455-a8de-9e59040e7347 5ca83367-6e45-459f-a27b-476b1d01c936 0
powercfg /setactive SCHEME_CURRENT

That binds to the active power scheme only, so switching schemes silently reverts it. And a closed lid under sustained build load is a real thermal question on a laptop — pair it with the machine being docked and ventilated rather than setting it blind.

# macOS
sudo pmset -c sleep 0
# Linux -- or set logind's IdleAction to ignore
systemd-inhibit --what=idle:sleep --who=gh-runner --why=CI sleep infinity &

Disabling sleep on battery too is usually the wrong trade on a laptop; if the machine is a CI host on AC, standby-timeout-ac 0 is enough. Display sleep is harmless — it is system standby that kills jobs.

Registering a runner by hand, without any of the tooling in this repo, is documented directly by GitHub — Settings -> Actions -> Runners -> New self-hosted runner on the target repo. Everything below is the same registration, done by entrypoint.sh inside a disposable container instead of by hand on a long-lived machine.

GitHub CLI. pools.sh, startRunners.sh/stopRunners.sh and discoverPools.sh all shell out to gh on the host to list repos and check which runners are actually online -- a separate installation from the gh baked into the runner container, which only helps workflows running inside a job. Install it and authenticate before using any of these scripts:

winget install --id GitHub.cli   # Windows; see https://cli.github.com for other platforms
gh auth login                    # in a new shell, so PATH picks up the install

On Windows, invoke every script here with an explicit path to Git Bash rather than bare bash -- if WSL is installed (likely, since Docker Desktop's recommended backend needs it), bash on PATH resolves to C:\WINDOWS\system32\bash.exe, WSL's launcher into a separate Linux environment with its own PATH. gh installed on Windows via winget is invisible there, so a script fails with a plain gh: command not found that has nothing to do with whether gh is actually installed. Confirm which one bash means with Get-Command bash, and if it points into system32 or WindowsApps, use the real path instead:

& 'C:\Program Files\Git\bin\bash.exe' ./discoverPools.sh leonarduk

Without it, these scripts fail fast with gh: command not found rather than silently doing nothing -- but on Windows that error can end up wherever stderr goes for whatever launched bash, so it is easy to miss if you are not looking for it.

Setup

You need a personal access token that can register runners:

  • Classic PAT — repo scope, or
  • Fine-grained PAT — the target repositories, Administration: read & write

One PAT covering several repositories can serve several pools. The container exchanges it for a short-lived registration token at start-up; the PAT itself is never baked into the image.

Quick start already walked through this via pools.sh -- what follows is the same thing one level down, with the raw docker compose command pools.sh up runs for you, for when something needs debugging directly.

git clone https://github.com/leonarduk/local-github-runner
cd local-github-runner
printf '%s' 'ghp_your_token_here' > pat.secret   # gitignored
cp .env.example .env                             # set RUNNER_HOST_LABEL
GITHUB_REPOSITORY=owner/repo docker compose up -d --build
docker compose logs -f

A healthy start looks like:

runner-entrypoint: requesting a registration token for owner/repo
runner-entrypoint: configuring bedroom-4f2c1a9b3e77-1
runner-entrypoint: waiting for a job

The runner then appears under the repo's Settings → Actions → Runners as idle. compose.yaml declares two of them; override that for a one-off with --scale runner=3.

GITHUB_REPOSITORY has no default and the up fails without it. That is deliberate: it used to default to the one repo this was written for, which is exactly the kind of thing that survives a copy to a new machine and quietly registers runners against the wrong repository.

Check the pool is really up

docker compose up returning 0 means the containers started, not that any runner registered. Registration happens later, inside the container, and can fail on its own -- a PAT without the right scope, a typo in GITHUB_REPOSITORY, a repo the token cannot see -- leaving containers that restart forever while GitHub shows nothing. Ask GitHub rather than Docker:

gh api repos/OWNER/REPO/actions/runners \
  --jq '.runners[] | "\(.name)  \(.status)  busy=\(.busy)"'

Expect one online line per container, each prefixed with the host label from .env. Fewer than you scaled to means some are still registering or some are failing; docker compose logs -f says which. Runners listed offline with no container behind them are strays from a container that was killed rather than stopped -- clear them with gh api -X DELETE repos/OWNER/REPO/actions/runners/ID.

Several repos from one checkout

The compose project name is the only thing separating one pool from another. Two pools sharing it are the same pool — an up for the second repo silently reconfigures the first repo's containers rather than adding to them. So set it alongside the repository:

GITHUB_REPOSITORY=leonarduk/jobtrack COMPOSE_PROJECT_NAME=gh-runner-jobtrack docker compose up -d

Only put per-machine facts in .env — RUNNER_HOST_LABEL, and GITHUB_REPOSITORY if this host serves exactly one repo. Every pool started from this directory reads that same file, so anything that differs per pool belongs on the command line.

Check what is running, across all pools, with docker ps --filter name=gh-runner. Docker Desktop groups them the same way, by compose project -- each expandable row is one pool, runner-1/runner-2 its containers (see the screenshot in Quick start step 5).

Several machines

Nothing is host-specific except two gitignored files, so a second machine is a clone plus those:

  1. git clone this repo.
  2. Write the PAT to pat.secret — it never travels through git.
  3. cp .env.example .env and set RUNNER_HOST_LABEL to that machine's name, e.g. bedroom.
  4. Bring up whichever pools that machine should serve. Two ways:
    • cp pools.conf.example pools.conf, edit it to list the repos this host serves, then ./startRunners.sh.
    • Or ./discoverPools.sh <owner>, which needs no local config at all -- it finds every private repo under <owner> carrying a .local_runner marker file and brings each one up. Either way, ./pools.sh up <name> <owner/repo> [count] remains the one-off, no-config primitive underneath both.

RUNNER_HOST_LABEL becomes both a runner label and the runner-name prefix, so the GitHub runner list and gh api .../actions/runners say which box a runner is on — container hostnames are random hex and no help. It also lets a workflow pin a job to one machine with runs-on: [self-hosted, bedroom].

Runners on different machines serving the same repo simply join the same pool: GitHub hands each queued job to whichever is free. There is no coordination between hosts and none is needed -- the screenshot in Quick start step 5 is one repo with runners from two machines, bedroom and steves-big-laptop, sitting in the same list.

The image is built per machine (--build). There is no registry involved, so a second host costs one ~1.4GB build rather than any shared infrastructure.

Point a workflow at it

A runner sits idle until a job asks for it. In the target repo's workflow:

-    runs-on: ubuntu-latest
+    runs-on: [self-hosted, linux, x64]

Every job in the workflow needs the label, or the untouched ones keep exhausting the account's Actions minutes as before. Keep the ubuntu-latest line commented directly above each replacement, so reverting to hosted runners is a one-line edit at the point of use rather than an archaeology exercise.

Running the same tests on Windows instead (or as well)? windows/README.md covers runs-on: [self-hosted, windows, x64], set up the same way but without a container.

Last verified: 2026-09-05. Re-run and update this date when making changes that could affect runner behavior.

Verified end to end against leonarduk/spring-professional-udemy-practice-tests: both its jobs ran on a pool from this image and passed, in 7s and 32s, having previously failed in about 2s without starting.

GitHub's own reference for this syntax and the default label set is Choosing the runner for a job.

Dedicated pools for a specific kind of job

Every pool so far shares one shape: a repo's normal CI, where a runner picks up a job, runs it in minutes, and frees up again. Some jobs do not behave like that -- a long-running pipeline that itself opens pull requests is the one that surfaced this. issue-worm-pro's pipeline clones a repo, builds a venv and runs pytest inside a container job, which already needs more than the 1g/512-pid defaults below are sized for. Worse, the PR it opens needs runners of its own, for python-ci.yml and the review workflows, and a Changes Requested requeue loop waits on those. Point enough concurrent worm jobs at the same pool that serves that repo's CI, and they can occupy every runner while the CI they are blocked on queues behind them -- a self-inflicted deadlock, not an external one.

The fix is a separate pool with its own label, never the CI pool: worm jobs run where python-ci.yml and the review workflows cannot be queued behind them, because they are physically on different runners.

RUNNER_EXTRA_LABELS (compose.yaml) is what makes a pool reachable that way. It appends one or more comma-separated labels after the usual self-hosted,linux,x64,docker,<host> set, so a workflow can target runs-on: [self-hosted, issue-worm] and land only on pools carrying that label -- CI's runs-on: [self-hosted, linux, x64] never matches it, and nothing else needs to change. Unset, it reproduces today's label set exactly -- no trailing comma, nothing appended.

POOL_MEM_LIMIT / POOL_PIDS_LIMIT (compose.yaml) override that pool's mem_limit / pids_limit away from the 1g/512 defaults, which a clone + venv + pytest build is liable to exceed. Unset, both stay at today's values.

All three are per-pool, so they are set the same way as [count]: as trailing, optional, positional arguments to ./pools.sh up/reset, or as the fourth through sixth columns of a pools.conf line (- skips one to reach a later column). They are deliberately not .env settings -- .env holds facts about the machine, and these differ per pool on the same machine.

# one-off, via pools.sh
./pools.sh up issue-worm leonarduk/issue-worm-pro 2 issue-worm 2g 1024

# or in pools.conf, alongside the plain CI pool for the same repo
# name                  owner/repo                     count  label        mem   pids
issue-worm-pro          leonarduk/issue-worm-pro        4
issue-worm              leonarduk/issue-worm-pro        2     issue-worm   2g    1024

Then point the worm workflow(s) at the label, leaving every other job in that repo (and every other repo's CI) targeting the plain pool:

-    runs-on: ubuntu-latest
+    runs-on: [self-hosted, issue-worm]

This mitigates the starvation risk; it does not eliminate queuing outright. Two pools drawing from the same finite host memory still compete for it, and undersizing either pool's [count] just moves the queue rather than removing it -- size each pool to the concurrency it actually needs, the same way as any other pool (see Host prerequisites, "Give it enough memory").

What the image provides

Base ubuntu:24.04
Runner v2.337.0, SHA256-verified at build, --disableupdate
Arch linux/amd64 and linux/arm64 (via TARGETARCH)
Tooling git, curl, jq, tar, unzip, zstd, sudo, python3 3.12, uuidgen, pkill
gh v2.100.0, SHA256-verified at build — preinstalled on GitHub-hosted images, so workflows assume it
JS actions bundled Node 20 and Node 24, verified working
Tool cache /opt/hostedtoolcache, writable — actions/setup-python needs this

Two deliberate choices worth knowing about, because both look like oversights:

  • sudo is installed, passwordless. Workflows install tools with sudo mv, matching what a GitHub-hosted runner allows. Removing sudo, or setting no-new-privileges in compose, breaks those steps.
  • gh is baked in rather than installed per job. Ubuntu's own package is too old for the --json flags workflows use, so repos had started curling the release tarball at job time without verifying it. Installing it here, checksummed, removes that. Workflows that guard on gh already being on PATH become no-ops; the binary carries no credentials, so each workflow still supplies its own GH_TOKEN.
  • RUNNER_MANUALLY_TRAP_SIG=1. Makes the runner handle SIGTERM itself and finish or cancel the running job cleanly. Without it, docker stop mid-job leaves the job hung until GitHub times it out.

Maintenance

Bumping the runner version means editing RUNNER_VERSION and the two checksums in the Dockerfile together — they are per-version, and a mismatched pair fails the build rather than silently skipping verification. The SHAs are published in each release's notes at https://github.com/actions/runner/releases.

--disableupdate keeps the pinned version honest. If GitHub starts refusing jobs from a version this old, drop that flag and rebuild the image to match.

A version bump has to reach every machine, since each builds its own image. Nothing enforces that; docker images local-github-runner:local on each host is how you tell.

Teardown

COMPOSE_PROJECT_NAME=gh-runner-repo docker compose down

Each container deregisters itself on the way out, so nothing should be left behind — a down of a two-runner pool completes in about five seconds with both registrations gone.

That relies on entrypoint.sh signalling Runner.Listener directly rather than only run.sh: a signal to the wrapper alone never reaches the process RUNNER_MANUALLY_TRAP_SIG=1 arms, and the container is then SIGKILLed with deregister() unreached. If a container is killed outright anyway — docker kill, a host crash — its runner lingers as "offline" under Settings → Actions → Runners and needs removing there, or with gh api -X DELETE repos/OWNER/REPO/actions/runners/ID. The container itself recovers unaided: the leftover .runner from its previous life is cleared at start-up before it reconfigures.

Known limits

  • No Docker-in-Docker. Adding it means mounting the host's Docker socket, which hands any job root on the host — do not do that on the strength of this README alone.
  • actions/cache has no backing store, so cache steps are no-ops that cost a little time. Worth knowing if a workflow starts depending on a warm cache.
  • pip install only works after actions/setup-python. The base is Ubuntu 24.04, whose system interpreter refuses installs under PEP 668 (externally-managed-environment). setup-python puts its own interpreter on PATH first — but a workflow that drops that step keeps working on a GitHub-hosted runner and fails here.
  • mem_limit: 1g / pids_limit: 512 are conservative defaults. If a build is OOM-killed, raise them per pool with POOL_MEM_LIMIT / POOL_PIDS_LIMIT rather than removing the limits outright -- see Dedicated pools for a specific kind of job.
  • This container image is Linux only. A workflow pinned to windows-* will never match it -- see windows/README.md for the separate, non-containerised path that does. macos-* is not covered by anything in this repo.
  • Pool names use the repo name, not owner/repo. Two repos of the same name under different owners would collide.

Further reading

Everything above is GitHub's own supported mechanism -- a self-hosted runner registered the normal way -- wrapped in a container and a compose file for convenience. None of it depends on this repo; every piece is documented independently, and it is worth reading the source rather than taking this README's word for the security-critical parts:

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages