Skip to content

Repository files navigation

Basewright

Automated provisioning of database instances on servers that already exist.

CI License Python Status

Why · How it works · The one rule · Quickstart · Operating it · Engines · Writing a profile · Runbook · Contributing


Why Basewright

An install that finished is not the same thing as an instance fit for production.

Package managers already install databases. The hard part was never apt install postgresql-16. The hard part is deciding whether this machine can carry this engine at all, sizing memory and storage for the hardware in front of you, doing it identically on the fortieth server as on the first, and being able to say six months later why shared_buffers is 8 GB on that host.

Basewright treats those four things as the product and the installation itself as a mechanical step at the end.

The workflow it replaces is a familiar one: the infrastructure team builds a VM and hands over an IP; someone from the database team then logs in, inspects the machine by hand, decides which engine version fits, and installs it — differently every time, depending on who did it and how busy they were. Nothing afterwards records what was decided or why. Reproducing the environment means reading someone's shell history.

A provisioning job that exits zero tells you a package was installed. It tells you nothing about whether the instance is sane. Basewright refuses to conflate the two: a host that fails preflight is reported as refused, with a reason, never as a degraded partial install.

An install that finished, contrasted with an instance fit for production

How it works

Five steps, always in this order, each separately runnable:

gather, preflight, plan, apply, verify — only apply changes the target
Step What it does Changes the target?
gather Collects facts: OS, CPU, RAM, disks, mounts, ports, existing engines No
preflight Runs the gate rules against facts and request; PASS / WARN / BLOCK No
plan Renders the full intended end state, every value annotated with its rule No
apply Executes the plan, idempotently Yes
verify Reads the live instance back and compares it to the plan No

plan is the centre of the product, and it is a file. It can be reviewed by a second person, committed to Git, attached to a change request, and diffed against the plan from three months ago. A provisioning tool where the reasoning lives only in the operator's head is the thing being replaced.

The artifact exists and its contract is frozen. A change to it is a version rather than a patch, and it is now at version two: writing apply against the plan alone found the one thing it did not carry, which was how the instance gets created (ADR-0022). Here is what it carries, and which step reads each part:

The sections of plan.json, what each carries, and which step reads it

And here is one, rendered for the person who has to approve it. Every value carries the rule that produced it and the reasoning that rule ships with:

basewright plan --from test/golden/postgresql/plan/typical.json
A rendered plan: request, host, preflight, parameters with their reasons, layout, initialization, changes, secrets and the verdict

Three properties make that artifact worth having:

  • A block produces no plan. It produces a refusal naming the failing rule, the observed value, the required value, and what would have to change. There is no partial plan and no --force.
  • Warnings must be acknowledged explicitly before apply will run. Silent warnings become invisible within a month.
  • The plan is deterministic. The same facts and the same profiles produce byte-identical output, which is what makes golden-file review of tuning decisions possible.

verify closes the loop: apply promised something, verify proves it, and it can be run again months later against any host Basewright provisioned. A verify failure is loud, because it means the instance is no longer what the documentation claims it is.

The one rule

Core logic never branches on an engine name. Engines are data.

There is no if engine == "postgresql" in the planner, the gate engine or the reporter. Everything engine-specific lives in a profile: a directory of declarative files plus one thin apply role.

The eight files of a profile, what each declares, and which step reads it

Adding an engine means adding a directory. It never means editing the core. If the core would need to know the engine name to behave correctly, the missing information belongs in the profile schema instead — the schema gets extended, not the conditionals.

The rule is enforced rather than merely intended. CI scans every line of basewright/, comments and docstrings included, and fails on an engine name. Profiles are validated against a JSON Schema in which an unknown key is an error, so a profile cannot smuggle in behaviour the core does not understand. A second engine ships specifically to prove the abstraction holds, because one engine proves nothing.

Every sizing rule carries its own justification, and that justification is rendered into the plan next to the computed value:

- id: pg.shared_buffers
  parameter: shared_buffers
  expr: "0.25 * mem_total"
  min: "128MB"
  max: "8GB"
  why: "25% of RAM is the standard starting point; capped at 8GB because beyond that
        the OS page cache is the better place for the memory."

A number without a reason is exactly the situation Basewright exists to end. Here is one real value, from the rule that produced it to the line it occupies in a plan:

A cache size, from its expression through rounding and bounds to the plan

Quickstart

The whole loop, and CI runs it on every pull request. A bare ubuntu:24.04 container is read, gated, planned for, provisioned from its own plan, provisioned again with nothing left to do, and then asked whether it is what the plan promised.

Every command below is copy-pasteable after make install, runs against documents committed under test/, and is the command that produced the picture underneath it: tools/render_assets.py runs them to make the images, and a test holds the text in this file against the commands it ran. Neither can drift from the other.

gather reads what a host reported and normalizes it into the model every rule is written against:

basewright gather --facts test/fixtures/hosts/typical.json
basewright gather summarising a host from a facts document

Where that document comes from is a playbook. ansible/playbooks/gather.yml reaches the host, reads it, writes the document, and then hands it straight back to the verb above:

ansible-playbook ansible/playbooks/gather.yml -l db-01.example.invalid

The host below is a real one — a container the role went and read during the test run that produced this page, committed as it came off the wire. It is not a fixture somebody wrote to make a point:

basewright gather --facts test/fixtures/hosts/collected.json
basewright gather summarising a host the collecting playbook actually read

The playbook is the entry point and the CLI never reaches a machine — not basewright gather --host db-01 with Ansible underneath it, which is what most tools in this space do:

Semaphore runs the playbook, the playbook reads the host and writes facts.json, and the CLI reads that document

That split is the reason the deciding half stays a pure function of a document: it has no connection handling, no inventory, no host key policy, and its tests need no network. The argument, including the case for the obvious alternative, is in ADR-0020.

The playbook takes one optional input, and it is the only thing it is ever told about what is being provisioned. Nineteen of the twenty shared rules ask about the machine alone; the twentieth asks whether the host can reach the repository its packages would come from, and where those come from is written in a profile. So the collector is told which one, and uses it for that fact and nothing else (ADR-0021):

ansible-playbook ansible/playbooks/gather.yml -l db-01.example.invalid -e gather_engine=postgresql
The three answers reachable_repositories gives, and the verdict each one produces

Absent is not empty, and the difference is the whole design. A host nobody asked reports nothing and the rule skips; a host that was asked and reached nothing reports an empty list, and that is a refusal — this one on a container the role really did read:

basewright preflight --facts test/fixtures/hosts/asked.json --profile test/fixtures/profiles/exampledb
basewright preflight refusing a host that cannot reach the repository its packages come from

preflight puts twenty engine-independent rules, and every rule the profile adds, to that host. A refusal is a first-class outcome, so it is an answer rather than an error: it names the rule, what was found, what was required, and what would have to change.

basewright preflight --facts test/fixtures/hosts/crowded.json --engine postgresql
basewright preflight refusing a host, naming four blocking rules and what each one found

There is no flag that turns any of that into a plan. A host that passes still reports what it is not happy about, and those warnings are acknowledged before apply will run:

basewright preflight --facts test/fixtures/hosts/typical.json --engine postgresql
basewright preflight passing a host, with four warnings to acknowledge

plan runs those rules again, refuses outright if any of them blocks, and otherwise sizes every parameter, resolves the layout and works out everything apply would do. The rendering above is what it prints; --json writes the artifact itself.

A plan is named after a digest of its own content, which makes the name a checksum as well as a name. plan --from reads one back — which is how the person who applies a plan can be somebody other than the person who produced it — and says so when the two no longer agree:

basewright plan --from test/fixtures/plan/edited.json
basewright plan refusing a plan whose id no longer matches its content

apply is not a verb of this CLI at all, and never will be: applying is Ansible's job, and the plan is the boundary between them. It is a playbook, and it takes one input:

ansible-playbook ansible/playbooks/apply.yml -l db-01.example.invalid   -e basewright_plan_file=plan.json -e basewright_accept_warnings=true

Before it touches anything, four things have to be true: the plan is a plan and this build understands its contract, it has not been edited since it was approved, the host is still the machine it was built from, and its warnings have been acknowledged by somebody. Then the phases run, in the order the plan lists its own changes in:

The phases of applying a plan, in order, and which half of the tool runs each

That picture is drawn from the playbook rather than written beside it, so a phase added or reordered redraws it. The two halves interleave because the order belongs to the plan rather than to the roles — a profile whose vendor package makes the service account cannot have its directories owned before the packages are installed, which is a thing a real container found rather than a thing anybody predicted.

Applying is proved in CI rather than pictured, because a package version and a moment are not things a byte-for-byte image can carry. molecule test -s apply takes a bare Ubuntu container, reads it, plans for it, applies the plan, applies it again and finds nothing to do — and then runs the verb below against it.

verify closes the loop. It is a playbook too, and for the same reason: reading a live instance runs over SSH, so an engine's role asks the instance eleven questions and writes down what it said, and the CLI compares that with the plan (ADR-0024):

ansible-playbook ansible/playbooks/verify.yml -l db-01.example.invalid -e basewright_plan_file=plan.json
Semaphore runs the playbook, an engine's role reads the instance and writes observation.json, and the CLI judges that against the plan

The document that role writes is an artifact like the plan is, so judging it is a step of its own that reaches nothing. The reading below came off a container this repository's test run really provisioned, committed as it came back:

basewright verify --plan test/fixtures/plan/applied.json --observed test/fixtures/observations/observed.json
basewright verify proving that a running instance is what its plan promised

A verify that only ever passes has not been shown to be looking. So the scenario changes a parameter on the running instance behind the plan's back and insists the tool goes red — and this is that report, from the reading taken after the change:

basewright verify --plan test/fixtures/plan/applied.json --observed test/fixtures/observations/changed.json
basewright verify refusing an instance whose parameters no longer match its plan

Every check reports one of three things, and the third is the one worth spelling out. A check nobody managed to put to the instance is not a pass: it is reported as its own outcome, and the run does not verify the instance (ADR-0025). A run that proved nothing and exited zero would be the failure this whole tool exists to end, moved one step later.

Kind What it reads back
service The unit the plan named is active, and enabled so it survives a reboot
port The instance is on the planned port and on no other
connection An authenticated connection is accepted
version The running server is the major version the plan asked for
parameters Every sized parameter reads back with the planned value
paths The instance stores its data and its write-ahead log where the plan put them
log There is a log, and it has been written to since the service started
backup The service account can write where the plan puts backups
auth No password-free rule is reachable from off the machine
account Every account that can log in has a password set
initialization The locale, encoding and checksum flag it was created with

The last one is the one that cannot be put right afterwards: those three are fixed when the instance is created, and changing any of them means dumping every database in it and reloading. The plan states them (ADR-0022) and this is what reads them back.

A profile can narrow a kind without the core learning anything about the engine. The port kind judges the port, because the port is what the plan carries; that this instance is bound to no address but loopback is profiles/postgresql/'s own decision, written as an expression in the file where somebody arguing with it would look.

How it is operated

Semaphore is the interface and there is no other one (ADR-0005). Four templates, one per verb, and deploy/semaphore/ ships their definitions as JSON with a setup guide beside them:

The four Semaphore templates, the survey fields each one asks for, and the playbook each one runs

That picture is drawn from the definitions rather than written beside them, so a survey field added to a template arrives in the documentation with it.

Two of the four rows are worth reading twice, because they are where the design shows.

Apply asks for a plan id and no host. A plan is built from one machine's facts and refuses to be applied to another, so the host is in the artifact already; asking for it again would be asking somebody to repeat something the plan carries, and to be the one who gets it wrong. Verify asks for a host and an instance and no plan id, because somebody asking whether a database is what it should be knows which database they mean and has no reason to know the id of a plan somebody else approved.

What travels between Plan and Apply is the id, and the id is a digest of the plan's own content — so it is a name and a checksum at once. The person who produces a plan and the person who applies it can be different people, on different days, which §12 of the brief calls a feature rather than an accident. Plans are kept in a durable store, addressed by that id:

basewright plan --plan-id 000000000000 --store test/fixtures/store
basewright plan refusing a plan id the store has not got, and listing what it does hold

The store keeps one more thing and deliberately not a third: a single line per instance naming the plan it was last applied from, which is what lets Verify take a host and an instance. It is overwritten every time, with no history and no who. An audit trail is a later phase, and half of one would look like a record while answering none of the questions a record is asked.

docs/dev/runbook.md is what to do when one of these goes red.

All five verbs exist:

basewright --help
basewright --help listing the gather, preflight, plan and verify verbs

Every image above is the real console output of the command printed beside it, and none of it can fall behind: tools/render_assets.py --check regenerates every one in CI and fails the build if what is committed differs by a single byte. Progress is tracked in docs/dev/STATUS.md.

What a run exits with

Semaphore is the interface (ADR-0005), and its view of a run is one bit: the task is green or the task is red. So the exit code is not how a failure gets reported — the report already does that, in the task log. What the number carries is what the person looking at a red task is supposed to do next, and there are three answers worth telling apart:

Code What happened What to do
0 The tool ran, and the answer is yes. Go on to the next step.
2 The tool ran, and the answer is no. Read the report: it names the rule, what was found, and the way out.
64 The request itself is malformed. Fix the command. Nothing was decided, so there is no report to read.
The three exit codes, what produces each one, and what an operator does about it

The line between the two non-zero answers is where a document stopped being readable: a file that could not be read at all is 64, and a file that was read and is not acceptable is 2. A missing facts document is a mistyped path; one that fails its contract is a real answer about a real file. A plan whose id no longer matches its content is 2 for the same reason — it is a plan, it is simply not the plan it claims to be.

A blocked gate, an unacceptable document and a verify mismatch are all 2. They differ in what happened, not in what to do about it, and the difference between them lives in the report rather than in a number. The set is closed, held by a test in both directions, and argued in ADR-0019.

There were four. 69 said the verb existed and was not built, ADR-0019 named it as the one member with an expiry date, and it went when verify landed. A set that loses a member is narrower than it was rather than different: nothing that read 0, 2 or 64 has to change, and there is now no invocation of this tool that can produce anything else.

What a host is, to a rule

Facts are normalized before anything reads them, so a change in how they are collected cannot ripple into the gate engine or the planner. The model carries what the rules need in order to reach a verdict, and nothing else:

The facts a host is described by, and which rule reads each one

The collector reports which packaging family the operating system belongs to; the core never infers it. Working out that one distribution is packaged like another is knowledge about operating systems, and it belongs where the observation is made — the same reason no engine name appears in the core.

Supported engines

Engine OS families Versions Status
PostgreSQL Debian 12, Ubuntu 22.04 / 24.04 16, 17; 15 warns, and not on 24.04 the whole loop
PostgreSQL RHEL / Rocky — planned
MySQL/MariaDB Debian / Ubuntu — planned
SQL Server Windows — planned

profiles/postgresql/ is eight declarative files and one template directory. It gates a host, sizes seventeen parameters, produces a complete plan, applies it through ansible/roles/postgresql/, and then reads the running instance back and proves it is what the plan promised. That role is the one directory in this repository where the word is allowed to appear. Nothing under basewright/ knows it, and a test reads every line of the core to keep that true.

The seven conventions a profile cannot avoid deciding — path layout, the service account, the locale, the authentication rules, the minimums that become blocks, the OS families and the port — are this repository's decisions rather than placeholders waiting for somebody (ADR-0026). docs/dev/STATUS.md states each one with the argument for it and the single line that changes it, because a value described as provisional is a value no reviewer argues with.

An engine nothing has a profile for is not a missing feature, it is a missing directory, and the refusal says so:

basewright preflight --facts test/fixtures/hosts/typical.json --engine mysql
basewright refusing an engine it has no profile for, and naming the ones it has

Writing a profile

The profile is the contribution surface: adding an engine, an OS family or a tuning rule is a change to profiles/, reviewable by a DBA rather than by a programmer. The guide is docs/dev/writing-a-profile.md, written alongside the schema it documents.

Checking a profile is one command, and it reports everything wrong at once — the file, the place inside it, and what to do about it:

python -m basewright.profiles test/fixtures/profiles/malformed
The profile loader refusing a profile, naming each defect and its remedy

The remedy in each of those messages is the schema's own description of the field, so the specification and the error message cannot drift apart: they are the same string. Beyond the schema, the loader checks the agreements between files that no single schema can see — that every file names the same engine, that the default version is one that exists, that every OS family used is declared and can be installed on, and that no identifier is used twice.

Repository layout

basewright/
├── basewright/                  # Python: the part that decides
│   ├── facts/                   # normalize raw facts into a typed model
│   ├── profiles/                # loader + JSON Schema validation
│   ├── preflight/               # gate engine, severity resolution
│   ├── planner/                 # sizing evaluation, layout resolution, plan assembly
│   ├── verify/                  # judging a reading against the plan it came from
│   ├── report/                  # human and JSON rendering, shared by plan and verify
│   ├── store.py                 # where a plan lives between planning and applying
│   └── cli.py                   # thin: basewright gather|preflight|plan|verify
├── ansible/                     # Ansible: the part that acts
│   ├── playbooks/               # one per verb: gather, preflight, plan, apply, verify
│   ├── roles/                   # common, then one thin role per engine
│   ├── plugins/                 # action/filter plugins bridging to the Python package
│   └── inventory/example/
├── profiles/                    # engine data — the extension point
├── schema/                      # JSON Schema for every profile file and for plan.json
├── deploy/semaphore/            # one environment, four templates, setup guide
├── test/
│   ├── unit/                    # pytest: facts, gates, sizing, rendering
│   ├── golden/                  # fixture facts → expected plan output
│   └── molecule/                # role tests in containers
└── docs/
    ├── adr/                     # architecture decisions
    └── dev/                     # writing-a-profile.md, STATUS.md, runbook.md

The design split is worth stating plainly: Ansible decides nothing, Python decides everything. Sizing arithmetic and gate evaluation in Jinja templates would be untestable and unreadable. Ansible executes a plan that has already been made.

Python decides and renders plan.json; Ansible reads it and acts

Development

make install       # install the package and development dependencies
make lint          # ruff, mypy, yamllint
make test          # pytest, unit and golden suites, with coverage
make schema        # validate every profile against the profile JSON Schema
make guard         # fail if an engine name leaks into the core or a shared role
make assets        # regenerate the diagrams and terminal captures in docs/assets/
make assets-check  # fail if a committed image is stale
make golden        # regenerate the golden plans, then read the diff
make golden-check  # fail if a committed golden plan is stale
make ansible-lint  # lint every playbook, role and scenario
make molecule      # run the role tests against real containers (slow, needs Docker)
make all           # everything CI runs on a pull request, except molecule

make molecule runs three scenarios. Two of them build a container per platform and bring systemd up inside it: one runs the collecting role against it for real, the other takes a bare container all the way to a verified instance and back over it a second time. The third is small and needs no init — it puts a secret through every store the shared role can write to, twice each, and insists on getting the same answer. Together they are slower than the rest of the suite put together, and they are the only things here that prove this code and a machine still agree, so they run on every pull request rather than on a schedule.

One command is worth knowing on its own, because it is what a profile author runs:

python -m basewright.profiles profiles/<engine>

It is not a fifth verb — the verbs act on a host, this acts on the repository.

Every image in this README is generated by tools/render_assets.py and by nothing else, in a light and a dark variant. Diagrams are drawn from the same description of the architecture the prose uses; terminal images are captured by running the command. Neither can drift, because CI regenerates both and compares them byte for byte with what is committed — a stale picture fails the build the way a stale test does.

Architecture decisions

Twenty-three decisions are recorded in docs/adr/, each with the context that forced it, what it costs, and the alternatives that were rejected. The four that shape everything else:

The twenty-seven decision records, grouped by the question each one answers
  • ADR-0001 — the plan comes before the change, and it is a file somebody can review.
  • ADR-0002 — engines are data, so adding one never edits the core.
  • ADR-0004 — two severities, and no way to override a block at run time. The most arguable of them, and the record makes the case against itself.
  • ADR-0008 — Python decides, Ansible acts, and the plan is the boundary.

The rest cover how a version is chosen (0003), why the interface is Semaphore (0005), how targets are reached (0006), how credentials are kept out of every artifact (0007) and where a generated one goes instead (0027), why every sizing rule carries its own justification (0009), what a second run is allowed to do (0010), where packages come from (0011), and the two boundaries that keep the scope finishable — 0012 and 0013.

Project status

Phase Contents Status
Foundation Schema, loader, fact model, gate engine, planner, report, CI complete
Phase A PostgreSQL on Debian/Ubuntu, end to end complete
Phase B Semaphore templates, plan storage, runbook complete
Phase C A second engine, proving the core needed no changes not started
Phase D Windows and SQL Server not started
Phase E Audit trail, plan diff against a live host, signed releases not started

Phase B as the brief writes it also contains RHEL and Rocky, and this repository does not ship them. The profile declares the Debian family because that is the family CI provisions on every pull request, and declaring another before it is tested would be a claim rather than a fact. Adding one is a support-matrix entry, a packages entry and a scenario — no core change, which is the whole point of engines being data.

Phases A and B are complete, and the loop closes. A bare ubuntu:24.04 container is read, gated, planned for, provisioned from its own plan, provisioned again with zero changes, and then asked whether it is what the plan promised — and CI does all of that on every pull request. Then the scenario changes a parameter on the running instance behind the plan's back and insists verify goes red, because a verify that only ever passes has not been shown to be looking.

Provisioning a database is a form somebody fills in, not a command somebody remembers: four Semaphore templates, plans kept in a store and found again by the id printed in a task log, and a runbook for when one of them goes red.

A generated password is written once, to a store, and never to the plan. There are two stores and the plan names neither: it carries the location a secret lives at and nothing else, which is what makes it safe to attach to a change request. file writes the password itself with mode 0600 on the control node and is the default; vault writes an ansible-vault file in the same place under a key that reaches the run from Semaphore's own secret store, and refuses to run if it was given no key. Semaphore's store was the obvious choice for the second one and it is not what shipped, because its API returns no secret values — so it could not hand back a password on the second apply of a plan, and would have had to roll a new one (ADR-0027). Getting a password back is ansible-vault view; that is the cost of the choice and docs/dev/STATUS.md says so under known gaps.

Seven conventions had to be settled before a profile could describe a real engine — path layout, service account, locale, authentication rules, the minimum resources a production instance may run on, the OS families supported, and the port convention. This repository decides all seven (ADR-0026), and docs/dev/STATUS.md states each one with the argument for it and the line that changes it.

Six of them are forced rather than chosen: the paths are where the vendor's packaging puts things, the account is the one the package creates, the encoding is the only defensible answer, the families are the ones that are tested. The minimums are the row that needed somebody willing to defend it, because they become block thresholds with no run-time override — 2 cores, 2 GB, and a floor per path.

An estate that disagrees changes a line, not a fork. That is what engines-as-data buys: every one of the seven is a single value in a YAML file under profiles/, with the reasoning for the original still written beside it. The two most likely to be changed — what the backup mount is called, and whether the write-ahead log gets a mount of its own — are named as such on the status page rather than left to be discovered.

What Basewright is not

Permanently out of scope. This list is the reason the tool stays finishable:

  • Creating VMs, networks, storage or DNS. Basewright starts at a reachable host.
  • Backup scheduling, verification or restore. That is a separate tool's job, and the tool is Fleetward — which restores a backup into a throwaway container and smoke-tests it, rather than believing a job that exited zero. One boundary and two tools: Basewright hands over an instance that matches a plan somebody approved, and Fleetward watches it for the rest of its life. Neither reaches into the other's job.
  • A monitoring stack. An exporter can be installed as an optional role; Prometheus and Grafana are not Basewright's to run.
  • Application schema deployment or data migration.
  • Managed cloud databases. RDS, Azure SQL and Cloud SQL are provisioned by their own APIs and have no server to inspect.
  • Major-version upgrades of an existing instance. Different problem, different risks.
  • A custom web UI. Semaphore already provides scheduling, RBAC, a secret store, task history and logs. Building a second one is how this project would die.

License

Apache-2.0. See LICENSE.

About

Automated provisioning of database instances on servers that already exist. Gathers facts, gates the host, renders a reviewable plan where every value carries the rule that produced it, then verifies the result.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages