Skip to content

Add a MariaDB backup and restore guide - #1093

Open
ideaship wants to merge 2 commits into
mainfrom
docs/mariadb-guide
Open

ideaship wants to merge 2 commits into
mainfrom
docs/mariadb-guide

Conversation

@ideaship

Copy link
Copy Markdown
Contributor

Adds a MariaDB backup and restore section to the operations guide, and retires the
three subsections under MariaDB in infrastructure.md that it replaces.

What it adds

Four pages under docs/guides/operations-guide/mariadb/:

  • index.md — orientation, scoped to a single-shard deployment.
  • backup.mdx — which host holds the archives, taking full and incremental
    backups, confirming a run succeeded, copying archives off the hosts, and
    scheduling with a prune script.
  • connection-fix.md — on OSISM 10.2.0 and earlier, mariabackup locks one node
    and copies another, so an archive can record a binary log position belonging to a
    different node and restore cleanly to the wrong state. How to tell whether a
    deployment is exposed, the overlay that corrects it, and what to do about
    archives taken before it.
  • restore.md — the six-step procedure, including restoring from a full plus one
    increment, and reconciling against OVN, Ceph and the hypervisors afterwards.

The old sections gave both mariadb-recovery commands but delegated the restore
itself to the kolla-ansible manual, which documents the archive layout OSISM left
behind at 9.2.0.

What was verified, and how

  • The whole procedure was followed end to end on a live 2026.1 cluster by an
    agent working only from these pages — full backup, incremental backup, then a
    destructive restore from the full plus one increment. Correctness was established
    with a marker written between the full and the increment, so its survival proves
    the increment was applied rather than merely that a restore happened. Verified on
    each node directly rather than through the load balancer.
  • The prune script was run against real archives, including reproducing the
    ordering bug an earlier revision had: it announced it was keeping a full because
    an increment referenced it, then cut that increment in a later pass. The shipped
    version removes the now-unreferenced full instead.
  • Every command in the guide has been run on that cluster.

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

✅MegaLinter analysis: Success

Descriptor Linter Files Fixed Errors Max errors Warnings Elapsed time
✅ ACTION actionlint 5 0 0 0.04s
✅ JSON jsonlint 4 0 0 0.12s
✅ JSON prettier 4 0 0 0.34s
✅ JSON v8r 4 0 0 10.48s
✅ MARKDOWN markdownlint 175 0 0 2.95s
✅ MARKDOWN markdown-table-formatter 175 0 0 0.37s
✅ REPOSITORY betterleaks yes no no 0.83s
✅ REPOSITORY checkov yes no no 19.98s
✅ REPOSITORY git_diff yes no no 0.06s
✅ REPOSITORY secretlint yes no no 1.71s
✅ REPOSITORY trufflehog yes no no 4.94s
✅ SPELL codespell 185 0 0 0.76s
✅ SPELL lychee 185 0 0 18.68s
✅ YAML prettier 6 0 0 0.41s
✅ YAML v8r 6 0 0 8.35s
✅ YAML yamllint 6 0 0 0.53s

See detailed reports in MegaLinter artifacts

Your project could benefit from a custom flavor, which would allow you to run only the linters you need, and thus improve runtime performances. (Skip this info by defining FLAVOR_SUGGESTIONS: false)

  • Documentation: Custom Flavors
  • Command: npx mega-linter-runner@10.1.0 --custom-flavor-setup --custom-flavor-linters ACTION_ACTIONLINT,JSON_JSONLINT,JSON_V8R,JSON_PRETTIER,MARKDOWN_MARKDOWNLINT,MARKDOWN_MARKDOWN_TABLE_FORMATTER,REPOSITORY_CHECKOV,REPOSITORY_GIT_DIFF,REPOSITORY_BETTERLEAKS,REPOSITORY_SECRETLINT,REPOSITORY_TRUFFLEHOG,SPELL_LYCHEE,SPELL_CODESPELL,YAML_PRETTIER,YAML_YAMLLINT,YAML_V8R

MegaLinter is provided by OX Security
Show us your support by starring ⭐ the repository

@ideaship ideaship self-assigned this Sep 28, 2026
@ideaship
ideaship marked this pull request as ready for review September 29, 2026 08:14
@jklare
jklare self-requested a review September 29, 2026 08:15
Comment thread docs/guides/operations-guide/mariadb/connection-fix.md Outdated
Comment thread docs/guides/operations-guide/mariadb/connection-fix.md Outdated
Comment thread docs/guides/operations-guide/mariadb/connection-fix.md Outdated
Comment thread docs/guides/operations-guide/mariadb/restore.md Outdated
Comment thread docs/guides/operations-guide/mariadb/restore.md Outdated
Comment thread docs/guides/operations-guide/mariadb/restore.md Outdated
@jklare

jklare commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

LGTM, thanks for writing and testing this. I will give it a try before approving and merging.

The section has four pages.

index.md orients the reader and scopes the guide to a single-shard
deployment, which is what OSISM deploys by default.

backup.mdx covers which host holds the archives and how to find it,
taking full and incremental backups, confirming that a run succeeded,
copying the archives off the database hosts, and scheduling runs with a
prune script to bound what they accumulate. It leads with a defect: on
OSISM 10.2.0 and earlier, mariabackup locks one node and copies
another, so an archive can record a binary log position belonging to a
different node and restore cleanly to the wrong state.

connection-fix.md documents that defect, both where the internal API
address is a keepalived VIP and where it is a BGP anycast address
announced by every control node, how to tell whether a deployment is
exposed, the configuration overlay that corrects it, what to do about
archives taken before it was applied, and how to revert the overlay
once an upgrade carries the upstream fix.

restore.md is the procedure: prerequisites, extracting and preparing an
archive before the outage, stopping the cluster, replacing the data,
recovering the cluster from the restored node, verifying before
returning to service, and reclaiming the scratch space. It also covers
restoring to an incremental backup, and reconciling the database
against OVN, Ceph and the hypervisors afterwards, since a restore moves
only MariaDB and leaves every backing store holding resources the
database no longer knows about.

The procedure was followed end to end on a live 2026.1 cluster,
including a destructive restore from a full plus one increment.

infrastructure.md still carries the sections this replaces; the next
commit retires them.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Roger Luethi <luethi@osism.tech>
The Backup and Restore subsections under MariaDB are replaced by the
guide added in the previous commit. It covers the same ground in more
detail, and for the archive layout OSISM has used since 9.2.0.

Replace them with a pointer to the new section. The Recovery subsection
stays: bringing back a cluster that stopped with its data intact is a
routine operation rather than part of backing up or restoring, and it
is handled separately. So does the rest of the MariaDB material, since
creating a database and user is not part of backing one up either.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Roger Luethi <luethi@osism.tech>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: In review

Development

Successfully merging this pull request may close these issues.

3 participants