chore: initial import of backup scripts and design docs
Existing borg-backup.sh/dump_db.sh, old/ reference scripts, and the brainstormed design spec for restore tooling + runbook.
This commit is contained in:
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,162 @@
|
||||
# Backup & Recovery System Design
|
||||
|
||||
Date: 2026-07-25
|
||||
|
||||
## Goal
|
||||
|
||||
Reliable, automated backup of `/home/srv/files/content` (including MariaDB
|
||||
running in Docker) to a local Borg repo + offsite Scaleway S3 mirror, plus a
|
||||
documented, scriptable recovery path so downtime is short and predictable.
|
||||
|
||||
## Current State
|
||||
|
||||
- `old/borg.sh`, `old/maria.sh` — scripts currently live on the server.
|
||||
- `borg-backup.sh`, `dump_db.sh` (repo root) — improved replacements, not yet
|
||||
deployed. Add: encrypted repo (`repokey-blake2` via `BORG_PASSCOMMAND`),
|
||||
flock-based locking, preflight checks, container health polling, atomic
|
||||
dump staging/swap, non-transactional table warnings, integrity checks
|
||||
(`borg check`, daily archive-only / weekly `--verify-data`), retry/timeout
|
||||
wrapping, log rotation.
|
||||
- No restore tooling and no scheduling exist yet.
|
||||
|
||||
`dump_db.sh` is kept as-is. `borg-backup.sh` had one functional change made
|
||||
during this design: it no longer stops the MariaDB container. `dump_db.sh`
|
||||
already produces a transactionally-consistent dump via
|
||||
`--single-transaction`, so a live logical dump is safe without stopping
|
||||
anything (this only covers InnoDB tables — the script already warns if it
|
||||
finds non-InnoDB tables). The remaining risk was borg also archiving
|
||||
MariaDB's raw on-disk data directory while it's being written to; that's
|
||||
resolved by excluding it from the archive entirely (drop a `.nobackup`
|
||||
marker file in it — `borg-backup.sh` already passes
|
||||
`--exclude-if-present .nobackup`), so only the logical dump ever gets
|
||||
backed up, and restore only ever needs the dump anyway. Net effect: daily
|
||||
backups run with **zero MariaDB downtime**. This exclusion step needs to be
|
||||
added to the one-time setup, documented in the runbook.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
cron (daily) → borg-backup.sh
|
||||
├─ dump_db.sh (mysqldump per-DB, atomic staging swap;
|
||||
│ container stays up throughout)
|
||||
├─ borg create (encrypted, zstd) → local repo
|
||||
│ (raw MariaDB data dir excluded via .nobackup;
|
||||
│ only the logical dump is archived)
|
||||
├─ borg prune + compact
|
||||
├─ borg check (daily fast / weekly full --verify-data)
|
||||
└─ rclone sync → Scaleway S3 (offsite mirror)
|
||||
|
||||
restore.sh (manual, run on demand during an incident)
|
||||
mode=full : fresh/broken server → extract full content dir from archive
|
||||
→ restore all DBs → start container
|
||||
mode=db : single database → restore one .sql from a chosen archive
|
||||
mode=file : single file/dir → borg extract path, no DB involved
|
||||
```
|
||||
|
||||
## Components
|
||||
|
||||
All at repo root, deployed to the server alongside the existing scripts.
|
||||
|
||||
### `borg-backup.sh` (existing, one change: no longer stops MariaDB)
|
||||
|
||||
Orchestrates the daily backup: dump → archive → prune → compact → check →
|
||||
offsite sync. The container is never stopped — the dump step alone
|
||||
produces a consistent backup, and the raw data directory is excluded from
|
||||
the archive. Already has locking, timeouts, health-check pings, and
|
||||
cleanup-on-exit logic (cleanup now only intervenes if the container
|
||||
happens to be down for an unrelated reason, restarting it defensively).
|
||||
|
||||
### `dump_db.sh` (existing, unchanged)
|
||||
|
||||
Runs inside `borg-backup.sh` step 1. Dumps each non-system database plus
|
||||
users/grants to a staging dir, verifies each dump file, then atomically
|
||||
swaps staging into place. Deployed at
|
||||
`/home/srv/files/content/mariadb/dump_db.sh`.
|
||||
|
||||
### `restore.sh` (new)
|
||||
|
||||
```
|
||||
restore.sh full [--archive NAME] [--force]
|
||||
restore.sh db <db_name> [--archive NAME]
|
||||
restore.sh file <path-within-target> [--archive NAME] [--dest DIR]
|
||||
restore.sh --dry-run <above args>
|
||||
restore.sh --list-archives
|
||||
```
|
||||
|
||||
- Defaults `--archive` to the most recent archive in the repo.
|
||||
- `full`: refuses to run if `$TARGET` is non-empty unless `--force` is
|
||||
passed (protects against overwriting a working system by accident).
|
||||
Extracts the full archive over `$TARGET`, then restores every `.sql`
|
||||
dump found under `mariadb/dump/` using root DB credentials, then starts
|
||||
the `mariadb` container and waits for it to become healthy.
|
||||
- `db <name>`: extracts only `mariadb/dump/<name>.sql` from the archive,
|
||||
prompts for the database name to be typed again as confirmation (it's
|
||||
destructive — drops/recreates the DB), restores it with root credentials.
|
||||
- `file <path>`: `borg extract` of a single path from the archive into
|
||||
`--dest` (default: a scratch dir under `/tmp`), no DB or container
|
||||
involvement — safe, non-destructive.
|
||||
- `--dry-run`: prints exactly what would run (archive chosen, paths
|
||||
extracted, commands) without touching anything.
|
||||
- Same logging conventions as `borg-backup.sh` (`/var/log/borg/restore-*.log`),
|
||||
non-zero exit on any failure, no partial/silent restores — a failed step
|
||||
aborts before touching the "good" copy where possible.
|
||||
- Uses the same `BORG_PASSCOMMAND` / passphrase file as backups.
|
||||
- Restore DB operations authenticate as `root` via `MYSQL_ROOT_PASSWORD` or
|
||||
a root password file (mirrors how `dump_db.sh` resolves credentials),
|
||||
since restore needs full privileges (CREATE/DROP) that the limited
|
||||
`backup` user used for dumping does not have.
|
||||
|
||||
### `RUNBOOK.md` (new)
|
||||
|
||||
Single operational document covering:
|
||||
|
||||
1. **One-time setup** (run once, by hand, not scripted): `borg init
|
||||
--encryption=repokey-blake2`, create `/root/.borg-passphrase`
|
||||
(chmod 600), create the `backup` MariaDB user with the grants
|
||||
`dump_db.sh` needs (SELECT, LOCK TABLES, SHOW VIEW, TRIGGER, PROCESS,
|
||||
RELOAD) and its password file `/root/.mariadb-backup.pw`, drop a
|
||||
`.nobackup` marker file into MariaDB's Docker volume data directory
|
||||
(excludes raw DB files from the archive — only the logical dump under
|
||||
`mariadb/dump/` gets backed up), configure `rclone` for the `scaleway`
|
||||
remote, set the real healthcheck URL.
|
||||
2. **Deploying the scripts**: where each file goes, permissions,
|
||||
`chmod +x`.
|
||||
3. **Scheduling**: cron entry (daily, off-peak hours) invoking
|
||||
`borg-backup.sh`, plus how to check the last run (`/var/log/borg/`,
|
||||
healthcheck dashboard).
|
||||
4. **Day-2 operations**: listing archives, checking repo size/health,
|
||||
rotating the passphrase, what to do if the lockfile is stale.
|
||||
5. **Recovery procedures** — one clearly-numbered walkthrough per
|
||||
scenario (full disaster, single DB, single file), written so someone
|
||||
unfamiliar with the internals can follow it step-by-step under
|
||||
pressure, including exact `restore.sh` invocations and what to verify
|
||||
afterward (container healthy, app responds, row counts sane).
|
||||
6. **Restore drill cadence**: recommend a quarterly test restore into a
|
||||
scratch location to confirm backups are actually usable.
|
||||
|
||||
## Error Handling & Safety
|
||||
|
||||
- `restore.sh` never overwrites live data without an explicit `--force`
|
||||
(full) or typed confirmation (db).
|
||||
- All destructive steps are logged before execution.
|
||||
- `--dry-run` available for every mode.
|
||||
- Failures abort immediately (`set -euo pipefail`), matching the existing
|
||||
scripts' style; partial state is called out explicitly in the log.
|
||||
|
||||
## Testing
|
||||
|
||||
No automated test framework — these are ops scripts against real
|
||||
infrastructure. Verification consists of:
|
||||
|
||||
- `restore.sh --dry-run` runs against the real repo to sanity-check
|
||||
archive selection and command construction.
|
||||
- Quarterly real restore drill into a scratch directory (documented in
|
||||
the runbook), confirming dumps are complete and importable.
|
||||
|
||||
## Out of Scope
|
||||
|
||||
- No changes to `dump_db.sh`, and no changes to `borg-backup.sh` beyond
|
||||
removing the DB stop/start (see Components above).
|
||||
- No monitoring/alerting beyond the existing healthcheck ping.
|
||||
- No multi-server / multi-target generalization — this is specific to
|
||||
`/home/srv/files/content`.
|
||||
Reference in New Issue
Block a user