All checks were successful
CD / Build and push images (push) Successful in 1m5s
CD / Deploy to Test (push) Successful in 10s
CD / Smoke tests against Test (push) Successful in 1m7s
CD / Promote to Int (push) Successful in 11s
CI / Lint, typecheck, test (push) Successful in 3m18s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Successful in 5m14s
CI / Import/export fidelity gate (push) Successful in 45s
New scheduled workflow (monthly + on demand) runs deploy/backup/drill.sh: it reads the drilled stage's backups volume strictly read-only, restores the latest successful set into a throwaway Postgres and volumes under a unique drill prefix via the backup image's restore path, boots the api against the result, and verifies readyz (database + migrations), row counts, rendered content in the page cache, a public API request, and a media byte-check against the attachments table — then tears everything down, also on failure. Each run reports its outcome as a comment on the pinned "Restore drills" issue (#98). docs/operations/restore-runbook.md carries the manual procedure, which doubles as the Prod relocation path; pre-go-live the drill restores the Test set (switch the source volume at go-live, #89 — off-host fetch from the BASEL mirror stays with #84). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
63 lines
3.0 KiB
Markdown
63 lines
3.0 KiB
Markdown
# Restore runbook (ADR 0015, issue #87)
|
||
|
||
How to restore a Dorfteich stage from a nightly backup set — manually in an
|
||
incident, automatically as the monthly drill. The same procedure doubles as
|
||
the **Prod relocation procedure**: restore the latest set on the new host.
|
||
|
||
A restore set is one backup id `YYYYMMDD-HHMMSS`: `db-<id>.dump`
|
||
(`pg_dump -Fc`) plus `files-<id>.tar.gz` (uploads + plugins volumes),
|
||
written nightly by the `backup` sidecar onto the `backups` volume, with
|
||
`status.json` describing the last run (deploy/monitoring.md).
|
||
|
||
## Manual restore (incident / relocation)
|
||
|
||
On the stage host, from the stage directory (`/home/DOCKER/dorfteich-<stage>/`):
|
||
|
||
1. **Pick the set.** `docker compose exec backup ls /backups` — usually the
|
||
id in `status.json` → `lastSuccess.backupId`.
|
||
2. **Run the automated runbook:** `./restore.sh <backup-id>`
|
||
(`deploy/backup/restore.sh`). It stops `web`/`api`/`collab` (the db stays
|
||
up), replays the dump with `pg_restore --clean --if-exists` and unpacks
|
||
the volume archive through the backup sidecar image, starts the stack,
|
||
and polls `/readyz`.
|
||
3. **Verify:** `/readyz` fully green, spot-check one page and one uploaded
|
||
file in the browser.
|
||
|
||
Consistency model (ADR 0015): the volume archive is taken minutes after the
|
||
dump — a page referencing a file uploaded in between shows a missing image,
|
||
never corruption.
|
||
|
||
**Relocation to a new host:** provision the stage directory (compose +
|
||
`.env`, deploy/stages.md), start only `db` and `backup`
|
||
(`docker compose up -d db backup`), copy the set into the backups volume
|
||
(`docker run --rm -v <src> -v <project>_backups:/backups …`), then steps 2–3.
|
||
|
||
## Automated monthly drill (`.gitea/workflows/drill.yml`)
|
||
|
||
Runs on the 1st of each month (and on demand via _Run workflow_): it
|
||
executes `deploy/backup/drill.sh`, which
|
||
|
||
- reads the drilled stage's backups volume **read-only** (stage volumes are
|
||
never touched — everything scratch lives under a unique
|
||
`dorfteich-drill-<timestamp>` prefix and is removed afterwards),
|
||
- restores the latest successful set into a throwaway Postgres + volumes
|
||
using the same backup-image code path as `restore.sh`,
|
||
- boots the api image against the result and checks: readyz database +
|
||
migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content
|
||
cache, a public API request answers 200, and one media file's bytes on
|
||
the volume match its database row,
|
||
- reports the outcome as a comment on the pinned **Restore drills** issue
|
||
(#98), then tears the scratch environment down (also on failure).
|
||
|
||
Pre-go-live the drill restores the **Test** stage's set
|
||
(`DRILL_SOURCE_VOLUME: dorfteich-test_backups`); at go-live (#89) point it
|
||
at the Prod backups volume. Manual invocation on the stage host:
|
||
|
||
```sh
|
||
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
|
||
```
|
||
|
||
A drill failure means the current backup set is **not restorable** — treat
|
||
it like a failed backup: check the sidecar logs and `status.json`, fix, and
|
||
re-run the drill the same day.
|