dorfteich/docs/operations/restore-runbook.md
Claude Fable 5 d95c18e9e8
All checks were successful
CD / Build and push images (push) Successful in 1m5s
CD / Deploy to Test (push) Successful in 10s
CD / Smoke tests against Test (push) Successful in 1m7s
CD / Promote to Int (push) Successful in 11s
CI / Lint, typecheck, test (push) Successful in 3m18s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Successful in 5m14s
CI / Import/export fidelity gate (push) Successful in 45s
Automate the monthly restore drill with a scratch-stack workflow (#87)
New scheduled workflow (monthly + on demand) runs deploy/backup/drill.sh:
it reads the drilled stage's backups volume strictly read-only, restores
the latest successful set into a throwaway Postgres and volumes under a
unique drill prefix via the backup image's restore path, boots the api
against the result, and verifies readyz (database + migrations), row
counts, rendered content in the page cache, a public API request, and a
media byte-check against the attachments table — then tears everything
down, also on failure. Each run reports its outcome as a comment on the
pinned "Restore drills" issue (#98). docs/operations/restore-runbook.md
carries the manual procedure, which doubles as the Prod relocation path;
pre-go-live the drill restores the Test set (switch the source volume at
go-live, #89 — off-host fetch from the BASEL mirror stays with #84).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-11 20:58:06 +02:00

63 lines
3.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Restore runbook (ADR 0015, issue #87)
How to restore a Dorfteich stage from a nightly backup set — manually in an
incident, automatically as the monthly drill. The same procedure doubles as
the **Prod relocation procedure**: restore the latest set on the new host.
A restore set is one backup id `YYYYMMDD-HHMMSS`: `db-<id>.dump`
(`pg_dump -Fc`) plus `files-<id>.tar.gz` (uploads + plugins volumes),
written nightly by the `backup` sidecar onto the `backups` volume, with
`status.json` describing the last run (deploy/monitoring.md).
## Manual restore (incident / relocation)
On the stage host, from the stage directory (`/home/DOCKER/dorfteich-<stage>/`):
1. **Pick the set.** `docker compose exec backup ls /backups` — usually the
id in `status.json``lastSuccess.backupId`.
2. **Run the automated runbook:** `./restore.sh <backup-id>`
(`deploy/backup/restore.sh`). It stops `web`/`api`/`collab` (the db stays
up), replays the dump with `pg_restore --clean --if-exists` and unpacks
the volume archive through the backup sidecar image, starts the stack,
and polls `/readyz`.
3. **Verify:** `/readyz` fully green, spot-check one page and one uploaded
file in the browser.
Consistency model (ADR 0015): the volume archive is taken minutes after the
dump — a page referencing a file uploaded in between shows a missing image,
never corruption.
**Relocation to a new host:** provision the stage directory (compose +
`.env`, deploy/stages.md), start only `db` and `backup`
(`docker compose up -d db backup`), copy the set into the backups volume
(`docker run --rm -v <src> -v <project>_backups:/backups …`), then steps 23.
## Automated monthly drill (`.gitea/workflows/drill.yml`)
Runs on the 1st of each month (and on demand via _Run workflow_): it
executes `deploy/backup/drill.sh`, which
- reads the drilled stage's backups volume **read-only** (stage volumes are
never touched — everything scratch lives under a unique
`dorfteich-drill-<timestamp>` prefix and is removed afterwards),
- restores the latest successful set into a throwaway Postgres + volumes
using the same backup-image code path as `restore.sh`,
- boots the api image against the result and checks: readyz database +
migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content
cache, a public API request answers 200, and one media file's bytes on
the volume match its database row,
- reports the outcome as a comment on the pinned **Restore drills** issue
(#98), then tears the scratch environment down (also on failure).
Pre-go-live the drill restores the **Test** stage's set
(`DRILL_SOURCE_VOLUME: dorfteich-test_backups`); at go-live (#89) point it
at the Prod backups volume. Manual invocation on the stage host:
```sh
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
```
A drill failure means the current backup set is **not restorable** — treat
it like a failed backup: check the sidecar logs and `status.json`, fix, and
re-run the drill the same day.