dorfteich/docs/operations/restore-runbook.md
Claude Fable 5 6e99cb35fc
All checks were successful
Restore drill / Restore the latest backup into a scratch stack (push) Successful in 28s
CD / Build and push images (push) Successful in 1m4s
CD / Deploy to Test (push) Successful in 11s
CD / Smoke tests against Test (push) Successful in 1m9s
CD / Promote to Int (push) Successful in 12s
CI / Lint, typecheck, test (push) Successful in 3m19s
CI / Build container images (push) Has been skipped
CI / Import/export fidelity gate (push) Successful in 45s
CI / Auth e2e pack (push) Successful in 5m9s
Trigger the restore drill on demand via drill-* tags (#87)
Gitea 1.22 cannot dispatch workflows through the API or UI (that lands in
1.23), so pushing a drill-* tag is the on-demand path next to the monthly
schedule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-11 21:01:00 +02:00

64 lines
3.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Restore runbook (ADR 0015, issue #87)
How to restore a Dorfteich stage from a nightly backup set — manually in an
incident, automatically as the monthly drill. The same procedure doubles as
the **Prod relocation procedure**: restore the latest set on the new host.
A restore set is one backup id `YYYYMMDD-HHMMSS`: `db-<id>.dump`
(`pg_dump -Fc`) plus `files-<id>.tar.gz` (uploads + plugins volumes),
written nightly by the `backup` sidecar onto the `backups` volume, with
`status.json` describing the last run (deploy/monitoring.md).
## Manual restore (incident / relocation)
On the stage host, from the stage directory (`/home/DOCKER/dorfteich-<stage>/`):
1. **Pick the set.** `docker compose exec backup ls /backups` — usually the
id in `status.json``lastSuccess.backupId`.
2. **Run the automated runbook:** `./restore.sh <backup-id>`
(`deploy/backup/restore.sh`). It stops `web`/`api`/`collab` (the db stays
up), replays the dump with `pg_restore --clean --if-exists` and unpacks
the volume archive through the backup sidecar image, starts the stack,
and polls `/readyz`.
3. **Verify:** `/readyz` fully green, spot-check one page and one uploaded
file in the browser.
Consistency model (ADR 0015): the volume archive is taken minutes after the
dump — a page referencing a file uploaded in between shows a missing image,
never corruption.
**Relocation to a new host:** provision the stage directory (compose +
`.env`, deploy/stages.md), start only `db` and `backup`
(`docker compose up -d db backup`), copy the set into the backups volume
(`docker run --rm -v <src> -v <project>_backups:/backups …`), then steps 23.
## Automated monthly drill (`.gitea/workflows/drill.yml`)
Runs on the 1st of each month — and on demand by pushing a `drill-*` tag
(`git tag drill-$(date +%s) && git push origin --tags`; Gitea 1.22 has no
workflow-dispatch button yet). It executes `deploy/backup/drill.sh`, which
- reads the drilled stage's backups volume **read-only** (stage volumes are
never touched — everything scratch lives under a unique
`dorfteich-drill-<timestamp>` prefix and is removed afterwards),
- restores the latest successful set into a throwaway Postgres + volumes
using the same backup-image code path as `restore.sh`,
- boots the api image against the result and checks: readyz database +
migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content
cache, a public API request answers 200, and one media file's bytes on
the volume match its database row,
- reports the outcome as a comment on the pinned **Restore drills** issue
(#98), then tears the scratch environment down (also on failure).
Pre-go-live the drill restores the **Test** stage's set
(`DRILL_SOURCE_VOLUME: dorfteich-test_backups`); at go-live (#89) point it
at the Prod backups volume. Manual invocation on the stage host:
```sh
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
```
A drill failure means the current backup set is **not restorable** — treat
it like a failed backup: check the sidecar logs and `status.json`, fix, and
re-run the drill the same day.