dorfteich/docs/operations/restore-runbook.md
Claude Fable 5 d95c18e9e8
All checks were successful
CD / Build and push images (push) Successful in 1m5s
CD / Deploy to Test (push) Successful in 10s
CD / Smoke tests against Test (push) Successful in 1m7s
CD / Promote to Int (push) Successful in 11s
CI / Lint, typecheck, test (push) Successful in 3m18s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Successful in 5m14s
CI / Import/export fidelity gate (push) Successful in 45s
Automate the monthly restore drill with a scratch-stack workflow (#87)
New scheduled workflow (monthly + on demand) runs deploy/backup/drill.sh:
it reads the drilled stage's backups volume strictly read-only, restores
the latest successful set into a throwaway Postgres and volumes under a
unique drill prefix via the backup image's restore path, boots the api
against the result, and verifies readyz (database + migrations), row
counts, rendered content in the page cache, a public API request, and a
media byte-check against the attachments table — then tears everything
down, also on failure. Each run reports its outcome as a comment on the
pinned "Restore drills" issue (#98). docs/operations/restore-runbook.md
carries the manual procedure, which doubles as the Prod relocation path;
pre-go-live the drill restores the Test set (switch the source volume at
go-live, #89 — off-host fetch from the BASEL mirror stays with #84).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-11 20:58:06 +02:00

3.0 KiB
Raw Blame History

Restore runbook (ADR 0015, issue #87)

How to restore a Dorfteich stage from a nightly backup set — manually in an incident, automatically as the monthly drill. The same procedure doubles as the Prod relocation procedure: restore the latest set on the new host.

A restore set is one backup id YYYYMMDD-HHMMSS: db-<id>.dump (pg_dump -Fc) plus files-<id>.tar.gz (uploads + plugins volumes), written nightly by the backup sidecar onto the backups volume, with status.json describing the last run (deploy/monitoring.md).

Manual restore (incident / relocation)

On the stage host, from the stage directory (/home/DOCKER/dorfteich-<stage>/):

  1. Pick the set. docker compose exec backup ls /backups — usually the id in status.jsonlastSuccess.backupId.
  2. Run the automated runbook: ./restore.sh <backup-id> (deploy/backup/restore.sh). It stops web/api/collab (the db stays up), replays the dump with pg_restore --clean --if-exists and unpacks the volume archive through the backup sidecar image, starts the stack, and polls /readyz.
  3. Verify: /readyz fully green, spot-check one page and one uploaded file in the browser.

Consistency model (ADR 0015): the volume archive is taken minutes after the dump — a page referencing a file uploaded in between shows a missing image, never corruption.

Relocation to a new host: provision the stage directory (compose + .env, deploy/stages.md), start only db and backup (docker compose up -d db backup), copy the set into the backups volume (docker run --rm -v <src> -v <project>_backups:/backups …), then steps 23.

Automated monthly drill (.gitea/workflows/drill.yml)

Runs on the 1st of each month (and on demand via Run workflow): it executes deploy/backup/drill.sh, which

  • reads the drilled stage's backups volume read-only (stage volumes are never touched — everything scratch lives under a unique dorfteich-drill-<timestamp> prefix and is removed afterwards),
  • restores the latest successful set into a throwaway Postgres + volumes using the same backup-image code path as restore.sh,
  • boots the api image against the result and checks: readyz database + migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content cache, a public API request answers 200, and one media file's bytes on the volume match its database row,
  • reports the outcome as a comment on the pinned Restore drills issue (#98), then tears the scratch environment down (also on failure).

Pre-go-live the drill restores the Test stage's set (DRILL_SOURCE_VOLUME: dorfteich-test_backups); at go-live (#89) point it at the Prod backups volume. Manual invocation on the stage host:

SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh

A drill failure means the current backup set is not restorable — treat it like a failed backup: check the sidecar logs and status.json, fix, and re-run the drill the same day.