dorfteich/docs/operations/restore-runbook.md
Claude Fable 5 18239e2fa9
All checks were successful
CD / Deploy to Test (push) Successful in 12s
CD / Smoke tests against Test (push) Successful in 1m15s
CD / Promote to Int (push) Successful in 12s
CI / Lint, typecheck, test (push) Successful in 6m19s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Successful in 8m13s
CI / Import/export fidelity gate (push) Successful in 54s
CI / Lint, typecheck, test (pull_request) Successful in 6m20s
CI / Build container images (pull_request) Successful in 1m12s
CI / Auth e2e pack (pull_request) Successful in 8m24s
CI / Import/export fidelity gate (pull_request) Successful in 58s
CD / Build and push images (push) Successful in 22s
#221: offline update path incl. migrations, rehearsed with rollback
Adds docs/operations/update-runbook.md (obtain, verify by digest, back
up, apply, verify, roll back) with the migration behaviour stated
explicitly: a failed migration rolls back its own transaction but is
recorded in _prisma_migrations and blocks every further migrate deploy
(P3009) — including a re-deployed old image — until migrate resolve
--rolled-back; semantically irreversible migrations have exactly one way
back, the pre-update backup set. No rolling updates on a compose stage.
Rehearsed in the isolated environment of #220: regular update to a v2
image set, then a deliberate failed-update (P3018 division by zero,
schema change proven rolled back) with image-rollback-alone shown
insufficient and the documented recovery executed. Protocol:
docs/vs-nfd/98-update-rollback-protokoll.md. ADR 0024 decisions 5+6
recorded as executed; operations handbook and restore runbook updated;
plan checkbox P1-3 ticked.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AUtYMxwTCMHG9mVHnwbFg8
2026-07-31 17:26:50 +02:00

4.6 KiB
Raw Blame History

Restore runbook (ADR 0015, issue #87)

How to restore a Dorfteich stage from a nightly backup set — manually in an incident, automatically as the monthly drill. The same procedure doubles as the Prod relocation procedure: restore the latest set on the new host.

A restore set is one backup id YYYYMMDD-HHMMSS: db-<id>.dump (pg_dump -Fc) plus files-<id>.tar.gz (uploads + plugins volumes), written nightly by the backup sidecar onto the backups volume, with status.json describing the last run (deploy/monitoring.md).

Manual restore (incident / relocation)

On the stage host, from the stage directory (/srv/DOCKER/dorfteich-<stage>/):

  1. Pick the set. docker compose exec backup ls /backups — usually the id in status.jsonlastSuccess.backupId.
  2. Run the automated runbook: ./restore.sh <backup-id> (deploy/backup/restore.sh). It stops web/api/collab (the db stays up), resets the public schema and replays the dump with pg_restore (#288 — objects created after the backup do not survive) and unpacks the volume archive through the backup sidecar image, starts the stack, and polls /readyz.
  3. Verify: /readyz fully green, spot-check one page and one uploaded file in the browser.

Consistency model (ADR 0015): the volume archive is taken minutes after the dump — a page referencing a file uploaded in between shows a missing image, never corruption.

Restore is also the designated way back from an update whose migrations succeeded but must be undone — the pre-update backup set is step 3 of docs/operations/update-runbook.md (#221).

Relocation to a new host: provision the stage directory (compose + .env, deploy/stages.md), start only db and backup (docker compose up -d db backup), copy the set into the backups volume (docker run --rm -v <src> -v <project>_backups:/backups …), then steps 23.

In-app restore (issue #103)

With a Nextcloud target configured, Site Admins can restore without shell access: Admin → System → Backups → Restore lists local and remote sets; after a type-to-confirm prompt the backup sidecar orchestrates the whole restore (maintenance mode → download + verify → terminate connections → pg_restore + volume extract → api restart). Progress lands in restore-status.json next to status.json; the public GET /api/v1/backup/restore-status endpoint keeps answering while everything else serves 503 maintenance_mode. This runbook stays the disaster path for when the app itself is gone.

Fetching a set from Nextcloud (total loss)

When the host is gone but the off-host copies exist, rebuild from the Nextcloud bundle alone:

  1. Download dorfteich-backup-<id>.tar.gz from the configured Nextcloud folder (browser or curl -u <user>:<app-password> -O https://cloud.example.com/remote.php/dav/files/<user>/<folder>/dorfteich-backup-<id>.tar.gz).
  2. Unpack it: tar -xzf dorfteich-backup-<id>.tar.gzdb-<id>.dump, files-<id>.tar.gz, and a manifest.json describing the set.
  3. Copy the two artifacts into the (fresh) stack's backups volume — see Relocation to a new host above — and run ./restore.sh <id>.

Automated monthly drill (.gitea/workflows/drill.yml)

Runs on the 1st of each month — and on demand by pushing a drill-* tag (git tag drill-$(date +%s) && git push origin --tags; Gitea 1.22 has no workflow-dispatch button yet). It executes deploy/backup/drill.sh, which

  • reads the drilled stage's backups volume read-only (stage volumes are never touched — everything scratch lives under a unique dorfteich-drill-<timestamp> prefix and is removed afterwards),
  • restores the latest successful set into a throwaway Postgres + volumes using the same backup-image code path as restore.sh,
  • boots the api image against the result and checks: readyz database + migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content cache, a public API request answers 200, and one media file's bytes on the volume match its database row,
  • reports the outcome as a comment on the pinned Restore drills issue (#98), then tears the scratch environment down (also on failure).

Pre-go-live the drill restores the Test stage's set (DRILL_SOURCE_VOLUME: dorfteich-test_backups); at go-live (#89) point it at the Prod backups volume. Manual invocation on the stage host:

SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh

A drill failure means the current backup set is not restorable — treat it like a failed backup: check the sidecar logs and status.json, fix, and re-run the drill the same day.