dorfteich/docs/operations/restore-runbook.md
Claude Fable 5 5cef359b8f
All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Nextcloud backup target: admin-configured, manual + scheduled uploads, in-app restore (#103)
Off-host backups for every self-hoster, configured entirely in the admin
UI — supersedes the host-specific mirror plan behind #84.

shared:
- webdav.ts (new package entry like token-crypto): minimal WebDAV client
  with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT
  (streamed), GET, DELETE; Nextcloud DAV path derived from the plain
  server URL, explicit DAV bases pass through
- backup-status.ts: additive remote-upload status in status.json, the
  restore-status.json contract (running/succeeded/failed + staleness
  bound), the backup_command/backup_maintenance NOTIFY channels, and the
  one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz)
- backup-set.ts moved here from apps/backup (api lists local sets)

backup sidecar:
- reads the backup.* instance settings directly from the database (admin
  changes apply next run; local retention row overrides the env) and the
  app password from the secret store
- after each successful set: bundle dump + files archive + manifest into
  ONE self-contained tar.gz, upload via WebDAV per schedule
  (off/daily/weekly; manual runs always upload), prune remote bundles —
  never the newest — and record the outcome in status.json; upload
  failures alert via a new backupUploadFailed mail (de+en)
- command listener on backup_command (run / restore) with a serial queue
  against the nightly timer
- restore orchestrator: restore-status.json → maintenance NOTIFY →
  grace → (remote: download + manifest-verify bundle) → terminate other
  DB connections → shared perform-restore path (same code as restore.sh)
  → final status + maintenance exit

api:
- MaintenanceGuard (global, registered before the setup gate): 503
  maintenance_mode while restore-status says running; health endpoints
  and the new public GET /backup/restore-status stay exempt; a stale
  running state (crashed sidecar) unblocks after 30 min
- MaintenanceStateService watches the file and restarts the api after a
  successful restore (fresh caches, migrate-on-start for older dumps);
  main.ts refuses to touch the database while a restore runs — a
  container restarting mid-restore must not race pg_restore with
  migrate deploy
- worker sweeps (conversion, mail outbox, scheduler) catch transient
  database failures instead of dying on an unhandled rejection — the
  restore's connection termination crashed the api in verification
- backup admin endpoints under /admin/system/backup: settings (live
  connection test before save, password write-only into the secret
  store), nextcloud/test, sets (local via the ro backups mount + remote
  via WebDAV), run + restore (type-to-confirm backstop, source
  validation) — commands travel as NOTIFY payloads; audit actions
  backup.settings_changed/run_triggered/restore_requested
- readyz: new warning-level backup_remote check while a target is
  configured (26 h daily / 170 h weekly bound)

collab:
- maintenance listener: on enter, persist + close every live session and
  refuse new connections until exit (failsafe timeout 30 min) — no
  in-memory document may write pre-restore content back afterwards

web:
- Admin → System backup section: status card with remote facts and a
  "Back up now" button, the Nextcloud settings form with test button,
  and the restore picker (local + remote sets, type-to-confirm)
- global maintenance screen: any 503 maintenance_mode flips the SPA to a
  status page polling the exempt endpoint, reloading when the instance
  returns

Verified end-to-end against a live stack (fresh DB, native api + sidecar,
fake WebDAV server): configure → test → manual backup → bundle upload →
readyz/sets/status surfaces → remote restore with maintenance gate,
marker rollback and api restart; suites: shared 21, backup 9, collab 11,
api 58 files green, lint + i18n:check + typecheck clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-12 10:39:18 +02:00

4.3 KiB
Raw Blame History

Restore runbook (ADR 0015, issue #87)

How to restore a Dorfteich stage from a nightly backup set — manually in an incident, automatically as the monthly drill. The same procedure doubles as the Prod relocation procedure: restore the latest set on the new host.

A restore set is one backup id YYYYMMDD-HHMMSS: db-<id>.dump (pg_dump -Fc) plus files-<id>.tar.gz (uploads + plugins volumes), written nightly by the backup sidecar onto the backups volume, with status.json describing the last run (deploy/monitoring.md).

Manual restore (incident / relocation)

On the stage host, from the stage directory (/home/DOCKER/dorfteich-<stage>/):

  1. Pick the set. docker compose exec backup ls /backups — usually the id in status.jsonlastSuccess.backupId.
  2. Run the automated runbook: ./restore.sh <backup-id> (deploy/backup/restore.sh). It stops web/api/collab (the db stays up), replays the dump with pg_restore --clean --if-exists and unpacks the volume archive through the backup sidecar image, starts the stack, and polls /readyz.
  3. Verify: /readyz fully green, spot-check one page and one uploaded file in the browser.

Consistency model (ADR 0015): the volume archive is taken minutes after the dump — a page referencing a file uploaded in between shows a missing image, never corruption.

Relocation to a new host: provision the stage directory (compose + .env, deploy/stages.md), start only db and backup (docker compose up -d db backup), copy the set into the backups volume (docker run --rm -v <src> -v <project>_backups:/backups …), then steps 23.

In-app restore (issue #103)

With a Nextcloud target configured, Site Admins can restore without shell access: Admin → System → Backups → Restore lists local and remote sets; after a type-to-confirm prompt the backup sidecar orchestrates the whole restore (maintenance mode → download + verify → terminate connections → pg_restore + volume extract → api restart). Progress lands in restore-status.json next to status.json; the public GET /api/v1/backup/restore-status endpoint keeps answering while everything else serves 503 maintenance_mode. This runbook stays the disaster path for when the app itself is gone.

Fetching a set from Nextcloud (total loss)

When the host is gone but the off-host copies exist, rebuild from the Nextcloud bundle alone:

  1. Download dorfteich-backup-<id>.tar.gz from the configured Nextcloud folder (browser or curl -u <user>:<app-password> -O https://cloud.example.com/remote.php/dav/files/<user>/<folder>/dorfteich-backup-<id>.tar.gz).
  2. Unpack it: tar -xzf dorfteich-backup-<id>.tar.gzdb-<id>.dump, files-<id>.tar.gz, and a manifest.json describing the set.
  3. Copy the two artifacts into the (fresh) stack's backups volume — see Relocation to a new host above — and run ./restore.sh <id>.

Automated monthly drill (.gitea/workflows/drill.yml)

Runs on the 1st of each month — and on demand by pushing a drill-* tag (git tag drill-$(date +%s) && git push origin --tags; Gitea 1.22 has no workflow-dispatch button yet). It executes deploy/backup/drill.sh, which

  • reads the drilled stage's backups volume read-only (stage volumes are never touched — everything scratch lives under a unique dorfteich-drill-<timestamp> prefix and is removed afterwards),
  • restores the latest successful set into a throwaway Postgres + volumes using the same backup-image code path as restore.sh,
  • boots the api image against the result and checks: readyz database + migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content cache, a public API request answers 200, and one media file's bytes on the volume match its database row,
  • reports the outcome as a comment on the pinned Restore drills issue (#98), then tears the scratch environment down (also on failure).

Pre-go-live the drill restores the Test stage's set (DRILL_SOURCE_VOLUME: dorfteich-test_backups); at go-live (#89) point it at the Prod backups volume. Manual invocation on the stage host:

SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh

A drill failure means the current backup set is not restorable — treat it like a failed backup: check the sidecar logs and status.json, fix, and re-run the drill the same day.